Courseiva

Google Cloud Generative AI Leader Generative AI Leader (Generative AI Leader) — Questions 151–225

1008 questions total · 14pages · All types, answers revealed

Page 2

Page 3 of 14

Page 4
151
MCQeasy

A developer wants to quickly prototype a multimodal application that can process images and text using the Gemini API without incurring any cost. Which access tier should they use?

A.Colab Enterprise
B.Cloud Shell
C.Vertex AI Gemini API
D.Google AI Studio
AnswerD

Google AI Studio provides free-tier access to the Gemini API for prototyping multimodal image and text applications, satisfying the stem's no-cost constraint. Vertex AI and paid Gemini API tiers require billing, so they fail the zero-cost requirement.

Why this answer

Google AI Studio (option D) is the correct choice because it provides a free, browser-based environment specifically designed for prototyping with the Gemini API, including multimodal capabilities for images and text, without requiring any payment or cloud billing setup. This makes it ideal for quick experimentation and cost-free development.

Exam trap

The trap here is that candidates often confuse Google AI Studio with Vertex AI Gemini API, assuming both require billing, or they overlook that Colab Enterprise and Cloud Shell are not designed for direct, cost-free Gemini API prototyping.

How to eliminate wrong answers

Option A is wrong because Colab Enterprise is a paid Google Cloud service that requires billing activation and is intended for production-grade notebook environments, not for free prototyping. Option B is wrong because Cloud Shell provides a command-line interface and basic cloud resources but does not natively offer the Gemini API or a multimodal prototyping interface without additional setup and potential costs. Option C is wrong because the Vertex AI Gemini API is a paid enterprise service that requires a Google Cloud project with billing enabled, and while it offers the same API, it is not cost-free for prototyping.

152
MCQeasy

To ensure that a generative AI model uses the most current information from the web for answering user queries, which Vertex AI feature should be enabled?

A.Grounding with Google Search
B.Safety filters
C.Context caching
D.Model tuning
AnswerA

Grounding with Google Search connects the model to live web results at query time, so answers reflect current information rather than the training cutoff. This satisfies the freshness requirement without retraining or fine-tuning the underlying model.

Why this answer

Grounding with Google Search is the correct feature because it enables the model to retrieve and reference real-time information from the web, ensuring responses are based on the most current data available. This is achieved by integrating Google Search results directly into the model's generation process, allowing it to cite live sources and reduce hallucinations from outdated training data.

Exam trap

Google Cloud often tests the distinction between features that improve output quality through external data retrieval (Grounding) versus those that modify the model's internal behavior (tuning, caching, filtering), leading candidates to confuse safety or optimization features with live data access.

How to eliminate wrong answers

Option B is wrong because safety filters are designed to block harmful or inappropriate content, not to fetch current web information. Option C is wrong because context caching stores frequently accessed context to reduce latency and cost, but it does not provide live web data. Option D is wrong because model tuning adjusts the model's parameters on a specific dataset to improve performance on a task, but it does not enable real-time web retrieval.

153
MCQmedium

A company wants to measure the ROI of a GenAI-based report generation tool. Which metric is MOST directly tied to business value?

A.Model accuracy on a test set
B.Number of API calls made per month
C.Percentage of reports that require human editing
D.Average time saved per report by analysts
AnswerD

Analyst time saved per report translates directly into labour cost avoided, the clearest financial proxy for business value. It ties GenAI output to measurable productivity gains, unlike technical metrics such as token latency or model accuracy.

Why this answer

The average time saved per report by analysts is the metric most directly tied to business value because it quantifies the productivity gain from the GenAI tool. Time savings can be translated into cost reduction or increased capacity, making it a clear ROI indicator. Other metrics like model accuracy or API calls are technical or operational, not direct business outcomes.

Exam trap

The trap is selecting technical metrics like model accuracy or API calls as ROI indicators, when the exam expects a business-centric metric such as time saved that directly impacts cost or productivity.

How to eliminate wrong answers

Option A is wrong because model accuracy on a test set is a technical performance metric; it does not directly measure business value or ROI, as a highly accurate model may still not save time or money. Option B is wrong because the number of API calls per month is an operational usage metric, not a business outcome; it indicates activity but not value generated. Option C is wrong because the percentage of reports requiring human editing is a quality metric that affects efficiency, but it is one step removed from direct business value compared to time saved, which directly impacts labor costs and throughput.

154
Multi-Selecteasy

A mid-size accounting firm wants to deploy a generative AI assistant that summarizes client meeting notes and drafts follow-up emails. The partners require that client financial data never leaves the firm's Google Cloud project boundary for third-party training, and that usage costs stay predictable month to month. Which two Google Cloud practices should the firm adopt? (Choose two.)

Select 2 answers
A.Disable all Cloud Logging on the assistant so that prompt and response content is never recorded anywhere.
B.Set Cloud Billing budgets and alerts on the project, and monitor token consumption per assistant feature.
C.Send every prompt to a third-party public chatbot API outside Google Cloud because it offers a free usage tier.
D.Publish the meeting notes to a public Cloud Storage bucket so the assistant can retrieve them without authentication.
E.Use Gemini through Vertex AI under Google Cloud's data governance terms, which keep customer prompts and responses within the project and out of foundation model training.
AnswersB, E

Budgets and alerts notify the firm when spending approaches a defined threshold, and tracking token consumption per feature shows which assistant capabilities drive cost. Together they give the partners the month-to-month predictability they asked for and enable early corrective action.

Why this answer

Running Gemini through Vertex AI keeps prompts and responses inside the firm's project under Google Cloud data governance, so client financial data is not used for foundation model training. Pairing that with budgets, alerts, and per-feature token monitoring delivers the cost predictability the partners require, and neither practice adds infrastructure the firm must operate.

Exam trap

The trap here is assuming that any generative AI endpoint is equally safe and predictable, when data governance terms and cost controls depend entirely on which service and configuration the firm chooses.

155
Multi-Selectmedium

A product team is using a generative AI model to summarize customer feedback from multiple sources. The summaries are sometimes missing key themes or including irrelevant details. The team wants to improve the quality of the summaries without retraining the model. (Choose two.)

Select 2 answers
A.Provide a structured prompt that specifies the sections to include, such as top themes, sentiment, and representative quotes.
B.Reduce the maximum output tokens so the model is forced to include only the most important points.
C.Increase the temperature to encourage the model to explore more diverse themes.
D.Use few-shot examples that show correctly formatted summaries with the desired level of detail.
E.Fine-tune the model on a dataset of customer feedback summaries.
AnswersA, D

A structured prompt tells the model exactly what to cover, reducing omissions of key themes and irrelevant additions. By defining sections like top themes, sentiment, and quotes, the team guides the model's focus. This is a prompt-engineering technique that requires no retraining. It directly addresses the missing-theme and irrelevant-detail problems by making the expected output explicit.

Why this answer

A structured prompt and few-shot examples both guide the model to include the right themes and exclude irrelevant details. The structured prompt defines the required sections, while examples show the desired format and depth. Together they improve summarization quality without retraining, directly addressing the team's concerns about missing themes and extraneous content.

Exam trap

The trap here is thinking that output length limits or higher randomness can fix summarization quality, when the real need is explicit guidance on content and format.

156
MCQhard

A financial services firm is deploying a generative AI chatbot to answer employee questions about internal policies. The firm must ensure that the chatbot does not reveal sensitive information from other departments. Which technique should they implement?

A.Fine-tune the model on all internal documents to improve accuracy.
B.Set the temperature to a low value to reduce creative responses.
C.Increase the model's context window to include more documents.
D.Use retrieval-augmented generation with document-level access controls.
AnswerD

RAG can retrieve documents from a source that enforces access permissions based on the user's identity. By integrating with an identity-aware search or filtering retrieved documents by department, the chatbot only supplies context the user is authorized to see. This prevents cross-department information leakage while still providing accurate answers.

Why this answer

Retrieval-augmented generation can be combined with access control lists so that the retrieval step only returns documents the requesting user is permitted to view. This ensures the model's context is limited to authorized information, preventing leakage across departments. Fine-tuning or context expansion without access controls would not solve the security requirement.

Exam trap

The trap here is thinking that fine-tuning or larger context windows can enforce security, when in fact access control must be applied at the retrieval layer.

157
MCQmedium

A global retailer's legal team is reviewing a proposed generative AI solution that drafts personalized marketing copy for customers in the EU. The team wants to confirm whether the solution can meet the EU AI Act's transparency obligations for AI-generated content while still using Google Cloud managed services. Which approach best satisfies the transparency requirement?

A.Log all prompts and model responses in Cloud Logging and retain the logs for twelve months.
B.Require human review of every generated marketing message before it is sent to customers.
C.Use SynthID watermarking on generated images and add clear AI-generated labels to text and media outputs.
D.Restrict the solution to internal employees so that no customer-facing AI-generated content is produced.
AnswerC

SynthID embeds imperceptible watermarks in AI-generated images and other media, and clear labels inform users that content is AI-generated. Together these mechanisms directly address the EU AI Act's transparency expectations for synthetic content, while remaining compatible with managed Google Cloud generative AI services such as Vertex AI and Gemini models.

Why this answer

The EU AI Act requires that AI-generated content be clearly identifiable as such. SynthID watermarking provides a durable, machine-detectable signal embedded in generated media, while explicit labels inform human readers. Using these together on Vertex AI and Gemini outputs delivers the required transparency without blocking the personalization workflow or forcing impractical manual review of every message.

Exam trap

The trap here is assuming that logging or human review alone equals regulatory transparency, when the obligation is specifically to disclose AI-generated content to the audience.

158
Multi-Selectmedium

A media company is evaluating Google Cloud generative AI offerings to build a production application that summarizes long articles and generates headlines. They want to use Gemini models with enterprise controls and need to understand which capabilities are provided by Vertex AI. (Choose two.)

Select 2 answers
A.Automated generation of print-ready magazine layouts from article text.
B.Access to Gemini models through a managed API with enterprise security and data governance controls.
C.Guaranteed placement of generated headlines on the publication's homepage.
D.Built-in tools for tuning and evaluating Gemini models against the company's own summarization quality criteria.
E.Automatic translation of every generated article into all supported languages without configuration.
AnswersB, D

Vertex AI provides managed access to Gemini models with enterprise-grade security, identity, and data governance controls, which is essential for a production media application handling proprietary content. It supports VPC Service Controls, customer-managed encryption keys, and audit logging. This capability aligns with the requirement to use Gemini under enterprise controls rather than through a consumer interface.

Why this answer

Vertex AI delivers managed access to Gemini with enterprise security and governance, plus tuning and evaluation tools to align models with custom summarization criteria. These two capabilities directly support a production media application that needs controlled model access and measurable quality. The remaining options describe publishing outcomes or unrealistic automation that Vertex AI does not provide.

Exam trap

The trap here is confusing content generation with content publishing, assuming Vertex AI also handles layout, homepage placement, or configuration-free translation.

159
MCQeasy

A developer is using Vertex AI Studio to test prompts for a text generation model. They want the model to follow a specific output format (JSON). Which prompt engineering approach is most effective?

A.Set stop sequences to '}'.
B.Include a few-shot example of the exact JSON format in the prompt.
C.Set the system instruction to 'Always output JSON.'
D.Set temperature to 0 to make output deterministic.
AnswerB

Few-shot prompting supplies concrete input-output pairs demonstrating the exact JSON schema, so the model infers field names, nesting and types rather than guessing. This constrains generation to the required structure more reliably than describing the format in prose alone.

Why this answer

Including a few-shot example of the exact JSON format in the prompt provides the model with a concrete pattern to follow, which is the most reliable method for enforcing structured output in generative models. Few-shot prompting leverages in-context learning, where the model uses the provided example to infer the desired schema and formatting rules, reducing ambiguity and improving adherence to the specified JSON structure.

Exam trap

Google Cloud often tests the misconception that system instructions or hyperparameter tuning alone can enforce output format, when in practice, few-shot examples are the most direct and reliable method for guiding model behavior in structured generation tasks.

How to eliminate wrong answers

Option A is wrong because setting stop sequences to '}' would prematurely terminate generation at the first closing brace, which may cut off nested JSON objects or arrays, and does not guarantee the model outputs valid JSON from the start. Option C is wrong because a system instruction like 'Always output JSON' is a high-level directive that models often fail to follow precisely without explicit formatting examples, as they may still produce markdown, extra text, or malformed JSON. Option D is wrong because setting temperature to 0 makes output deterministic but does not enforce a specific output format; the model could still generate non-JSON text or deviate from the required schema, as temperature controls randomness, not structure.

160
MCQmedium

A developer is building a code‑generation assistant using the Codey API on Vertex AI. The assistant should generate Python functions based on natural language descriptions. However, the generated code sometimes contains syntax errors. Which parameter adjustment would MOST directly help reduce syntax errors?

A.Lower the temperature (e.g., from 0.8 to 0.2)
B.Increase the context window
C.Set top-k to 1
D.Increase the max output tokens
AnswerA

Lowering the temperature sharpens the model's token probability distribution, so sampling favours high-confidence tokens rather than risky alternatives. This directly reduces the erratic token choices that produce malformed Python syntax, satisfying the stem's requirement to cut syntax errors without altering the prompt or model.

Why this answer

Reducing temperature makes the model more deterministic, which typically reduces creative but incorrect outputs like syntax errors. Prompt engineering can also help, but adjusting temperature is the simplest direct fix. Increasing max tokens or changing top-k does not directly address syntax correctness.

161
MCQmedium

A company is evaluating Google Cloud vs AWS Bedrock for building a multimodal application that needs to understand images, video, and text in a single model. Which unique Google Cloud capability supports this requirement?

A.Gemini API's native multimodal support across text, images, video, and audio
B.Use of Hugging Face models on Vertex AI
C.Integration with Amazon Bedrock's Titan model
D.BigQuery ML's image analysis functions
AnswerA

The Gemini API natively accepts text, images, video and audio within a single model, so one request can reason across mixed media without separate pipelines. Bedrock composes modality-specific models, lacking this unified native multimodal capability.

Why this answer

The Gemini API on Google Cloud offers native multimodal support, meaning a single model can process and reason across text, images, video, and audio without stitching together separate models. This is a distinguishing capability for building applications that must understand multiple modalities in one model. It directly satisfies the requirement for a multimodal application handling images, video, and text.

Exam trap

The trap is picking a generic 'use models on Vertex AI' answer or an AWS option — the exam tests that Gemini's native multimodality across text, image, video, and audio is the unique Google Cloud capability for unified multimodal applications.

How to eliminate wrong answers

Option B is wrong because using Hugging Face models on Vertex AI typically involves deploying separate task-specific models, not a single natively multimodal model, and it is not a unique Google Cloud capability in the way Gemini is. Option C is wrong because Amazon Bedrock's Titan model is an AWS offering, not a Google Cloud capability, and the question asks for a Google Cloud differentiator. Option D is wrong because BigQuery ML's image analysis functions are limited to specific ML tasks on data in BigQuery and do not provide a unified multimodal model for images, video, and text.

162
MCQhard

A team is using Vertex AI Pipelines to deploy a generative AI model for real-time inference. The model sometimes generates harmful content. They want to implement a safety filter that checks the output before returning it to the user, but they need to minimize latency. Which approach best balances safety and performance?

A.Use a secondary lightweight classifier to filter outputs in real-time.
B.Retrain the model on every flagged harmful output.
C.Manually review all outputs before delivery.
D.Disable safety checks to improve latency.
AnswerA

A lightweight secondary classifier inspects each generated output and blocks harmful content before it reaches the user. Its low compute overhead adds minimal latency compared with a full model-based filter, satisfying the stem's need to balance safety against real-time inference performance.

Why this answer

Deploying a secondary lightweight classifier (e.g., a distilled BERT or a small logistic regression model) as a post-processing filter allows real-time inference with minimal latency overhead. This approach decouples safety from the primary generative model, enabling fast rejection of harmful outputs without retraining or blocking the main inference pipeline.

Exam trap

Google Cloud often tests the misconception that safety must be integrated into the generative model itself (e.g., via retraining or fine-tuning), when in practice a separate, lightweight post-processing filter is the standard for low-latency production systems.

How to eliminate wrong answers

Option B is wrong because retraining the model on every flagged harmful output is computationally expensive, introduces significant latency, and can lead to catastrophic forgetting or overfitting to specific examples, making it impractical for real-time inference. Option C is wrong because manual review of all outputs introduces unacceptable latency and does not scale, violating the requirement to minimize latency. Option D is wrong because disabling safety checks entirely eliminates the safety requirement, which is explicitly needed, and would expose users to harmful content, failing the core objective.

163
Multi-Selectmedium

A data science team is building a document question-answering assistant on Vertex AI. Users report that answers are sometimes fabricated when the retrieved passages do not contain the answer. Which TWO techniques should the team apply to reduce hallucinations? (Choose two.)

Select 2 answers
A.Raise the maximum output token limit so the model has more space to explain its reasoning and avoid mistakes.
B.Instruct the model to answer only from the provided context and to respond with a fixed phrase such as 'not found in the provided documents' when the context is insufficient.
C.Increase the temperature so the model explores alternative answers when the retrieved context seems incomplete.
D.Return citations to the specific retrieved passages used and require the answer to reference them.
E.Fine-tune the model on the full document corpus so all answers are stored in the model weights.
AnswersB, D

An explicit grounding instruction with a defined abstention phrase gives the model a safe path when evidence is missing. It reduces fabrication because the model is told to prefer the supplied passages and to admit when they do not contain the answer. This is a low-cost prompt change that directly targets unsupported claims.

Why this answer

Grounding instructions with an abstention path and mandatory citations both push the model to rely on retrieved passages and to expose when evidence is missing. Higher temperature and larger output limits increase unsupported text, while fine-tuning stores a stale snapshot without providing query-time provenance. Together, the two prompt and output controls reduce fabrication.

Exam trap

The trap here is believing that more randomness or a larger output budget improves reasoning, when both actually increase the chance of unsupported claims.

164
MCQmedium

A research team uses a generative AI model to analyze historical texts. They want to provide users with insight into the model's reasoning process. Which explainability technique should they implement?

A.Chain-of-thought reasoning
B.Grounding
C.Confidence indicators
D.Safety filters
AnswerA

Chain-of-thought prompting makes the model emit intermediate reasoning steps before its conclusion, exposing the inferential path users can inspect. This satisfies the stem's explainability requirement by revealing how historical texts were interpreted, unlike post-hoc methods that only approximate feature importance.

Why this answer

Chain-of-thought reasoning provides step-by-step explanations of how the model arrived at its conclusion, enhancing transparency.

165
Multi-Selectmedium

A media company wants to use Google Cloud generative AI to produce short video clips from text prompts for internal storyboarding. The team needs a managed model that can generate video from text and wants to evaluate it before committing to production. Which two Google Cloud offerings or capabilities should they use? (Choose two.)

Select 2 answers
A.Cloud Translation API
B.Document AI
C.Veo on Vertex AI
D.Vertex AI Model Garden
E.Imagen on Vertex AI
AnswersC, D

Veo is Google Cloud's generative video model available through Vertex AI, and it can create short video clips from text prompts. Using it gives the media company the text-to-video capability it needs without building or hosting a video generation model itself. It fits the storyboarding use case directly.

Why this answer

Veo on Vertex AI provides the actual text-to-video generation, while Vertex AI Model Garden gives the team a managed place to discover and evaluate that model before production adoption. Together they cover both the creative capability and the evaluation step. The remaining services address images, documents, or language translation and cannot generate video.

Exam trap

The trap here is confusing image generation with video generation, or assuming any Vertex AI model catalog entry can produce motion output.

166
MCQmedium

A media company wants to add generative AI features to its mobile app but must control costs and prevent unexpected spend as usage grows. Leadership wants visibility into consumption by product team. Which governance approach should the GenAI Leader recommend on Google Cloud?

A.Rely on the monthly consolidated billing report and review Vertex AI charges with finance after each invoice arrives.
B.Apply a single organization-wide quota for Vertex AI and share one service account across all product teams.
C.Turn off Vertex AI API access for the organization and require teams to request manual approval for each generative feature.
D.Issue each product team its own Google Cloud project and apply quotas and budget alerts on the Vertex AI usage in those projects.
AnswerD

Separating product teams into distinct projects creates a natural billing and quota boundary, so consumption is attributable per team and can be capped. Quotas limit request or token throughput to prevent runaway spend, while budget alerts notify owners before limits are breached. This gives leadership both the per-team visibility and the preventive control the scenario requires.

Why this answer

Project-level separation combined with quotas and budget alerts gives each product team an attributable cost boundary and a hard ceiling on consumption. The other approaches either report spend only after it happens, blur attribution by sharing identity, or block the capability entirely instead of governing it.

Exam trap

The trap here is confusing cost reporting with cost control, when only quotas and alerts act before spend occurs.

167
MCQmedium

A developer is building a code generation assistant using Codey. They notice that the generated code sometimes contains deprecated API calls. What is the most likely cause?

A.The top-p sampling is too low, limiting the model's vocabulary
B.The temperature setting is too high, causing creative but incorrect outputs
C.The context window is too short to include relevant API documentation
D.Codey's training data has a knowledge cutoff date before the deprecation
AnswerD

Codey's training corpus has a fixed knowledge cutoff, so APIs deprecated after that date remain represented as current in its learned patterns. The model reproduces those stale calls because it has no post-cutoff knowledge, satisfying the deprecated-API symptom described.

Why this answer

Codey, like all large language models, is trained on a static dataset with a specific knowledge cutoff date. If the training data predates the deprecation of certain APIs, the model will not be aware of the newer, recommended alternatives and will continue to generate code using the deprecated calls. This is a fundamental limitation of the model's training data recency, not a parameter tuning issue.

Exam trap

A common trap in Google Gen AI exams is that candidates focus on hyperparameter tuning (temperature, top-p) as the cause for deprecated API calls, but the real issue is the training data's knowledge cutoff. For Codey, this is especially important when using older model versions.

How to eliminate wrong answers

Option A is wrong because top-p sampling controls the cumulative probability threshold for token selection, affecting output diversity, not the model's knowledge of deprecated APIs. Option B is wrong because temperature controls the randomness of token selection; while high temperature can increase creativity, it does not cause the model to generate deprecated APIs it is unaware of from its training data. Option C is wrong because the context window length determines how much input text the model can consider at once, but it does not affect the model's inherent knowledge of API deprecation dates; even with a long context, the model will still generate deprecated calls if its training data lacks that information.

168
MCQmedium

An e-commerce company is using Vertex AI PaLM 2 for Text (via Model Garden) to generate product descriptions. They have an existing pipeline that calls the model with a prompt including product attributes. Recently, they migrated to the Gemini API. The team notices that the Gemini model sometimes outputs descriptions that are factually inconsistent with the input (e.g., wrong color or size). This was less frequent with PaLM 2. They have not changed the prompts. What is the most likely cause and solution?

A.Revert to PaLM 2 since it was more reliable for this task.
B.Add negative prompts to discourage incorrect facts.
C.Adjust the prompt to be more explicit about adhering to the input data, and reduce the temperature.
D.Increase the model's temperature to make outputs more deterministic.
AnswerC

Lowering temperature sharpens the token probability distribution, reducing creative drift away from supplied attributes, while explicit instructions reinforce grounding on the input. Together these address the factual inconsistency introduced by Gemini's changed decoding behaviour and instruction-following, without altering the existing pipeline's product-attribute prompt structure.

Why this answer

The core issue is that the prompt, originally optimized for PaLM 2, may not be sufficiently explicit for the Gemini model's different instruction-following behavior. By making the prompt more explicit about adhering strictly to the input data and reducing the temperature (e.g., to 0.2 or lower), the model's output becomes more deterministic and less prone to hallucinating incorrect attributes. This directly addresses the factual inconsistency without changing the model family, leveraging Gemini's ability to follow detailed instructions when properly guided.

Exam trap

The trap here is that candidates assume model migration is the root cause and choose to revert (Option A), when in fact the real issue is prompt adaptation and hyperparameter tuning for the new model's behavior.

How to eliminate wrong answers

Option A is wrong because reverting to PaLM 2 ignores the fact that the prompt was not optimized for Gemini; the issue is prompt engineering and hyperparameter tuning, not model reliability. Option B is wrong because negative prompts are not a standard mechanism in Gemini or PaLM 2 for text generation; they are used in image generation models (e.g., Imagen) to avoid certain concepts, not to enforce factual consistency in text outputs. Option D is wrong because increasing temperature would make outputs more random and less deterministic, worsening the factual inconsistency problem, not solving it.

169
MCQeasy

A developer wants to generate Python code using Google Cloud's generative AI. Which model should they invoke?

A.Chirp
B.Codey
C.Imagen
D.Meena
AnswerB

Codey is Google Cloud's foundation model family purpose-built for code, trained on source code and tuned for generation, completion and debugging tasks. Invoking it satisfies the stem's Python code generation requirement, whereas general-purpose text models like PaLM or Gemini lack that code-specialised tuning.

Why this answer

Codey is Google Cloud's family of models specifically designed for code generation, completion, and chat, built on the PaLM 2 architecture and fine-tuned on code-heavy datasets. For a developer needing to generate Python code, Codey is the correct choice because it is purpose-built for code-related tasks, unlike other models that specialize in different modalities.

Exam trap

The trap here is that candidates may confuse Chirp (audio) or Imagen (image) with code generation because all are Google Cloud generative AI offerings, but each is specialized for a distinct modality, and the question explicitly asks for code generation.

How to eliminate wrong answers

Option A is wrong because Chirp is Google Cloud's speech-to-text model, designed for audio transcription, not code generation. Option C is wrong because Imagen is a text-to-image generation model, focused on creating visual content from text prompts, not code. Option D is wrong because Meena is a general-purpose conversational AI model (predecessor to LaMDA) optimized for open-domain dialogue, not for generating syntactically correct Python code.

170
Multi-Selectmedium

An organization is planning to roll out a generative AI internal knowledge base assistant to employees. They want to ensure adoption and manage change effectively. Which two change management practices should they prioritize? (Choose TWO)

Select 2 answers
A.Provide training sessions on how to write effective prompts
B.Roll out to all employees on day one to maximize impact
C.Mandate usage of the assistant for all employees
D.Start with a pilot group of AI champions to gather feedback
E.Disable the assistant after two weeks if usage is low
AnswersA, D

Training empowers users to get better results, increasing satisfaction and adoption.

Why this answer

Effective prompt engineering is critical for generative AI assistants; without training, employees may produce vague or poorly structured prompts, leading to irrelevant or low-quality responses, which undermines adoption. Providing training on prompt writing directly addresses the skill gap and empowers users to leverage the assistant effectively.

Exam trap

The Generative AI Leader exam often tests the distinction between 'maximizing immediate impact' (Option B) and 'phased adoption with feedback loops' (Option D), where candidates mistakenly choose a rapid full rollout thinking it drives faster adoption, ignoring the proven change management principle of starting small to build advocacy and refine the tool.

171
MCQmedium

A developer uses the Gemini API to summarize long articles. The summaries often miss key points from the end of the article. Which technique specifically addresses this length-based loss of information?

A.Increase the max output tokens to 2048
B.Break the article into sections and ask the model to summarize each section, then combine
C.Truncate the article to the first 2000 tokens
D.Use a different model with a larger context window
AnswerB

This structured approach ensures each part is summarized, mitigating attention drop-off.

Why this answer

The Gemini API, like many LLMs, has a limited context window and exhibits a 'lost-in-the-middle' effect where information at the beginning and end of long inputs is retained better, but the middle and far end can be dropped. By breaking the article into sections, summarizing each independently, and then combining those summaries, you ensure that key points from the end are captured in their own focused summary, bypassing the length-based information loss. This technique is a form of 'chunking' and 'recursive summarization' that directly addresses the model's tendency to lose context over long sequences.

Exam trap

Google often tests the misconception that simply increasing token limits or using a larger context window solves all length-related issues, when in fact the underlying attention mechanism and positional biases require explicit chunking strategies to reliably capture information from all parts of a long input.

How to eliminate wrong answers

Option A is wrong because increasing max output tokens only controls the length of the generated summary, not the input context; the model still processes the full article and may drop end content due to its fixed context window. Option C is wrong because truncating the article to the first 2000 tokens removes the end of the article entirely, which is the opposite of what is needed to capture key points from the end. Option D is wrong because using a different model with a larger context window does not guarantee the end content will be retained; even with larger windows, models can still suffer from the 'lost-in-the-middle' phenomenon where information in the middle and end is less reliably recalled.

172
MCQmedium

A media company wants to use a generative AI model to create marketing copy that includes citations to original sources. Which feature should they enable to ensure the model provides accurate attributions?

A.Confidence indicators
B.Grounding
C.Chain-of-thought reasoning
D.Safety filters
AnswerB

Grounding connects the model to authoritative external data sources, so generated marketing copy cites verifiable original material rather than relying on parametric memory alone. This directly satisfies the stem's requirement for accurate source attributions, since citations must reference retrieved documents the model can actually point to.

Why this answer

Grounding allows the model to cite sources, improving explainability and trustworthiness by connecting outputs to verifiable information.

173
MCQmedium

A company plans to commercially use images generated by a text-to-image model. What should they check to avoid copyright issues?

A.The training data provenance and the model's license terms
B.The model's accuracy on standard benchmarks
C.The model's latency and throughput
D.The model's bias evaluation results
AnswerA

Commercial reuse depends on the model's licence and the provenance of its training images, since these determine whether outputs can be lawfully exploited. Checking both satisfies the stem's copyright constraint before any generated image is published or sold.

Why this answer

Copyright issues arise from the training data and the model's license. The training data provenance determines if the model was trained on copyrighted works without permission, and the license terms specify whether commercial use is allowed. Without checking both, the company risks infringing on original creators' rights.

Exam trap

Google often tests the misconception that technical performance metrics (accuracy, latency, bias) are relevant to legal compliance, when in fact copyright issues hinge on data provenance and licensing.

How to eliminate wrong answers

Option B is wrong because model accuracy on standard benchmarks (e.g., FID score, CLIP score) measures image quality and alignment, not copyright compliance. Option C is wrong because latency and throughput are performance metrics for deployment, unrelated to legal or copyright considerations. Option D is wrong because bias evaluation results address fairness and ethical concerns, not copyright ownership or licensing.

174
MCQhard

A model response deployed on Vertex AI includes safety attributes with a toxicity score of 0.9 and an insult score of 0.3. The application must reject any prediction where the toxicity score exceeds 0.8. Based on the response, what action should the application take?

A.Reject the prediction because the toxicity score exceeds 0.8.
B.Retry the request with a lower temperature.
C.Display the prediction because the insult score is below 0.8.
D.Log the prediction but still display it.
AnswerA

The safety attributes returned by Vertex AI include a toxicity score, and the application's stated threshold is 0.8. Since the exhibited score exceeds that limit, the prediction breaches the configured policy and must be rejected before reaching the user.

Why this answer

The response from the model includes a safety attribute with a toxicity score of 0.9, which exceeds the application's threshold of 0.8. Vertex AI safety attributes provide scores for categories like toxicity, insult, and sexual content, and the application must enforce its own rejection logic based on these scores. Since the toxicity score is above the defined threshold, the application should reject the prediction to comply with safety policies.

Exam trap

Google's exam often tests the distinction between different safety attribute categories (e.g., toxicity vs. insult) and the importance of applying the correct threshold to the correct score. Candidates may mistakenly focus on a lower-scoring attribute instead of the one specified in the policy.

How to eliminate wrong answers

Option B is wrong because retrying with a lower temperature would not change the underlying safety attributes; temperature controls randomness in token generation, not the safety scores, and the model's output is already generated. Option C is wrong because the application's rejection criterion is based on the toxicity score, not the insult score; even if the insult score is below 0.8, the toxicity score of 0.9 triggers rejection. Option D is wrong because the application must reject the prediction when the toxicity score exceeds 0.8, not log and display it; logging without rejection violates the stated policy.

175
MCQmedium

A social media company uses a generative AI to moderate user comments. They need to filter hate speech, violence, and sexual content. What is the most efficient way to implement content safety in Vertex AI?

A.Hire human moderators to manually review all comments
B.Use a third-party API for content moderation
C.Train a custom content classifier from scratch using Vertex AI AutoML
D.Use Google's pre-built safety filters provided with Vertex AI
AnswerD

Vertex AI's pre-built safety filters apply configurable thresholds across harm categories including hate speech, violence and sexual content, requiring no custom model training. This satisfies the efficiency requirement: the filters operate on requests and responses directly, giving immediate coverage rather than building and maintaining bespoke classifiers.

Why this answer

Google's pre-built safety filters in Vertex AI are specifically designed for content moderation tasks like hate speech, violence, and sexual content detection. They are immediately available, require no custom training, and integrate directly with Vertex AI's generative AI workflows, making them the most efficient choice for a social media company needing rapid deployment.

Exam trap

The Generative AI Leader exam often tests the misconception that custom training (AutoML) is always better for domain-specific tasks, but here the pre-built filters are already optimized for the exact content categories needed, making custom training unnecessary and inefficient.

How to eliminate wrong answers

Option A is wrong because hiring human moderators is not efficient at scale; it introduces latency, high cost, and inconsistency, and does not leverage AI automation. Option B is wrong because using a third-party API introduces additional latency, cost, and potential data privacy concerns, and it does not integrate natively with Vertex AI's generative AI pipeline. Option C is wrong because training a custom content classifier from scratch using Vertex AI AutoML requires significant labeled data, time, and compute resources, which is inefficient compared to using pre-built, optimized safety filters.

176
MCQhard

A company with limited AI expertise wants to adopt gen AI. They need a solution that integrates with existing data and applications. Which Google Cloud offering is best?

A.Apigee
B.Colab Enterprise
C.BigQuery ML
D.Vertex AI Agent Builder
AnswerD

Vertex AI Agent Builder provides pre-built agents and connectors that ground generative output in the company's existing data sources and applications, so limited in-house AI expertise is not a barrier. It directly satisfies the integration constraint, unlike raw model APIs requiring custom orchestration.

Why this answer

Vertex AI Agent Builder is a Google Cloud offering designed to help organizations build and deploy generative AI agents and applications that integrate with enterprise data and existing applications, with minimal AI expertise required. It provides pre-built connectors, grounding with enterprise data, and managed infrastructure, making it the best fit for a company with limited AI expertise that needs integration.

Exam trap

The trap is confusing BigQuery ML (which is for traditional ML in SQL) with Vertex AI Agent Builder (which is for gen AI agents), or assuming Colab Enterprise is a turnkey integration solution when it is a developer notebook environment.

How to eliminate wrong answers

Option A is wrong because Apigee is an API management platform, not a generative AI application builder, and does not provide gen AI agent capabilities. Option B is wrong because Colab Enterprise is a managed notebook environment for data scientists and ML engineers, requiring significant AI expertise and not focused on integrating gen AI into existing applications. Option C is wrong because BigQuery ML allows SQL-based model training and prediction within BigQuery, but it is not a gen AI agent builder and does not provide the application integration and agent orchestration needed.

177
Multi-Selecthard

A company is deploying a generative AI application using Vertex AI. They need to minimize latency for real‑time inference while maintaining high quality. Which TWO actions are most effective?

Select 2 answers
A.Batch multiple inference requests together
B.Use a smaller model like Gemini Flash
C.Use Gemini Pro instead of Gemini Flash to ensure quality
D.Reduce the max output tokens to the minimum acceptable length
E.Increase the temperature to 1.0 for more creative outputs
AnswersB, D

Gemini Flash trades parameter count for throughput, so each decoding step costs fewer FLOPs and returns tokens faster than Pro-class models. That directly cuts per-token inference latency, meeting the real-time constraint while quality remains adequate for many production tasks.

Why this answer

Option B is correct because Gemini Flash is a lower-latency, lighter-weight model in the Gemini family, purpose-built for high-throughput, real-time tasks, so it directly reduces inference latency while still delivering strong quality for most generative AI use cases. Option D is correct because output tokens are generated sequentially (autoregressive decoding), so capping max output tokens to the shortest acceptable length proportionally cuts generation time and thus end-to-end latency. Option A is not appropriate because batching requests improves throughput and cost efficiency, not per-request real-time latency, and can actually add queuing delay.

Option C is wrong because Gemini Pro is larger and slower than Flash, trading latency for quality rather than minimizing latency. Option E is wrong because raising temperature to 1.0 only affects randomness/creativity of sampling and does not reduce latency.

Exam trap

The trap is assuming that larger models always provide better quality and that batching reduces latency; in reality, smaller models and shorter outputs reduce latency, while batching increases it.

178
Multi-Selectmedium

A data science team wants to decide between using a pre-built API (e.g., Vertex AI Gemini API) and fine-tuning a custom model for a specific business task. Which TWO factors are most important in making this build versus buy decision?

Select 2 answers
A.Color scheme of the user interface
B.Developer preference for programming languages
C.Number of users who will interact with the system
D.Cost of inference per query for pre-built API vs fine-tuned model
E.Availability of high-quality labeled data for the specific task
AnswersD, E

Per-query inference cost differs structurally: a pre-built API charges per call, whereas a fine-tuned model adds hosting or endpoint charges regardless of traffic. Comparing these at expected volume reveals which option is economically viable for the task.

Why this answer

Option D is correct because the build-versus-buy decision hinges on the total cost of ownership, and inference cost per query directly determines whether paying per-token/per-request for a pre-built API like Vertex AI Gemini is cheaper than hosting and serving a fine-tuned model at scale. Option E is correct because fine-tuning a custom model requires a sufficient volume of high-quality, task-specific labeled data; if such data is unavailable, the pre-built API is the practical choice, making data availability a decisive factor. The unmarked options do not belong: A (UI color scheme) is a cosmetic design choice unrelated to model economics or feasibility, B (developer language preference) is a minor implementation detail that does not determine build-versus-buy value, and C (number of users) is a demand metric that only indirectly matters through its effect on inference cost and data volume, which are already captured by D and E.

Exam trap

The trap here is focusing on superficial factors like UI design or developer preferences, which are irrelevant to the core economic and technical trade-offs in build vs. buy decisions for AI models.

179
MCQmedium

A global nonprofit organization is deploying a generative AI chatbot to provide educational content in multiple languages to underserved communities. They operate in regions with limited internet connectivity. The chatbot must work offline or with minimal data usage. The team has a moderate budget and limited technical staff. Which deployment strategy should they use?

A.Fine-tune an open-source model and host it on a cloud VM with auto-scaling
B.Deploy a distilled version of the model on edge devices using TensorFlow Lite
C.Host a large foundation model on Google Cloud and use a mobile app to send API requests
D.Deploy a distill of a smaller model on Google Cloud VM instances
AnswerB

TensorFlow Lite converts the model to a compact flatbuffer executed on-device, so inference runs locally without network calls — satisfying the offline and minimal-data constraint. Distillation shrinks the model to fit edge hardware, and the free runtime suits their moderate budget and limited technical staff.

Why this answer

Deploying a distilled version of the model on edge devices using TensorFlow Lite directly addresses the constraints of offline operation, minimal data usage, and limited technical staff. Distillation reduces model size and computational requirements, enabling inference on local hardware without cloud dependency, which is critical for underserved regions with intermittent connectivity.

Exam trap

The trap here is that candidates confuse 'distillation on edge' with 'distillation on cloud VMs' (Option D), overlooking that edge deployment is the only way to guarantee offline functionality, while cloud VMs still require network access for inference.

How to eliminate wrong answers

Option A is wrong because hosting a fine-tuned model on a cloud VM with auto-scaling requires constant internet connectivity for the chatbot to function, which fails the offline requirement. Option C is wrong because using a large foundation model via API requests from a mobile app incurs high data usage and relies on continuous cloud access, contradicting the need for minimal data usage and offline capability. Option D is wrong because deploying a distilled model on Google Cloud VM instances still requires internet connectivity for inference, missing the offline requirement, and does not leverage edge deployment for local processing.

180
Multi-Selectmedium

A data scientist wants to apply reinforcement learning from human feedback (RLHF) to improve a chatbot's helpfulness. Which TWO steps are part of the RLHF process? (Select 2)

Select 2 answers
A.Use prompt engineering to tune the model without retraining
B.Collect human-annotated demonstrations of ideal responses
C.Collect human rankings or preferences on multiple model outputs
D.Deploy the model in A/B testing to gather implicit feedback
E.Train a reward model based on human preferences
AnswersC, E

RLHF begins with a supervised fine-tuning stage, then humans rank or compare multiple model outputs to express preferences. These preference rankings become the training signal for the reward model, making this a core data-collection step.

Why this answer

Option C is correct because RLHF fundamentally relies on gathering human preference data, typically by having annotators rank or compare multiple candidate responses from the model, which produces the preference signal used to define what 'helpful' means. Option E is correct because those human preferences are used to train a separate reward model that learns to predict human judgments, and this reward model then supplies the scalar reward signal that the policy (the chatbot) is optimized against with reinforcement learning. Options A, B, and D are not part of the core RLHF pipeline: prompt engineering (A) tunes behavior without retraining and involves no reward learning, collecting ideal demonstrations (B) is supervised fine-tuning rather than preference-based RLHF, and A/B testing in deployment (D) is an evaluation/production feedback technique, not a training step in the RLHF process.

181
MCQmedium

A company is piloting a GenAI feature for internal knowledge base search. During the pilot, users report that the AI sometimes gives incorrect answers based on outdated documents. What is the MOST effective way to address this issue?

A.Add a system instruction to the prompt telling the model to only answer if it is confident
B.Decrease the temperature parameter of the model to 0 to reduce randomness
C.Implement Retrieval-Augmented Generation (RAG) with the knowledge base documents indexed in a vector store and ensure the index is updated when documents change
D.Fine-tune the model on the current knowledge base to improve accuracy
AnswerC

RAG grounds responses in the indexed knowledge base at query time, so stale documents are replaced once the vector index refreshes. This directly addresses the outdated-document constraint, unlike fine-tuning, which bakes stale content into model weights and cannot be updated cheaply.

Why this answer

Option C is correct because RAG retrieves the most current, relevant documents from a vector store at query time and feeds them to the model, so answers are grounded in up-to-date knowledge base content. Updating the index when documents change ensures the model never relies on stale information. This directly addresses the root cause: outdated source documents.

Exam trap

Generative AI Leader often tests the misconception that fine-tuning or prompt engineering can solve stale-data problems — the correct answer is almost always RAG with a maintained index when the issue is outdated source content.

How to eliminate wrong answers

Option A is wrong because instructing the model to 'only answer if confident' does not fix outdated data — the model may be confidently wrong based on stale documents. Option B is wrong because lowering temperature reduces randomness but does not update the knowledge source; the model can still produce incorrect answers from outdated context. Option D is wrong because fine-tuning on the current knowledge base is expensive, slow to update, and still becomes stale as documents change; RAG is the standard pattern for dynamic knowledge retrieval.

182
Multi-Selecteasy

A startup is building a multimodal application that needs to process both images and text. They want to prototype quickly for free before moving to production with enterprise controls. Which TWO services should they use? (Choose 2)

Select 2 answers
A.Google AI Studio for prototyping
B.Vertex AI for prototyping
C.Google Colab for production
D.Vision AI and Natural Language AI separately
E.Vertex AI for production deployment
AnswersA, E

AI Studio provides free access to Gemini models for rapid prototyping of multimodal applications.

Why this answer

Google AI Studio is correct because it provides a free, browser-based environment for prototyping multimodal applications that process both images and text, using models like Gemini. It allows rapid experimentation without cost or infrastructure setup, making it ideal for quick prototyping before moving to production.

Exam trap

Google often tests the distinction between prototyping and production services, trapping candidates who confuse Vertex AI (production) with AI Studio (prototyping) or who think Colab is suitable for production deployment.

183
MCQmedium

A healthcare startup is using a large language model (LLM) to generate discharge summaries. To comply with regulations, they need to ensure that a human reviews all AI-generated summaries before they are sent to patients. Which Google Cloud feature should they use to enforce this workflow?

A.Cloud Audit Logs
B.Vertex AI Human-in-the-Loop (HITL)
C.Vertex AI Model Registry
D.Vertex AI Evaluation Service
AnswerB

Vertex AI Human-in-the-Loop enforces a mandatory review step before AI-generated discharge summaries reach patients, directly satisfying the regulatory constraint. It pauses the pipeline for clinician approval, so no summary is dispatched without human sign-off, unlike post-hoc logging or evaluation tools that cannot block delivery.

Why this answer

Human oversight is a key requirement for high-stakes AI systems. Vertex AI Human-in-the-Loop (HITL) provides a managed workflow to route predictions for human review, approval, or override before final output.

184
MCQeasy

A marketing team wants to generate blog post ideas and draft outlines. They need a solution that works within Google Docs and leverages a foundation model without leaving the document editor. Which service should they use?

A.Vertex AI Studio
B.Gemini for Google Workspace (Duet AI)
C.Model Garden
D.Vertex AI Agent Builder
AnswerB

Gemini for Google Workspace embeds generative AI directly inside Docs, satisfying the constraint of working without leaving the editor. It uses Google's foundation models to draft outlines and brainstorm ideas in-context, unlike standalone chatbot interfaces or API-based services that require switching applications or custom development.

Why this answer

Gemini for Google Workspace (formerly Duet AI) is the correct choice because it is the only service that integrates generative AI directly into Google Docs, allowing users to generate blog post ideas and outlines without leaving the document editor. It leverages a foundation model (Gemini) natively within the Workspace environment, enabling seamless, in-context assistance for content creation tasks.

Exam trap

The trap here is that candidates often confuse Vertex AI Studio (a model development platform) with Gemini for Google Workspace (a productivity assistant), assuming any Google Cloud generative AI service can be used directly within Docs, but only the Workspace-integrated tool provides the seamless, in-editor experience described.

How to eliminate wrong answers

Option A is wrong because Vertex AI Studio is a standalone platform for building, testing, and customizing generative AI models, not an integrated assistant within Google Docs; it requires leaving the document editor to interact with the model. Option C is wrong because Model Garden is a repository of pre-trained foundation models and does not provide a direct, in-editor assistant for generating content within Google Docs. Option D is wrong because Vertex AI Agent Builder is designed for creating conversational agents and search applications, not for inline content generation within a document editor like Google Docs.

185
MCQeasy

A developer wants to use a pre-trained model to identify objects in images. Which Google Cloud AI API should they use?

A.Speech-to-Text
B.Translation API
C.Natural Language AI
D.Vision AI
AnswerD

Vision AI provides pre-trained models purpose-built for image analysis, including object detection and label recognition, so no custom training is needed. It directly satisfies the developer's requirement to identify objects in images using an existing model, unlike APIs scoped to text, speech or video.

Why this answer

Vision AI provides pre-trained models for object detection, image classification, etc. Natural Language AI is for text, Speech-to-Text for audio, and Translation for text translation.

186
MCQmedium

A software vendor is building a product that must call a Gemini model through a stable, versioned API with enterprise controls such as VPC Service Controls, and must run on Google Cloud infrastructure. Which Google Cloud offering should the vendor use?

A.Gemini for Google Workspace
B.Google AI Studio
C.The Gemini API in Vertex AI
D.Vertex AI Feature Store
AnswerC

The Gemini API in Vertex AI exposes Google's foundation models through a versioned enterprise endpoint that supports Google Cloud controls including VPC Service Controls, IAM, and audit logging. It runs on Google Cloud infrastructure and is designed for building production applications, so it satisfies the stability and governance requirements described.

Why this answer

The Gemini API in Vertex AI is the enterprise path to Google's foundation models: it offers versioned endpoints, runs on Google Cloud, and integrates with IAM, VPC Service Controls, and audit logging. Workspace is an end-user assistant, AI Studio targets prototyping without enterprise governance, and Feature Store serves ML features rather than generative model inference.

Exam trap

The trap here is confusing the convenient prototyping API with the enterprise API; only the Vertex AI endpoint provides the versioning and Google Cloud governance controls a shipped product requires.

187
MCQmedium

Refer to the exhibit. A developer executed the command to list endpoints. They notice that two models are deployed to the same endpoint. What is the most likely reason for this configuration?

A.It is a canary deployment with traffic splitting
B.The endpoint is misconfigured and will cause conflicts
C.The models are from different frameworks
D.It is a batch prediction endpoint
AnswerA

A canary deployment routes a small slice of live traffic to a new model version while the rest continues to the stable one, so both models must sit behind the same endpoint for traffic splitting to work. This matches the observed configuration without implying an A/B test or blue-green swap.

Why this answer

A is correct because deploying two models to the same endpoint with traffic splitting is a standard canary deployment strategy. In this configuration, a small percentage of inference requests are routed to the new model while the majority go to the stable model, allowing validation of the new model's performance before full rollout. This is commonly supported by Google Cloud's Vertex AI, where you can deploy multiple models to an endpoint and assign traffic percentages to each model variant (e.g., 90% to the stable model and 10% to the canary model).

Exam trap

Google Cloud often tests the misconception that deploying two models to the same endpoint is always an error, when in fact it is a deliberate pattern for canary testing or A/B testing with traffic splitting.

How to eliminate wrong answers

Option B is wrong because deploying two models to the same endpoint with traffic splitting is a deliberate, supported configuration, not a misconfiguration; conflicts are avoided by routing traffic based on defined weights. Option C is wrong because models from different frameworks can be deployed to the same endpoint without issue, as the serving layer handles framework-specific inference containers independently. Option D is wrong because batch prediction endpoints typically use a single model or a single job configuration, not multiple models deployed simultaneously with traffic splitting.

188
Multi-Selecthard

An organization is building a generative AI application on Vertex AI. Which THREE actions should they take to ensure responsible AI practices?

Select 3 answers
A.Disable content filtering
B.Implement human review for sensitive outputs
C.Conduct fairness evaluation
D.Create a safety policy and enforce via content filtering
E.Use only Google's foundation models
AnswersB, C, D

Human review places a person in the loop to inspect sensitive outputs before they reach users, catching harmful or inaccurate content that automated filters miss. This directly satisfies the responsible AI requirement for oversight and accountability in the generative AI application built on Vertex AI.

Why this answer

Option B is correct because implementing human review for sensitive outputs ensures that high-risk or ambiguous AI-generated content is validated by a person before it reaches users, which is a core responsible AI safeguard for generative applications on Vertex AI. Option C is correct because conducting fairness evaluation (for example, using Vertex AI's model evaluation tools to assess bias across demographic groups) helps detect and mitigate discriminatory or skewed model behavior, directly supporting responsible AI. Option D is correct because creating a safety policy and enforcing it via content filtering operationalizes responsible AI by defining prohibited content and using Vertex AI safety filters (such as configurable hate speech, harassment, and dangerous content thresholds) to block violations.

Option A is not correct because disabling content filtering removes a key safety control and increases the risk of harmful outputs, which is contrary to responsible AI. Option E is not correct because using only Google's foundation models does not by itself guarantee responsible AI; responsibility depends on evaluation, policy enforcement, monitoring, and human oversight regardless of which models are used.

Exam trap

The trap here is that candidates might think disabling content filtering (Option A) improves performance, but the exam tests that responsible AI on Google Cloud's Vertex AI requires both automated filters and human oversight, as emphasized in Google's AI Principles.

189
MCQmedium

A retail company uses Vertex AI Agent Builder to create a virtual assistant for order tracking. Users frequently ask about delivery dates, but the assistant sometimes gives incorrect information. The team wants to improve accuracy without retraining the underlying model. Which technique should they apply?

A.Increase the temperature parameter for more deterministic outputs
B.Switch to a larger model size for better reasoning
C.Add more few-shot examples to the prompt template
D.Enable Grounding with Google Search or connect to a custom data store
AnswerD

Grounding anchors the agent's responses to retrieved, verifiable sources — Google Search or a custom data store holding live order data — rather than the model's frozen parameters. This corrects stale delivery-date answers without retraining, satisfying the accuracy constraint.

Why this answer

Grounding with Google Search or enterprise data sources (like order databases) ensures the agent retrieves real-time, accurate information instead of relying solely on the model's training data.

190
MCQeasy

A startup wants to embed generative AI features into their mobile app but has limited ML expertise. Which Google Cloud service is best suited for rapid integration with no ML training?

A.Vertex AI Model Garden
B.Vertex AI Agent Builder
C.Gemini API
D.Cloud Run with a custom container
AnswerC

The Gemini API provides pre-trained generative capabilities callable directly from application code, requiring no model training, tuning, or ML expertise. This satisfies the stem's constraints of limited ML expertise and rapid integration into a mobile app.

Why this answer

The Gemini API provides direct, no-code access to Google's most capable generative AI models via a simple REST API, requiring zero ML training or infrastructure setup. This makes it the fastest path for a startup with limited ML expertise to embed generative AI features like text generation, summarization, or chat into a mobile app.

Exam trap

Candidates often confuse the Gemini API with Vertex AI Model Garden. They choose Model Garden because it offers a broader platform, but the requirement for 'no ML training' and 'rapid integration' indicates that the direct API is the best fit.

How to eliminate wrong answers

Option A is wrong because Vertex AI Model Garden is a curated repository of foundation models that still requires users to deploy, fine-tune, or serve models via endpoints, which demands ML expertise and operational overhead. Option B is wrong because Vertex AI Agent Builder is designed for building conversational AI agents with grounding and tool integration, which involves more complex orchestration and is overkill for simply embedding generative AI features without ML training. Option D is wrong because Cloud Run with a custom container requires the startup to containerize and deploy their own model or inference code, which necessitates ML expertise and infrastructure management, contradicting the 'no ML training' requirement.

191
Multi-Selectmedium

A financial services firm is evaluating Google Cloud generative AI offerings for an internal knowledge assistant. They need capabilities for controlling access to model endpoints and for monitoring model usage and safety signals. (Choose two.)

Select 2 answers
A.Cloud Monitoring dashboards and alerting for Vertex AI metrics
B.BigQuery ML model export to Cloud Storage
C.Cloud Storage bucket lifecycle rules on training data
D.IAM roles and policies on Vertex AI resources
E.Vertex AI Feature Store for online serving
AnswersA, D

Cloud Monitoring collects Vertex AI metrics such as request counts, latency, and error rates, and can alert on anomalies. This supports the requirement to monitor model usage and safety-related signals, enabling the firm to detect unexpected traffic or failures and respond operationally.

Why this answer

IAM roles and policies restrict and grant access to Vertex AI endpoints and related resources, satisfying the access-control requirement. Cloud Monitoring dashboards and alerts surface usage, latency, and error signals for those endpoints, satisfying the monitoring requirement. The remaining options address model export, feature serving, and storage lifecycle, none of which cover access control or operational monitoring.

Exam trap

The trap here is selecting data or ML lifecycle tools that sound adjacent to governance but do not actually control endpoint access or report usage metrics.

192
Multi-Selecthard

A company is transitioning a generative AI pilot to production. They need to ensure cost predictability and scalability. Which THREE actions should they take?

Select 3 answers
A.Provision reserved throughput for all requests
B.Implement response caching for common queries
C.Select the smallest model that meets quality requirements
D.Use batch API for non-real-time requests
E.Conduct A/B testing on model versions
AnswersB, C, D

Caching avoids redundant API calls, reducing token usage and cost.

Why this answer

Response caching for common queries reduces latency and API costs by serving repeated requests from a cache instead of invoking the generative AI model each time. This directly improves cost predictability (fewer model invocations) and scalability (reduced load on the model endpoint), making it a core optimization for production deployments.

Exam trap

Google often tests the misconception that reserved throughput (Option A) is always the best way to ensure cost predictability, when in fact it can increase costs for spiky workloads and ignores the scalability benefits of caching and batch processing.

193
MCQhard

A financial services firm needs to use Gemini for analyzing customer transaction data. They require that all data remain within their VPC and that model inference logs be auditable. Which access tier should they choose?

A.Colab Enterprise
B.Gemini API without Vertex AI
C.Vertex AI
D.Google AI Studio
AnswerC

Vertex AI runs Gemini within the customer's Google Cloud project boundary, so transaction data stays inside their VPC and inference logging flows to Cloud Audit Logs. The consumer Gemini app cannot satisfy the data-residency and auditability constraints, making Vertex AI the required access tier.

Why this answer

Vertex AI provides enterprise controls like VPC-SC, data isolation, and audit logging, while Google AI Studio is a prototyping environment without these guarantees.

194
MCQmedium

A financial services firm uses a generative AI model on Vertex AI to answer employees' HR policy questions. The model sometimes invents policy details. The firm wants answers grounded in the official HR handbook and needs to cite the source section. Which solution should they implement?

A.Use retrieval-augmented generation with Vertex AI Search over the HR handbook and instruct the model to cite retrieved sections.
B.Fine-tune the model on the HR handbook text so the knowledge is embedded in the weights.
C.Increase the context window by using a model with a larger token limit and paste the entire handbook into every prompt.
D.Lower the temperature to zero and rely on the model's pretrained knowledge of HR policies.
AnswerA

RAG retrieves the most relevant handbook passages and injects them into the prompt, so answers are based on official text. Instructing the model to cite the retrieved section provides traceability. When the handbook is updated, reindexing the source keeps answers current without retraining. This directly addresses both grounding and citation requirements.

Why this answer

Retrieval-augmented generation over the HR handbook supplies authoritative passages and enables citations, keeping answers grounded and current. Fine-tuning embeds knowledge but does not guarantee citation or easy updates. Low temperature does not add missing knowledge, and stuffing the full handbook into the prompt is inefficient and less precise than retrieval.

Exam trap

The trap here is believing that fine-tuning or a larger context window provides reliable grounding and citations, when retrieval is designed specifically for that purpose.

195
Multi-Selectmedium

A global retailer wants to build a generative AI application on Google Cloud that answers employee questions about HR policies. The team must ensure the application uses their private policy documents, keeps answers grounded in those documents, and avoids exposing confidential content to unauthorized staff. Which two Google Cloud components should they include in the architecture? (Choose two.)

Select 2 answers
A.Vertex AI Pipelines for nightly retraining of the Gemini model on HR documents
B.Vertex AI Model Registry for versioning the HR policy document set
C.Vertex AI Search configured with the HR policy documents as a data store
D.Gemini models on Vertex AI for generating grounded responses
E.Vertex AI Feature Store for storing employee tenure and department features
AnswersC, D

Vertex AI Search can index the private HR policy documents and serve relevant passages to the model at query time, which keeps answers grounded in the company's own content. Its enterprise connectors and access controls also help restrict which documents are retrievable, supporting the confidentiality requirement for unauthorized staff.

Why this answer

A grounded HR assistant needs a retrieval layer and a generation layer. Vertex AI Search indexes the private policy documents, retrieves relevant passages per question, and supports access controls, while Gemini models on Vertex AI generate answers from those passages with grounding. Feature Store, Pipelines, and Model Registry address structured features, workflow orchestration, and model versioning, none of which provide document grounding or content access control.

Exam trap

The trap here is treating model training and versioning services as substitutes for retrieval, when grounding private documents requires an index that is queried at runtime rather than weights updated by retraining.

196
MCQmedium

A healthcare organization wants to build a generative AI application that summarizes patient notes. They are concerned about the model generating harmful or inappropriate content. Which Google Cloud service should they use to filter out such content in real time?

A.Vertex AI Pipelines
B.Vertex AI Search
C.Vertex AI Model Garden
D.Vertex AI Safety Filters
AnswerD

Vertex AI Safety Filters are designed to detect and block harmful or inappropriate content in generative AI outputs. They can be configured with thresholds for categories like hate speech, harassment, and sexually explicit content. This service directly addresses the healthcare organization's need to filter out such content in real time, ensuring patient note summaries remain safe and compliant.

Why this answer

Vertex AI Safety Filters are specifically built to detect and filter harmful content in generative AI outputs. For a healthcare application summarizing patient notes, using these filters ensures that inappropriate or harmful content is blocked in real time. Other services like Model Garden, Search, or Pipelines serve different purposes and do not provide content moderation.

Exam trap

The trap here is assuming that any Vertex AI service can filter content, when only Safety Filters provide that specific capability.

197
MCQeasy

A company uses a generative AI model to answer customer queries. The model sometimes returns outdated information. Which technique should they apply to ensure responses rely on current data?

A.Fine-tune the model on historical data.
B.Extend the context window to include more tokens.
C.Increase the model's temperature to encourage novelty.
D.Use grounding with a refreshed knowledge base.
AnswerD

Grounding retrieves relevant passages from a refreshed knowledge base at query time and injects them into the prompt, so answers reflect current data rather than stale parametric training. This directly satisfies the requirement that responses rely on up-to-date information.

Why this answer

Grounding with a refreshed knowledge base is the correct technique because it directly connects the generative AI model to an external, up-to-date data source at inference time. This ensures responses are based on current information without retraining the model, addressing the problem of outdated outputs by retrieving fresh data from a vector database or API in real time.

Exam trap

Google often tests the misconception that fine-tuning is the primary method to update model knowledge, when in fact grounding with a refreshed knowledge base is the correct approach for real-time data currency without retraining.

How to eliminate wrong answers

Option A is wrong because fine-tuning on historical data would embed outdated information further into the model's parameters, worsening the problem of stale responses. Option B is wrong because extending the context window only allows the model to process more tokens in a single prompt, but does not introduce new or current data; it merely expands the capacity for existing input. Option C is wrong because increasing the temperature encourages more random or creative outputs, which does not guarantee factual accuracy or recency; it can actually increase hallucinations.

198
MCQeasy

A marketing team wants to use a generative AI model to create blog post drafts from short product descriptions. They need the model to produce varied, creative text each time they run it. Which configuration parameter should they adjust to control the randomness of the output?

A.Temperature
B.Top-K
C.Top-P
D.Max output tokens
AnswerA

Temperature controls the randomness of the model's output. A higher temperature (e.g., 0.9) makes the output more diverse and creative, while a lower temperature makes it more deterministic. For generating varied blog post drafts, increasing the temperature is appropriate to encourage creative variation.

Why this answer

Temperature is the parameter that scales the logits before softmax, directly controlling the randomness of the output distribution. Higher values flatten the distribution, making less likely tokens more probable, which increases creativity and variation. For generating diverse blog drafts, the team should increase the temperature setting.

Exam trap

The trap here is confusing temperature with top-K or top-P, which also affect diversity but do not directly control the degree of randomness in the same way.

199
MCQeasy

A developer is using the Gemini API to generate creative product taglines. The taglines are often bland and uncreative. The developer wants more variety and novelty in the outputs. Which parameter adjustment would most effectively increase the diversity of the generated taglines?

A.Decrease top_p from 1.0 to 0.5.
B.Set frequency_penalty to 2.0.
C.Increase temperature from 0.2 to 0.9.
D.Decrease temperature from 0.7 to 0.2.
AnswerC

Temperature scales the sampling distribution's randomness; raising it from 0.2 to 0.9 flattens the probability curve, so lower-probability tokens are selected more often. That directly increases tagline variety and novelty rather than reinforcing the bland high-probability output.

Why this answer

Increasing temperature from 0.2 to 0.9 raises the randomness of token sampling, which directly increases the diversity and novelty of generated text. A low temperature (e.g., 0.2) makes the model highly deterministic, always picking the most probable next token, leading to bland outputs. A higher temperature (e.g., 0.9) allows less probable tokens to be selected more often, producing more creative and varied taglines.

Exam trap

The trap here is that candidates often confuse temperature with top_p, incorrectly assuming that lowering top_p increases diversity, when in fact it restricts the token pool and reduces variety.

How to eliminate wrong answers

Option A is wrong because decreasing top_p from 1.0 to 0.5 reduces the nucleus of tokens considered for sampling, which actually decreases diversity by cutting off the long tail of less probable tokens. Option B is wrong because setting frequency_penalty to 2.0 penalizes token repetition too aggressively, which can suppress natural language patterns and may reduce overall output quality without directly increasing novelty. Option D is wrong because decreasing temperature from 0.7 to 0.2 makes the model more deterministic, reducing randomness and thus decreasing diversity, which is the opposite of what the developer wants.

200
MCQhard

A healthcare company is using a generative AI model to draft patient education materials. The model sometimes includes outdated medical advice. The team wants to ensure the content reflects the latest clinical guidelines. They have a database of current guidelines and want to integrate it into the generation process without retraining the model. Which approach should they use?

A.Retrieval-augmented generation (RAG) with the guidelines database as the retrieval source.
B.Increasing the model's temperature to encourage more up-to-date responses.
C.Fine-tuning the model on the latest guidelines.
D.Using a chain-of-thought prompt to ask the model to reason about the latest guidelines.
AnswerA

RAG dynamically retrieves relevant passages from the guidelines database and provides them as context to the model. This ensures the generated content is grounded in the latest clinical guidelines without modifying the model's weights. It is ideal for frequently updated information and avoids the cost of retraining.

Why this answer

Retrieval-augmented generation retrieves relevant, current information from an external database and includes it in the prompt, grounding the model's output in the latest guidelines. This approach avoids retraining and ensures the content is based on verified, up-to-date sources. Other methods like fine-tuning, temperature adjustment, or chain-of-thought do not provide the necessary external knowledge.

Exam trap

The trap here is assuming that prompt engineering or parameter tuning can inject new factual knowledge into a model without external data.

201
MCQhard

A financial services firm is using Vertex AI to generate investment reports. They need to ensure that the model outputs are explainable and comply with regulatory requirements. Which Vertex AI feature should they use?

A.Vertex AI Model Registry
B.Vertex Explainable AI
C.Vertex AI Safety Settings
D.Vertex AI AutoML
AnswerB

Vertex Explainable AI provides feature attributions that show which input features influenced each prediction, satisfying the regulatory requirement for explainable outputs. It integrates directly with Vertex AI models, delivering explanations alongside predictions without altering the model architecture, which meets the financial firm's compliance constraint for investment reports.

Why this answer

Vertex Explainable AI provides feature attributions and explanations for model predictions, which is essential for financial services firms that must comply with regulatory requirements like the EU's GDPR or the US SEC's model risk management guidelines. It helps auditors and stakeholders understand why a model generated a specific investment report output, ensuring transparency and accountability in AI-driven decisions.

Exam trap

Google's certification exam often tests the distinction between safety/security features and explainability features, leading candidates to confuse Vertex AI Safety Settings (which block harmful content) with the need for regulatory compliance explanations.

How to eliminate wrong answers

Option A is wrong because Vertex AI Model Registry is a centralized repository for managing and deploying models, not a tool for generating explanations or ensuring regulatory compliance. Option C is wrong because Vertex AI Safety Settings are designed to filter harmful content and enforce safety policies, not to provide explainability for model outputs. Option D is wrong because Vertex AI AutoML automates model training and deployment but does not inherently provide explainability features; you would need to integrate Explainable AI separately for that purpose.

202
MCQhard

A company deploys a Gemini model on Vertex AI for a healthcare application. They need to ensure that the model does not generate medical advice and that responses are grounded in trusted medical sources. Which combination of safety measures should they implement?

A.Enable safety filters and use Vertex AI Grounding with a labeled medical dataset
B.Use Vertex AI Grounding with a public dataset and disable safety filters
C.Enable safety filters only, without grounding
D.Fine-tune the model on a curated medical dataset and disable safety filters for faster responses
AnswerA

Safety filters block the model from producing medical advice, while Vertex AI Grounding against a labelled medical dataset constrains responses to trusted sources, jointly satisfying both the no-advice and grounded-in-trusted-sources constraints in the healthcare scenario.

Why this answer

It combines two essential safety layers: safety filters block harmful content (including medical advice), and Vertex AI Grounding anchors responses to a labeled medical dataset, ensuring factual accuracy and compliance with healthcare regulations. This dual approach prevents the model from generating unverified or dangerous medical information while maintaining relevance to trusted sources.

Exam trap

The trap here is that candidates assume fine-tuning alone is sufficient for domain-specific safety, but without grounding and safety filters, the model can still hallucinate or generate unverified medical advice, which is a key distinction Google Cloud tests in the Generative AI Leader exam.

How to eliminate wrong answers

Option B is wrong because using a public dataset for grounding introduces unverified or non-authoritative medical information, and disabling safety filters removes the critical barrier against generating harmful or unlicensed medical advice. Option C is wrong because safety filters alone cannot ensure responses are grounded in trusted medical sources; they only block explicit content but do not prevent the model from fabricating medical facts. Option D is wrong because fine-tuning on a curated dataset does not guarantee real-time grounding in trusted sources, and disabling safety filters exposes the application to generating unverified medical advice, which is unacceptable in healthcare.

203
Multi-Selecthard

A global e-commerce company uses generative AI to generate product descriptions in multiple languages. They want to ensure consistency across markets while respecting cultural nuances. Which THREE strategies should they adopt?

Select 3 answers
A.Standardize all descriptions to a neutral tone to avoid cultural issues.
B.Develop region-specific prompt templates that incorporate local cultural references and legal requirements.
C.Engage local marketing teams to review and approve AI-generated descriptions before publication.
D.Use a single global model with a translation layer to convert English descriptions.
E.Use A/B testing to measure engagement metrics per region and iterate on prompts.
AnswersB, C, E

Region-specific prompt templates embed local cultural references and legal requirements into generation, ensuring each market's descriptions respect nuances while a shared template structure maintains cross-market consistency. This directly satisfies both the consistency and cultural-respect constraints in the stem.

Why this answer

Option B is correct because region-specific prompt templates let the generative AI encode local cultural references, idioms, and market-specific legal requirements (e.g., advertising claims, required disclaimers) directly into generation, which preserves consistency of brand intent while adapting to each locale. Option C is correct because human-in-the-loop review by local marketing teams catches culturally inappropriate phrasing, mistranslations, and compliance issues that an AI model may miss before content is published. Option E is correct because A/B testing per region provides quantitative engagement metrics (click-through rate, conversion rate, dwell time) that let the company iteratively refine prompts and validate that localized content actually resonates.

Option A is not appropriate because a single neutral tone strips out the cultural nuances the scenario explicitly wants to respect and can still be perceived as tone-deaf in some markets. Option D is not appropriate because a single global model with a translation layer typically produces literal translations that lose idiomatic and cultural context, and it does not address per-region legal requirements.

Exam trap

Generative AI Leader often tests the trade-off between global standardization and local adaptation, and candidates may choose neutral tone or simple translation as shortcuts, ignoring cultural nuances.

204
MCQmedium

A national retailer's generative AI pilot for product description writing succeeded technically, but six months later only two of forty merchandising teams use it. Interviews reveal that teams were never trained, the tool sits outside their existing content workflow, and no one owns adoption targets. Which action best addresses the root cause of this outcome?

A.Upgrade to a newer foundation model with higher benchmark scores for creative writing tasks.
B.Establish an adoption plan with named business owners, role-specific enablement, and integration of the tool into the existing content workflow.
C.Publish a company-wide mandate requiring all merchandising teams to use the tool for every product description.
D.Reduce the per-token cost of the generative AI service by switching to a smaller, cheaper model.
AnswerB

The interviews point to organizational causes: no training, poor workflow fit, and no ownership. An adoption plan that assigns business owners, delivers role-specific enablement, and embeds the tool where merchandisers already work addresses each cause directly, which is what turns a technically successful pilot into sustained business value across the remaining teams.

Why this answer

Adoption failures after a technically successful pilot are usually organizational, not technical. The interviews identify missing enablement, poor workflow fit, and absent ownership, so the effective response is an adoption plan that names business owners, trains users by role, and integrates the assistant into the tools merchandisers already use daily.

Exam trap

The trap here is assuming that low usage signals a model quality or cost problem, when the stated evidence points to enablement, workflow integration, and ownership gaps that no model change can fix.

205
MCQeasy

What is the primary function of embeddings in the context of generative AI?

A.To control the creativity of the model's output
B.To define the maximum length of the generated text
C.To convert tokens into a fixed-length vector that captures semantic meaning
D.To determine the next token in a sequence during generation
AnswerC

Embeddings map discrete tokens to dense fixed-length vectors positioned so semantically similar content sits nearby in vector space, enabling similarity search, clustering and retrieval. This vector representation, not tokenisation itself, satisfies the semantic-meaning requirement in the stem.

Why this answer

Embeddings are numerical representations of data that capture semantic meaning, enabling vector search and semantic similarity. They are not used for token generation directly, and they are not the same as temperature or context window.

206
MCQeasy

A project manager wants to reduce the cost of using Gemini API for batch processing of customer feedback. The team is on a tight budget. Which cost management strategy is MOST effective?

A.Use batch requests to group multiple prompts
B.Increase the temperature to max to reduce output length
C.Disable logging to reduce storage costs
D.Switch to the largest model available for better accuracy
AnswerA

Batch requests are cheaper than individual calls for large volumes.

Why this answer

Batch requests reduce per-token cost by grouping multiple prompts into one API call. Caching helps but is less impactful for varied feedback.

207
Multi-Selecthard

A machine learning engineer is tuning a large language model on Vertex AI for question answering. They want to evaluate the model's performance before deployment. Which THREE metrics should they consider?

Select 3 answers
A.Cost per training epoch
B.F1 score
C.Exact match (EM)
D.Training time per epoch
E.ROUGE-L score
AnswersB, C, E

F1 score is the harmonic mean of precision and recall, capturing the balance between correctly retrieved answers and missed ones. For question answering it exposes whether the tuned model trades precision for coverage, which accuracy alone would hide.

Why this answer

For a question-answering evaluation on Vertex AI, F1 score (B) is correct because it measures token-level overlap between the predicted answer and the ground-truth answer, balancing precision and recall and handling partial matches. Exact match (C) is correct because it reports the percentage of predictions that exactly equal the reference answer, a standard metric in QA benchmarks such as SQuAD. ROUGE-L score (E) is correct because it measures the longest common subsequence between generated and reference text, capturing fluency and recall-oriented overlap useful for generative QA outputs.

Cost per training epoch (A) and training time per epoch (D) are operational/training-efficiency measures, not model quality metrics, so they do not evaluate performance before deployment.

Exam trap

The trap here is that candidates confuse operational metrics (like cost or training time) with evaluation metrics that directly measure model output quality, leading them to select options that are irrelevant to performance assessment.

208
Multi-Selectmedium

A data science team is choosing between Gemini Pro and Gemini Flash for a real-time content moderation API. They need low cost and low latency, but can tolerate slightly lower accuracy. Which TWO Gemini variants should they consider?

Select 1 answer
A.Gemini Nano
B.Gemini Flash
C.Gemini 1.5 Pro
D.Gemini Ultra
E.Gemini Pro
AnswersB

Gemini Flash is designed for low latency and low cost, fitting the primary requirements.

Why this answer

Gemini Flash (B) is correct because it is the latency-optimized, lower-cost model tier designed for high-throughput, real-time tasks where slightly reduced accuracy is acceptable, matching the team's priority of low latency and low cost. Gemini Nano (A) is also a low-cost, low-latency variant, but it is intended for on-device/local use cases rather than a server-side real-time API, so it is not the best fit here. Gemini Pro (E), Gemini 1.5 Pro (C), and Gemini Ultra (D) are excluded because they are higher-capability, higher-cost, higher-latency models that conflict with the stated low-cost and low-latency requirements.

Because only Flash clearly satisfies all constraints, the question should be revised to have a single correct answer or to use distinct, non-overlapping model options.

Exam trap

Google often tests the distinction between model variants (Nano, Flash, Pro, Ultra) and their specific trade-offs, trapping candidates who confuse 'lowest cost' with 'best for real-time' or who assume Pro is a low-latency option when it is actually higher cost and higher latency than Flash.

209
MCQeasy

An organization wants to use AI-generated images commercially. According to Google's AI principles and copyright guidelines, what should they do FIRST?

A.Assume all AI-generated content is copyright-free
B.Only use images generated by models trained on public domain data
C.Add a copyright symbol to all AI-generated images
D.Verify the training data provenance and licensing of the foundation model
AnswerD

Commercial reuse depends on the foundation model's training data provenance and licensing, since copyright risk originates there rather than in the generated output. Verifying these terms first establishes whether commercial use is permissible at all, satisfying the requirement to check rights before generating or publishing any images.

Why this answer

The primary risk is copyright infringement from training data. Verifying that the model's training data was properly licensed is essential before using outputs commercially.

210
MCQmedium

A company develops a generative AI model for resume screening. They discover that the model is rejecting candidates from certain demographic groups disproportionately. Which step should they take first to address unfair bias?

A.Audit the training data for demographic representativeness
B.Reduce the model's complexity to avoid overfitting to biased patterns
C.Apply adversarial debiasing to the model
D.Collect more data from all demographic groups equally
AnswerA

Auditing training data for demographic representativeness identifies whether skewed historical hiring data caused the disparate rejection rates. Fixing the data source precedes mitigation techniques such as reweighting, threshold adjustment or post-processing, which treat symptoms rather than the underlying cause.

Why this answer

Before attempting to fix bias, it is essential to evaluate the training data for representativeness. Bias often originates from imbalanced or unrepresentative training data.

211
Multi-Selecteasy

Which TWO are advantages of using Retrieval-Augmented Generation (RAG) over fine-tuning?

Select 2 answers
A.No need to retrain the base model
B.Requires less data preparation
C.Lower inference latency
D.More secure because model weights are not modified
E.Better suited for rapidly changing knowledge bases
AnswersA, E

Correct: RAG works with the pre-trained model and a retrieval system.

Why this answer

RAG retrieves relevant external knowledge at inference time without modifying the base model's parameters, eliminating the need for retraining (option A). This contrasts with fine-tuning, which requires updating model weights through additional training cycles. Because the knowledge lives in an external, updatable index rather than in the model weights, RAG is also better suited for rapidly changing knowledge bases (option E): documents can be added, updated, or removed without any retraining, whereas fine-tuning would require a new training run each time the knowledge changes.

RAG thus preserves the original model while augmenting its output with up-to-date information.

Exam trap

Google often tests the misconception that RAG is always faster or simpler than fine-tuning, but candidates must remember that retrieval adds latency and requires careful data preprocessing, making options B and C tempting but incorrect.

212
MCQeasy

A developer is using the Gemini API to generate creative marketing copy. They want the output to be more diverse and unexpected. Which parameter should they increase?

A.Temperature.
B.Presence penalty.
C.Top-p.
D.Frequency penalty.
AnswerA

Temperature controls sampling randomness by scaling the probability distribution over candidate tokens before selection. Raising it flattens that distribution, so lower-probability tokens are chosen more often, directly producing the diverse, unexpected marketing copy the developer wants.

Why this answer

Increasing the temperature parameter makes the model's output probabilities more uniform, encouraging it to sample less likely tokens and produce more diverse, unexpected, and creative text. A higher temperature (e.g., >1.0) flattens the probability distribution, so the model is more likely to choose surprising word combinations rather than the most probable ones.

Exam trap

A common pitfall is confusing 'diversity' with 'avoiding repetition.' Increasing temperature or top-p increases randomness and diversity, while presence and frequency penalties reduce repetition. Candidates may incorrectly choose presence or frequency penalties thinking they increase diversity.

How to eliminate wrong answers

Option B (Presence penalty) is wrong because it penalizes tokens that have already appeared in the text, which reduces repetition but does not directly increase diversity or unexpectedness in the same way as temperature; it can still lead to predictable token choices. Option C (Top-p) is wrong because it controls the cumulative probability threshold for token sampling (nucleus sampling), which can limit diversity by cutting off the long tail of low-probability tokens, making outputs less unexpected. Option D (Frequency penalty) is wrong because it reduces the likelihood of tokens based on how often they have appeared, which primarily discourages repetition but does not flatten the probability distribution to encourage surprising token choices.

213
Multi-Selecteasy

A team is deciding between using fine-tuning and in-context learning for a document classification task. They have 500 labeled examples and need low latency. Which TWO statements are accurate?

Select 2 answers
A.In-context learning always has lower latency than fine-tuning
B.In-context learning can only be used with models that have a context window smaller than 1000 tokens
C.Fine-tuning eliminates the need for a validation dataset
D.Fine-tuning generally improves accuracy more than in-context learning when sufficient labeled data is available
E.In-context learning requires no training step and can be used immediately
AnswersD, E

Fine-tuning adjusts the model's weights on the 500 labelled examples, embedding task-specific patterns directly into parameters rather than supplying them at inference. This yields higher accuracy than in-context learning once labelled data is sufficient, satisfying the scenario's classification goal, though it does not address the low-latency constraint.

Why this answer

Option D is correct because fine-tuning updates the model's weights on the 500 labeled examples, allowing it to learn task-specific patterns and typically achieve higher accuracy than in-context learning when a sufficient amount of labeled data is available. Option E is correct because in-context learning places examples directly in the prompt at inference time, requiring no gradient updates or training step, so it can be used immediately with a pretrained model. Option A is incorrect because in-context learning increases prompt length and thus inference latency, so it does not always have lower latency than fine-tuning.

Option B is incorrect because in-context learning works with any model context window and is not limited to windows smaller than 1000 tokens. Option C is incorrect because fine-tuning still requires a validation dataset to tune hyperparameters and monitor for overfitting.

Exam trap

Google often tests the misconception that in-context learning always has lower latency than fine-tuning, but the trap is that latency depends on prompt length and model architecture, not just the absence of a training step.

214
MCQeasy

Which Google Cloud offering allows you to create a machine-readable document that describes a model's intended use, performance, and limitations?

A.People + AI Guidebook
B.Vertex AI Model Registry
C.Datasheets for Datasets
D.Model Cards
AnswerD

Model Cards are the structured, machine-readable artefact documenting a model's intended use, performance metrics, and known limitations. They satisfy the requirement for a standardised description that downstream consumers and governance tools can parse, unlike prose documentation or general transparency reports.

Why this answer

Model Cards are structured documents that describe a machine learning model's intended use, performance metrics, limitations, and ethical considerations. They are designed to be machine-readable and provide transparency for stakeholders. Google Cloud's Model Cards are part of Vertex AI and can be generated and stored alongside models in the Model Registry.

Exam trap

The trap is confusing Model Cards with Datasheets for Datasets or the Model Registry; candidates must remember that Model Cards describe models, while Datasheets describe datasets, and the Registry is a management tool.

How to eliminate wrong answers

Option A is wrong because the People + AI Guidebook is a set of design principles for human-centered AI, not a machine-readable document describing a specific model. Option B is wrong because Vertex AI Model Registry is a repository for managing models, not the document itself. Option C is wrong because Datasheets for Datasets describe datasets, not models.

215
MCQmedium

A media company wants its editorial staff to draft blog posts inside a web-based workspace where Gemini can summarize uploaded research PDFs, generate outlines, and cite files from the team's shared drive, all without writing code or managing any Google Cloud infrastructure. Which Google Cloud generative AI offering best fits this requirement?

A.Google Cloud Vertex AI Agent Builder with a custom tool
B.Gemini for Google Workspace
C.Gemini Enterprise with a connected data store
D.Vertex AI Studio with a tuned Gemini model endpoint
AnswerB

Gemini for Google Workspace embeds Gemini directly in Gmail, Docs, Drive, and related apps, letting non-technical staff summarize PDFs, draft content, and ground responses in files they can already access. It requires no infrastructure work, which matches the editorial team's need for a no-code workspace experience with shared-drive content available in context.

Why this answer

Gemini for Google Workspace is the offering designed to bring Gemini into the productivity apps employees already use, with enterprise-grade privacy protections and grounding in content the user can access. Because the scenario calls for drafting and summarizing inside a workspace with no coding or infrastructure, the embedded Workspace experience is the natural fit rather than developer consoles or agent platforms.

Exam trap

The trap here is assuming any Gemini-branded managed service delivers in-app Workspace drafting, when most Gemini offerings are developer or standalone-assistant surfaces rather than features embedded in Docs and Drive.

216
MCQmedium

A company wants to use AI-generated images commercially. They are concerned about copyright and IP issues. Which action should they take FIRST to mitigate legal risk?

A.Conduct a freedom-to-operate search each time an image is generated
B.Add a watermark to all generated images using SynthID
C.Ensure the AI model's training data does not contain copyrighted material without permission
D.Purchase a commercial license for the AI platform
AnswerC

Verifying training data licensing addresses the root cause: copyright exposure originates from the material used to train the model. Confirming permission or excluding protected works first prevents infringing outputs, which downstream filters or disclaimers cannot fully remedy.

Why this answer

The foundational legal risk in AI-generated imagery stems from the training data. If the model was trained on copyrighted works without permission, any output—even a novel image—can be considered a derivative work, exposing the company to infringement claims. Addressing the training data's compliance is the first and most critical step, as downstream mitigations (like watermarks or licenses) cannot retroactively fix an unlawfully trained model.

Exam trap

This exam often tests the misconception that purchasing a commercial license or adding a watermark is sufficient to avoid copyright liability, when in fact the primary legal risk originates from the training data's compliance with copyright law.

How to eliminate wrong answers

Option A is wrong because conducting a freedom-to-operate search after each generation is impractical at scale and legally insufficient; copyright law does not require a search for infringement, and the output's similarity to training data is often non-obvious, so a search cannot reliably detect latent infringement. Option B is wrong because adding a watermark (e.g., SynthID) only identifies the image as AI-generated but does not address the underlying copyright status of the training data or the output; it is a transparency tool, not a legal risk mitigator. Option D is wrong because purchasing a commercial license for the AI platform typically covers the platform's own IP rights (e.g., the model weights) but does not indemnify the user against third-party copyright claims arising from the training data; many licenses explicitly disclaim such liability.

217
MCQhard

Refer to the exhibit. A Vertex AI endpoint configured with the above deployment is returning HTTP 429 (Too Many Requests) errors during peak traffic. The current CPU utilization reaches 80% consistently. What should the team adjust to resolve this?

A.Increase maxReplicaCount to 10
B.Increase scaleTarget to 0.9
C.Change machineType to n1-highmem-2
D.Increase minReplicaCount to 2
AnswerA

Correct: Higher max allows more replicas to handle traffic spikes.

Why this answer

Increasing maxReplicaCount to 10 allows the Vertex AI endpoint to scale out to more instances during peak traffic, distributing the load and reducing HTTP 429 errors. Since CPU utilization is at 80%, the current maxReplicaCount is insufficient to handle the demand, and raising this limit enables the horizontal pod autoscaler to add replicas up to the new maximum, directly addressing the capacity bottleneck.

Exam trap

The Google Cloud Gen AI Leader exam often tests the distinction between scaling limits (min/maxReplicaCount) and scaling thresholds (scaleTarget), trapping candidates who confuse raising the scaling target with increasing capacity.

How to eliminate wrong answers

Option B is wrong because increasing scaleTarget (the CPU utilization threshold for scaling) to 0.9 (90%) would actually delay scaling, making the 429 errors worse as the endpoint would wait until CPU is even higher before adding replicas. Option C is wrong because changing machineType to n1-highmem-2 (a memory-optimized machine) does not address the CPU bottleneck; the issue is insufficient compute capacity, not memory pressure. Option D is wrong because increasing minReplicaCount to 2 ensures a baseline of 2 replicas but does not raise the upper scaling limit, so during peak traffic the endpoint still cannot scale beyond the current maxReplicaCount, leaving it vulnerable to overload.

218
MCQeasy

A developer wants to integrate Gemini multimodal capabilities (text + image) into a mobile app using Python. Which Google Cloud client library should they use?

A.Dialogflow CX
B.Vertex AI client library (google-cloud-aiplatform)
C.Cloud Vision API
D.Natural Language API
AnswerB

The google-cloud-aiplatform library exposes Vertex AI's Gemini multimodal endpoints, handling text and image inputs natively from Python. It satisfies the stem's requirement to integrate Gemini text-plus-image capabilities into a mobile app backend, unlike single-modality or non-Vertex client libraries.

Why this answer

The Vertex AI client library (google-cloud-aiplatform) provides the Generative AI SDK that supports multimodal capabilities, including the ability to send both text and image inputs to Gemini models. This library directly exposes the `GenerativeModel` class with methods like `generate_content()` that accept `Part` objects containing image data (e.g., `Part.from_image()` or `Part.from_uri()`), making it the correct choice for integrating Gemini multimodal features into a Python mobile app backend.

Exam trap

The trap here is that candidates confuse specialized single-modality APIs (Vision, Natural Language) with the unified multimodal API provided by Vertex AI, assuming that combining separate services is equivalent to Gemini's native multimodal reasoning.

How to eliminate wrong answers

Option A is wrong because Dialogflow CX is a conversational AI platform for building chatbots and virtual agents, not a library for directly accessing Gemini multimodal models; it lacks the low-level API to construct multimodal requests with image parts. Option C is wrong because Cloud Vision API is a specialized service for image analysis (e.g., object detection, OCR) and does not provide access to Gemini's generative multimodal capabilities or its text+image reasoning. Option D is wrong because Natural Language API is designed for text-only analysis (e.g., sentiment, entity extraction) and cannot process image inputs or generate multimodal responses.

219
MCQhard

A company uses Gemini 1.5 Pro to analyze customer call transcripts and generate summaries. They notice that the summaries occasionally include fabricated details that were not in the transcript. Which technique is specifically designed to reduce such hallucinations?

A.Ground the model responses by implementing Retrieval-Augmented Generation (RAG) with the transcript as a source
B.Decrease the temperature to 0.0
C.Use a system prompt that instructs the model not to make up information
D.Fine-tune the model on a larger dataset of call transcripts
AnswerA

RAG retrieves the actual transcript and supplies it as grounding context, so the model conditions its summary on real source text rather than parametric memory. This constrains generation to verifiable transcript content, directly reducing fabricated details.

Why this answer

Grounding with RAG retrieves factual data from trusted sources to condition the model's response, directly reducing hallucinations. Prompt engineering can help but does not guarantee factual accuracy; fine-tuning may reduce but not eliminate; temperature reduction makes output more deterministic but does not address factuality from external data.

220
MCQhard

A multinational corporation deploys a generative AI chatbot for customer support in the EU. They must ensure compliance with GDPR regarding user data used for fine-tuning. Which data governance practice is REQUIRED?

A.Store user data only in the EU region to comply with data residency requirements
B.Obtain explicit consent from users for their data to be used in fine-tuning
C.Implement a mechanism to delete specific user data from the fine-tuning dataset upon user request
D.Anonymize all user data before using it for fine-tuning
AnswerB

Consent is required, but it does not address the right to erasure after fine-tuning has occurred.

Why this answer

Under GDPR, using personal data for fine-tuning a generative AI model is a new processing purpose and requires its own lawful basis. Where consent is used, it must be freely given, specific, informed and unambiguous, and explicit consent is required for special-category or unexpected secondary uses. Option B identifies the mandatory legal requirement: obtaining explicit consent from users for their data to be used in fine-tuning.

Option C addresses a data subject right that may apply once data is already processed, but it is not the primary upfront governance practice required for using personal data in fine-tuning.

Exam trap

A common trap is conflating privacy-enhancing best practices (anonymization, regional storage, model unlearning) with strict GDPR legal requirements. Candidates should identify the mandatory lawful-basis requirement — explicit consent for using personal data in fine-tuning — rather than selecting a best-practice option.

How to eliminate wrong answers

Option A is wrong because GDPR does not mandate data storage solely within the EU; it allows transfers to third countries with adequate safeguards (e.g., Standard Contractual Clauses), so regional storage is not a strict requirement. Option B is wrong because explicit consent is one lawful basis for processing, but GDPR also permits legitimate interest or contractual necessity for fine-tuning, so consent is not always required. Option D is wrong because anonymization is a recommended privacy technique but not a mandated practice; GDPR requires data minimization and purpose limitation, but does not force anonymization before fine-tuning.

221
MCQmedium

A bank wants to use LLMs to generate responses for customer support chat. All conversations must be logged, and any PII must be masked. The solution must comply with financial regulations. Which combination of Vertex AI services should be used?

A.Deploy a custom model on Cloud Run and write a Cloud Function to mask PII.
B.Use Vertex AI Prediction with a custom container that masks PII before inference.
C.Use the Gemini API directly with a custom logging solution in Cloud Logging.
D.Use Vertex AI Agent Builder with Data Governance, which can automatically mask PII and log interactions.
AnswerD

Vertex AI Agent Builder with Data Governance satisfies both constraints directly: Data Governance applies automatic PII masking (DLP-based de-identification) to prompts and responses, while Agent Builder logs full conversation interactions for audit. This meets the bank's regulatory requirement for masked, retained chat records without custom pipeline work.

Why this answer

Vertex AI Agent Builder integrates with Data Governance to automatically mask PII and log interactions, meeting both the logging and compliance requirements without custom development. This managed service ensures adherence to financial regulations by providing built-in data loss prevention (DLP) capabilities and audit trails, unlike the other options which require manual or less integrated approaches.

Exam trap

Google Cloud often tests the misconception that custom development (e.g., Cloud Functions or custom containers) is necessary for PII masking and logging, when in fact managed services like Vertex AI Agent Builder with Data Governance provide a more compliant and integrated solution out of the box.

How to eliminate wrong answers

Option A is wrong because deploying a custom model on Cloud Run with a Cloud Function for PII masking introduces operational complexity and latency, and does not natively integrate with Vertex AI's logging or compliance features, risking gaps in regulatory adherence. Option B is wrong because using Vertex AI Prediction with a custom container that masks PII before inference still requires custom development for logging and does not leverage Vertex AI's built-in data governance, making it harder to ensure consistent compliance across all interactions. Option C is wrong because using the Gemini API directly with a custom logging solution in Cloud Logging lacks automatic PII masking and data governance, forcing manual implementation that is error-prone and may not meet strict financial regulations for auditability and data protection.

222
Multi-Selectmedium

A company is building a generative AI application that must adhere to strict data residency regulations. Which TWO Google Cloud features can help ensure that data does not leave a specific geographic region?

Select 2 answers
A.Using the regional endpoint for Vertex AI
B.Vertex AI Model Caching
C.Global load balancer with Cloud Armor
D.Cloud CDN for content delivery
E.Deploying models on dedicated VMs in a specific region
AnswersA, E

Regional endpoints ensure API calls stay within the region.

Why this answer

To ensure data residency, you can use the regional endpoint for Vertex AI (A) which ensures that API calls and data processing stay within the specified region. Additionally, deploying models on dedicated VMs in a specific region (E) ensures that compute resources and data do not leave that region. Option B (Vertex AI Model Caching) does not enforce residency as caching may use regional resources but does not guarantee data stays within a region.

Option C (Global load balancer with Cloud Armor) is a network security and load balancing service that operates globally. Option D (Cloud CDN) caches content at edge locations globally, which would move data outside the region.

223
MCQmedium

Refer to the exhibit. A team has deployed a model to an endpoint with the configuration shown. They notice that during peak traffic, the endpoint frequently returns 429 (Too Many Requests) errors. Which action should they take to resolve this issue?

A.Change MACHINE_TYPE to n1-highmem-4
B.Increase MIN_REPLICA_COUNT to 5
C.Decrease MAX_REPLICA_COUNT to 1
D.Disable autoscaling by setting MIN_REPLICA_COUNT equals MAX_REPLICA_COUNT
AnswerB

429 errors indicate the endpoint's replicas cannot absorb peak request volume. Raising MIN_REPLICA_COUNT to 5 keeps more serving replicas warm, increasing concurrent request capacity and distributing traffic so the quota per replica is no longer exceeded.

Why this answer

The 429 (Too Many Requests) errors indicate the endpoint is receiving more concurrent requests than its current replica capacity can handle. Increasing MIN_REPLICA_COUNT raises the baseline number of serving replicas, so the endpoint has more capacity to absorb peak traffic without throttling. This directly addresses the throughput bottleneck rather than changing the instance shape or disabling scaling.

Exam trap

The trap here is confusing vertical scaling (bigger machine type) with horizontal scaling (more replicas); 429 errors are a concurrency/throughput signal, not a memory or CPU-per-instance signal.

How to eliminate wrong answers

Option A is wrong because changing MACHINE_TYPE to n1-highmem-4 alters memory-to-CPU ratio but does not increase the number of serving replicas, so concurrent request capacity is not meaningfully expanded and 429s can persist. Option C is wrong because decreasing MAX_REPLICA_COUNT to 1 caps the endpoint at a single replica, reducing capacity and worsening throttling under peak load. Option D is wrong because setting MIN_REPLICA_COUNT equal to MAX_REPLICA_COUNT disables autoscaling entirely, freezing replica count and preventing the endpoint from scaling out to meet peak demand.

224
MCQmedium

A financial services company wants to automate contract analysis to extract key clauses and identify risky terms. They have thousands of PDF contracts and need a solution that can be quickly integrated into their existing document management system. Which Google Cloud service is MOST suitable?

A.Model Garden
B.Document AI (DocAI)
C.Vertex AI Studio
D.Vertex AI Agent Builder
AnswerB

Document AI provides pre-trained and custom processors that extract clauses and entities from PDFs, integrating through APIs into existing document management systems. This meets the requirement for rapid integration and accurate contract analysis at scale without building models from scratch.

Why this answer

DocAI is purpose-built for document understanding and structured extraction from PDFs. Vertex AI Studio focuses on prompt design for generative models, while Model Garden is for selecting models. Vertex AI Agent Builder is for building conversational agents, not document processing.

225
MCQhard

Refer to the exhibit. A data scientist is fine-tuning a model. The training loss and accuracy are improving each epoch. However, after training, the model performs poorly on a held-out validation set. What is the most likely issue?

A.Underfitting
B.Inappropriate learning rate
C.Data leakage
D.Overfitting
AnswerD

Overfitting occurs when the model memorises training data, so training loss and accuracy keep improving while generalisation to unseen data degrades. That mismatch between strong training metrics and poor held-out validation performance is precisely the stem's symptom.

Why this answer

The model's training loss and accuracy improve each epoch, but performance on the validation set is poor. This classic symptom indicates overfitting, where the model memorizes the training data (including noise) rather than learning generalizable patterns. In fine-tuning, this often occurs when the model is trained for too many epochs or the dataset is too small relative to model capacity.

Exam trap

Google Cloud often tests the distinction between overfitting and underfitting by presenting improving training metrics alongside poor validation performance, which candidates may misinterpret as a learning rate issue or data leakage if they do not recognize the hallmark divergence pattern.

How to eliminate wrong answers

Option A is wrong because underfitting would show poor performance on both training and validation sets, not improving training metrics. Option B is wrong because an inappropriate learning rate typically causes training instability (e.g., loss divergence or stagnation), not a clear divergence between training and validation performance. Option C is wrong because data leakage would cause both training and validation metrics to be artificially high (since validation data leaks into training), not a gap where training is good and validation is poor.

Page 2

Page 3 of 14

Page 4