Courseiva

Google Cloud Generative AI Leader Generative AI Leader (Generative AI Leader) — Questions 451–525

1008 questions total · 14pages · All types, answers revealed

Page 6

Page 7 of 14

Page 8
451
MCQmedium

A customer support team uses a foundation model via the Gemini API to answer billing questions. Responses are often vague and sometimes omit required steps. The team wants to improve output quality without fine-tuning the model. Which approach should they use?

A.Use prompt engineering with few-shot examples and a clear output format that lists the required steps in order.
B.Lower the top-p value so the model only considers the most probable tokens, which should make the answers more accurate.
C.Increase the maximum output tokens so the model has more room to produce a longer, more complete answer.
D.Raise the temperature setting so the model can explore more possible answers and include more detail in its responses.
AnswerA

Few-shot examples show the model the desired structure and level of detail, while an explicit format forces the required steps to appear. This steers the base model without any training, which matches the constraint of not fine-tuning. It directly addresses both vagueness and missing steps by making the expected answer shape part of the prompt.

Why this answer

Prompt engineering with few-shot examples and an enforced output format is the fastest way to raise quality without changing the model weights. Examples demonstrate the expected depth, and a structured template guarantees that required steps appear. Sampling parameters alter randomness, and token limits alter length, but neither supplies missing procedural knowledge or enforces completeness.

Exam trap

The trap here is assuming that sampling parameters such as temperature or top-p can add missing factual steps, when they only change randomness and cannot supply required content.

452
MCQmedium

A company wants to build a proof-of-concept generative AI application quickly. They have limited ML expertise and need to test multiple foundation models with different prompts. Which Vertex AI tool should they use?

A.Vertex AI Studio
B.Vertex AI Agent Builder
C.Model Garden
D.Vertex AI Training
AnswerA

Vertex AI Studio provides a no-code console for prompt design and rapid testing across multiple foundation models, including Gemini and partner models. This directly satisfies the limited ML expertise and multi-model prompt experimentation constraints, letting the team validate the proof-of-concept without building training or deployment pipelines.

Why this answer

Vertex AI Studio is designed for rapid prototyping with prompt design, model testing, and evaluation without requiring coding. Model Garden is for browsing models but not for interactive testing. Agent Builder is for building agents.

Custom training is for production.

453
MCQhard

A bank's risk team must review every AI-generated customer communication for compliance before it is sent. They want to enforce a policy that blocks any message containing prohibited financial advice and routes flagged messages to a human reviewer. Which Vertex AI capability should they configure?

A.Vertex AI Model Monitoring with skew and drift detection
B.A custom classifier deployed on Vertex AI Endpoints and invoked in the generation workflow
C.Vertex AI safety filters configured with adjustable harm thresholds
D.Vertex AI Model Evaluation with an automatic metric threshold
AnswerB

A custom classifier trained on examples of compliant and non-compliant communications can detect prohibited financial advice, and deploying it to a Vertex AI Endpoint lets the application call it after generation. The workflow can then block flagged text and route it to a human reviewer, satisfying the compliance policy.

Why this answer

The bank needs an inline, policy-specific gate on each generated message. A custom classifier trained on compliant and non-compliant examples, deployed to a Vertex AI Endpoint, can be called in the generation pipeline to score each message; the application then blocks or routes based on that score. Monitoring, standard safety filters, and offline evaluation do not provide real-time domain policy enforcement and routing.

Exam trap

The trap here is assuming that built-in safety filters cover organization-specific compliance policies like prohibited financial advice, when such policies require custom detection logic.

454
MCQhard

A healthcare provider uses a Gemini model on Vertex AI to summarize patient intake notes for clinicians. The summaries must consistently follow a fixed structure: chief complaint, history, medications, and plan. Early tests show the model sometimes reorders or omits sections. Which technique most reliably enforces the required structure?

A.Provide few-shot examples in the prompt that show the exact section order and format.
B.Fine-tune the model on a large corpus of unstructured clinical literature.
C.Increase the temperature so the model varies its section ordering.
D.Reduce the maximum output tokens to force the model to be concise.
AnswerA

Few-shot examples demonstrate the precise structure the model should follow, and in-context patterns strongly shape output format for that request. By showing several correctly ordered summaries, the prompt makes the expected sections and sequence explicit, which reliably improves adherence for a fixed template. It requires no model change and can be updated quickly as the template evolves, making it well suited to structured clinical summaries.

Why this answer

Few-shot examples are the most reliable prompt-level technique for enforcing a fixed output structure, because they show the model the exact sections and their order within the same request. They need no retraining and can be revised as the template changes. Temperature changes increase disorder, broad unstructured fine-tuning does not teach the template, and token limits can drop sections rather than order them.

Exam trap

The trap here is reaching for fine-tuning or length limits to fix a formatting problem, when demonstrating the desired structure with in-context examples is the direct and reliable control.

455
MCQeasy

What is the primary purpose of the temperature parameter when generating text with an LLM?

A.It sets the maximum number of tokens in the response
B.It controls the randomness or creativity of the output
C.It determines the number of most likely tokens considered at each step
D.It sets the cumulative probability threshold for token selection
AnswerB

Temperature scales the probability distribution over the model's next-token predictions before sampling. A low value sharpens the distribution, favouring high-probability tokens for deterministic, factual output; a high value flattens it, giving unlikely tokens more chance and producing varied, creative text. This directly governs the randomness the question asks about.

Why this answer

Temperature controls the randomness of token selection. Higher temperature increases creativity, lower temperature makes outputs more deterministic.

456
MCQmedium

A company wants to use Google DeepMind's advances in protein structure prediction to accelerate drug discovery. Which DeepMind achievement is most relevant to this goal?

A.AlphaFold
B.WaveNet
C.AlphaCode
D.AlphaGo
AnswerA

AlphaFold predicts three-dimensional protein structures directly from amino acid sequences, solving a decades-old challenge in structural biology. This capability lets researchers model target proteins and potential binding sites computationally, dramatically accelerating lead identification and drug discovery timelines compared with slower experimental methods such as X-ray crystallography.

Why this answer

AlphaFold solves protein structure prediction, a key enabler for drug discovery. AlphaGo is for board games, AlphaCode for programming, and WaveNet for audio.

457
MCQhard

A model generates biased output. Which technique is least effective?

A.Use adversarial debiasing
B.Apply safety filters
C.Set frequency penalty to 1.0
D.Fine-tune on diverse data
AnswerC

Frequency penalty only discourages repeated tokens; it cannot correct bias baked into training data or model weights. Bias mitigation requires data curation, fine-tuning or prompt design, so this setting leaves the underlying cause untouched and is therefore least effective.

Why this answer

Setting the frequency penalty to 1.0 is least effective for reducing biased output because frequency penalties reduce repetition of tokens based on their frequency in the generated text, not their association with protected attributes or fairness. This parameter controls lexical diversity, not demographic parity or representational harm, so it does not address the root cause of bias in the model's training data or inference logic.

Exam trap

Google often tests the misconception that any hyperparameter affecting output diversity (like frequency penalty) can mitigate bias, when in fact bias mitigation requires targeted techniques that address representation, fairness, or safety directly.

How to eliminate wrong answers

Option A is wrong because adversarial debiasing directly trains the model to minimize a discriminator's ability to predict protected attributes from the model's representations, actively reducing bias. Option B is wrong because safety filters can block or flag outputs containing harmful stereotypes or slurs, providing a post-hoc mitigation layer against biased content. Option D is wrong because fine-tuning on diverse data reweights the training distribution to include underrepresented groups, reducing statistical bias in the model's predictions.

458
MCQmedium

An enterprise wants to use a foundation model through Google Cloud but must ensure that their prompts and responses are not used to train the underlying model and that data is encrypted in transit and at rest. They also want to avoid managing infrastructure. Which approach best meets these requirements?

A.Use a third-party public chatbot website to test prompts and then copy the results into internal systems.
B.Fine-tune a model on the enterprise's confidential prompts so the model learns the company's style.
C.Use Vertex AI generative AI models through the Google Cloud console and API under the enterprise's own project and data-governance controls.
D.Download an open-source model and run it on a Compute Engine VM that the team patches and scales manually.
AnswerC

Vertex AI provides managed access to foundation models without infrastructure management, and Google Cloud's data processing terms state that customer prompts and responses are not used to train foundation models. Traffic is encrypted in transit and data at rest is encrypted by default, aligning with all stated requirements within the customer's project boundary.

Why this answer

Vertex AI offers managed access to generative models with enterprise controls inside the customer's Google Cloud project. Google Cloud's terms specify that customer data for foundation models is not used to train those models, and encryption in transit and at rest is provided by default. This combination satisfies the governance and no-infrastructure-management requirements.

Exam trap

The trap here is equating any model access with enterprise-grade governance, when public chatbots and self-managed deployments miss the training-data and infrastructure requirements.

459
Multi-Selectmedium

Which TWO options are benefits of using Vertex AI Model Garden compared to using raw pre-trained models from external sources? (Choose two.)

Select 2 answers
A.Lower cost compared to using generic APIs
B.Ability to fine-tune models on custom data
C.Integration with Vertex AI tools like evaluation and monitoring
D.Simplified deployment and scaling with Vertex AI endpoints
E.Guaranteed data privacy and no data sharing
AnswersC, D

Model Garden models are natively wired into Vertex AI evaluation and monitoring, giving drift detection and quality metrics without custom integration work. Raw external models lack this managed observability, which is the specific benefit here.

Why this answer

Option C is correct because Model Garden models are first-class Vertex AI resources, so they plug directly into Vertex AI Evaluation and Vertex AI Model Monitoring (including skew/drift detection) without custom glue code. Option D is correct because Model Garden provides one-click or API-driven deployment to Vertex AI Endpoints, which handles autoscaling, traffic splitting, and serving infrastructure that you would otherwise build yourself for an externally sourced model. Option A is not inherently true — Model Garden includes both open and partner/proprietary models, and pricing varies, so it is not a guaranteed cost benefit.

Option B is not unique to Model Garden, since raw pre-trained models from external sources can also be fine-tuned. Option E is not guaranteed, as data privacy depends on the specific model's terms and your configuration, not on Model Garden itself.

Exam trap

Google Cloud often tests the distinction between inherent platform benefits (like integration and managed deployment) versus features that are not exclusive to Model Garden (like fine-tuning or cost), leading candidates to mistakenly select options that are generally true for any model but not unique advantages of Model Garden.

460
MCQhard

A logistics company needs an assistant that answers driver questions about shipment status. The answers must reflect live data from an operational database and must include citations so dispatchers can verify each claim. The company wants a managed Google Cloud capability rather than custom retrieval code. Which capability should they use?

A.Increasing the Gemini model's temperature setting so responses draw more broadly on the model's internal knowledge.
B.Exporting the operational database nightly to a Cloud Storage bucket and pasting relevant rows into each prompt manually.
C.Fine-tuning the Gemini model on historical shipment records so it memorizes typical status patterns.
D.Grounding with Vertex AI Search data stores connected to the operational data, used through Vertex AI Agent Builder.
AnswerD

Vertex AI Search data stores can index enterprise sources, and when used for grounding they return responses with citations pointing back to the retrieved passages. Agent Builder orchestrates the assistant and surfaces those citations, which satisfies both the live-data and verifiability requirements using managed components instead of custom retrieval code written by the company.

Why this answer

The requirements are live data plus citations, delivered through a managed capability. Grounding with Vertex AI Search data stores retrieves from indexed enterprise sources and returns citations, and Agent Builder provides the managed orchestration layer that surfaces them. Temperature changes, manual exports, and fine-tuning cannot supply live lookups or verifiable source references, so they fail the scenario regardless of model quality.

Exam trap

The trap here is believing fine-tuning can substitute for retrieval, when tuning only adjusts model behavior and never grants access to current records or produces source citations.

461
Multi-Selecthard

Which THREE capabilities are provided by Vertex AI Agent Builder? (Choose three.)

Select 3 answers
A.Automated model hyperparameter tuning.
B.Integration with Dialogflow CX for conversational flows.
C.Support for multimodal (text, image, video) input processing in agents.
D.Creating custom agents with memory and tool integration.
E.Built-in grounding with Google Search to improve answer accuracy.
AnswersB, D, E

Agent Builder can leverage Dialogflow CX for advanced conversational design.

Why this answer

Vertex AI Agent Builder integrates with Dialogflow CX to enable the design of sophisticated conversational flows, including state management, conditional logic, and multi-turn interactions. This allows developers to build agents that can handle complex dialogues with branching paths, leveraging Dialogflow CX's visual flow builder and fulfillment capabilities within the Vertex AI ecosystem.

Exam trap

The Generative AI Leader exam often tests the distinction between Vertex AI Agent Builder's core capabilities (like Dialogflow CX integration, custom agents with memory/tools, and grounding with Google Search) and features that belong to other Vertex AI services, such as hyperparameter tuning in Vertex AI Training or multimodal support that excludes video in the agent builder context.

462
MCQeasy

A company is using a generative AI model to answer customer questions. The answers are often vague and do not use the company's specific terminology. They want to improve the relevance and specificity of the responses. Which technique should they use?

A.Using retrieval-augmented generation (RAG)
B.Increasing the model's temperature
C.Fine-tuning the model on company documents
D.Reducing the max output tokens
AnswerA

RAG combines a generative model with a retrieval system that fetches relevant documents from a knowledge base. By grounding responses in company-specific documents, the model can produce more accurate and specific answers that use the correct terminology, directly addressing the issue.

Why this answer

Retrieval-augmented generation (RAG) enhances a generative model by retrieving relevant information from a knowledge base and including it in the prompt. This grounds the model's responses in factual, company-specific content, improving relevance and ensuring the use of correct terminology. It is efficient because it does not require retraining the model.

Exam trap

The trap here is assuming that fine-tuning is always necessary to inject domain knowledge, when RAG can achieve similar results more quickly and with less data.

463
Multi-Selecthard

Which THREE techniques are commonly used to improve the overall quality and coherence of generative model outputs? (Choose three.)

Select 3 answers
A.Using self-consistency or iterative refinement to choose the best output.
B.In-context learning (few-shot prompting) with relevant examples.
C.Applying output safety filters to remove inappropriate content.
D.Prompt chaining to decompose complex tasks into simpler sub-tasks.
E.Random sampling to increase output diversity.
AnswersA, B, D

Iterative methods improve reliability and coherence by selecting the most consistent response.

Why this answer

Self-consistency (A) improves quality by generating multiple candidate outputs from the same prompt and selecting the most consistent or frequent answer, reducing variance and errors, and iterative refinement lets the model revise its own output via feedback loops. In-context learning (B) with relevant few-shot examples steers the model toward the desired format, style, and reasoning pattern, improving output quality without retraining. Prompt chaining (D) decomposes a complex task into simpler sub-tasks, so each step is handled with a focused prompt and yields more coherent final outputs.

By contrast, output safety filters (C) address appropriateness rather than quality or coherence, and random sampling (E) increases diversity but can reduce coherence, so neither is a quality-improvement technique.

Exam trap

Google often tests the distinction between techniques that improve output quality (e.g., self-consistency, prompt chaining, in-context learning) versus safety or diversity mechanisms, leading candidates to mistakenly select output filters or random sampling as quality-enhancing methods.

464
MCQhard

A retail company is building a generative AI chatbot to assist customers with product recommendations and order tracking. The chatbot uses Vertex AI with Gemini 1.5 Pro, and the development team has implemented a Retrieval-Augmented Generation (RAG) pipeline using Vertex AI Search for grounding. The pipeline uses a vector store containing product descriptions and order history. During testing, the team observes that the chatbot sometimes provides incorrect order statuses—for example, claiming an order is 'shipped' when it is actually 'pending'. The team suspects the issue is related to how context is retrieved and used. The RAG pipeline currently retrieves the top 5 chunks based on cosine similarity from the vector store, and passes them as context to the model. The team is considering several changes to improve factual accuracy. Which single action would most effectively reduce hallucinations in this scenario?

A.Switch from Vertex AI Search to a different vector database like Pinecone.
B.Reduce the model temperature to 0.0 to make outputs more deterministic.
C.Increase the similarity score threshold for retrieval to 0.85 to filter out less relevant chunks.
D.Increase the top-K retrieval value to 10 to provide more context to the model.
AnswerC

Increasing the similarity threshold to 0.85 filters out low-relevance chunks (e.g., order statuses from other customers), ensuring the model receives only highly relevant context, which directly reduces hallucinated order statuses in this RAG pipeline.

Why this answer

Increasing the similarity score threshold to 0.85 ensures that only highly relevant chunks are passed to the Gemini 1.5 Pro model, directly reducing the risk of the model generating responses based on irrelevant or low-confidence context. In a RAG pipeline using Vertex AI Search, low-similarity chunks can contain order statuses from different customers or products, leading to hallucinations like incorrect order statuses. Filtering out these less relevant chunks improves the factual grounding of the model's output.

Exam trap

Google Cloud often tests the misconception that simply adding more context (higher top-K) or making the model more deterministic (lower temperature) will fix hallucinations, when the real issue is the relevance and quality of the retrieved context in a RAG pipeline.

How to eliminate wrong answers

Option A is wrong because switching to a different vector database like Pinecone does not address the core issue of retrieval relevance; the problem lies in the similarity threshold and chunk selection, not the database technology. Option B is wrong because reducing temperature to 0.0 makes the model more deterministic but does not fix the underlying issue of irrelevant or incorrect context being retrieved; the model will still confidently generate incorrect answers based on poor context. Option D is wrong because increasing top-K to 10 would retrieve more chunks, potentially including even more low-relevance or noisy context, which could worsen hallucinations rather than improve factual accuracy.

465
MCQeasy

Refer to the exhibit. What is the most likely cause of this error?

A.The user does not have the required IAM role
B.The model is too large
C.The network is down
D.The project ID is incorrect
AnswerA

The error stems from an authorisation failure: the caller's identity lacks the IAM role granting the required permission on the resource. Without that binding, the API rejects the request regardless of valid credentials, so granting the missing role resolves it.

Why this answer

The error shown in the exhibit is an HTTP 403 Forbidden response, which indicates that the server understood the request but refuses to authorize it. In Google Cloud, this is most commonly caused by the user's identity lacking the necessary IAM role or permission to call the specific API or access the resource. Even if the project ID is correct and the network is functional, a missing IAM role (e.g., `aiplatform.user` or `roles/aiplatform.user`) will result in this exact error.

Exam trap

Google Cloud often tests the distinction between authentication (who you are) and authorization (what you can do), and the trap here is that candidates confuse a 403 Forbidden with a 404 Not Found or a network error, leading them to pick 'The project ID is incorrect' or 'The network is down' instead of recognizing the IAM permission failure.

How to eliminate wrong answers

Option B is wrong because model size does not cause an HTTP 403 error; a model that is too large would typically result in a 413 Payload Too Large or a resource-exhausted error, not an authorization failure. Option C is wrong because a network outage would produce a connectivity error (e.g., timeout, DNS resolution failure, or HTTP 502/503), not a 403 Forbidden response which requires a successful TCP connection and HTTP request to reach the server. Option D is wrong because an incorrect project ID would cause a 404 Not Found or a 400 Bad Request (e.g., 'Project not found'), not a 403 Forbidden; the 403 specifically indicates the request was received and the project exists, but the caller lacks authorization.

466
MCQeasy

A retail company wants to use a generative AI model to create unique product descriptions for thousands of items. They need the model to produce varied, creative text without requiring them to provide any examples. Which type of model should they use?

A.A convolutional neural network (CNN)
B.A recurrent neural network (RNN)
C.A decision tree classifier
D.A large language model (LLM) such as Gemini
AnswerD

LLMs like Gemini are pre-trained on vast text corpora and can generate varied, creative text from prompts without task-specific examples. They excel at open-ended generation tasks like product descriptions, making them ideal for this scenario. The model's broad training enables it to produce unique outputs for each product, aligning with the requirement for creativity and variety.

Why this answer

Large language models like Gemini are pre-trained on diverse text and can generate creative, varied outputs from prompts alone. They do not require task-specific examples, making them perfect for producing unique product descriptions at scale. Other model types are designed for different tasks and cannot fulfill the creative text generation requirement.

Exam trap

The trap here is assuming that any neural network can generate text, when only generative models like LLMs are designed for open-ended creation.

467
MCQmedium

A developer is using Vertex AI Gemini API for a chatbot. The chatbot sometimes outputs harmful content. What is the best first step to mitigate this?

A.Fine-tune the model on curated safe data
B.Add a human-in-the-loop review
C.Use safety filters and safety settings in the API request
D.Switch to a smaller model
AnswerC

Safety filters and safety settings are applied per API request, letting the developer block harmful categories before responses reach users. This is the fastest mitigation, requiring no retraining or architectural change to the Gemini chatbot.

Why this answer

The Vertex AI Gemini API provides built-in safety filters and configurable safety settings (e.g., `safety_settings` parameter with categories like `HARM_CATEGORY_HARASSMENT` and thresholds like `BLOCK_ONLY_HIGH`) that allow developers to block harmful outputs at inference time without retraining. This is the fastest and most direct first step to mitigate harmful content, as it requires no additional infrastructure or model modification.

Exam trap

Google Cloud often tests the misconception that the first step to mitigate harmful content is to fine-tune the model, when in reality the immediate, low-cost, and recommended first step is to leverage the API's built-in safety filters and settings.

How to eliminate wrong answers

Option A is wrong because fine-tuning on curated safe data is a resource-intensive, secondary step that does not address immediate harmful outputs during inference and may not cover all edge cases of harmful content. Option B is wrong because adding a human-in-the-loop review introduces latency and cost, and is a reactive measure rather than a proactive first step to block harmful content at the API level. Option D is wrong because switching to a smaller model does not inherently reduce harmful outputs; smaller models can still generate harmful content and may have reduced capabilities for safe response generation.

468
MCQeasy

A support team uses a generative AI assistant to answer customer questions. Agents report that the answers are often too long and include unnecessary background. The team wants the responses to be concise and directly address the question. Which prompt adjustment is most appropriate?

A.Add a system instruction that the model should always provide a detailed explanation for every answer.
B.Lower the model's temperature to zero so answers become deterministic.
C.Tell the model to answer in two sentences or fewer and to skip background unless the customer asks for it.
D.Increase the top-p value so the model considers more possible words.
AnswerC

Adding an explicit length and scope constraint directly addresses the verbosity problem. The model receives a clear instruction about how many sentences to use and what to omit. This is a simple prompt-engineering fix that does not require retraining or infrastructure changes. It also preserves the assistant's ability to answer follow-up questions when background is requested.

Why this answer

The most appropriate adjustment is an explicit prompt constraint that limits answer length and omits background unless requested. Verbosity is a content and instruction-following issue, so a clear directive in the prompt is the right control. Sampling parameters like temperature and top-p do not directly enforce conciseness, and asking for more detail would worsen the problem.

Exam trap

The trap here is confusing randomness controls such as temperature and top-p with instruction-following controls such as explicit length and scope constraints.

469
MCQhard

An organization wants to use a generative model to automatically generate legal contracts. The model must produce clauses that are not only grammatically correct but also legally enforceable and consistent with current jurisdiction laws. Which combination of techniques best ensures legal compliance?

A.Fine-tune a small model exclusively on legal contracts from a single jurisdiction and use it for generation.
B.Implement retrieval-augmented generation (RAG) with a vector database of all relevant laws.
C.Fine-tune a model on a diverse set of enforceable contracts and incorporate an external compliance verifier that uses rule-based checks.
D.Use a large instruction-tuned model with carefully engineered prompts describing jurisdiction details.
AnswerC

Fine-tuning on enforceable contracts aligns the model’s outputs with jurisdiction-specific clause patterns, while the external rule-based verifier enforces statutory constraints the model cannot guarantee. This satisfies the stem’s requirement for legally enforceable, jurisdiction-consistent clauses by combining learned drafting conventions with deterministic compliance checks, rather than relying on the model’s probabilistic recall of current law.

Why this answer

Fine-tuning on a diverse corpus of enforceable contracts teaches the model patterns of valid legal language across contexts, while an external rule-based compliance verifier checks generated clauses against jurisdiction-specific statutes and regulations. This hybrid approach combines generative fluency with deterministic legal validation, which is necessary because LLMs alone cannot guarantee current legal compliance. The verifier provides an auditable, updatable layer that can be revised as laws change.

Exam trap

Generative AI Leader often tests the belief that RAG or prompt engineering alone guarantees compliance, when deterministic external validation is required for regulated domains.

How to eliminate wrong answers

Option A is wrong because training on a single jurisdiction's contracts limits generalization and still cannot guarantee the model internalizes every current statute or amendment. Option B is wrong because RAG retrieves relevant laws but does not enforce them; the model may still generate non-compliant clauses despite having the correct text in context. Option D is wrong because prompt engineering alone cannot ensure legal enforceability or up-to-date jurisdiction compliance, as the model's parametric knowledge may be stale or incomplete.

470
MCQmedium

A researcher wants to detect whether text was generated by an AI model to identify potential misinformation. Which Google technology is specifically designed for this purpose?

A.Model Cards
B.PAIR Explorables
C.SynthID
D.Datasheets for Datasets
AnswerC

SynthID embeds imperceptible watermarks directly into AI-generated content, enabling later detection of synthetic origin. This satisfies the misinformation-detection goal because the signal survives common edits and transformations, unlike classifiers or metadata, which can be stripped or produce probabilistic guesses.

Why this answer

SynthID is a Google DeepMind technology specifically designed to watermark and detect AI-generated text, images, audio, and video. It embeds an imperceptible digital watermark directly into the output of generative models, allowing for reliable detection even after modifications like cropping or compression. This makes it the correct tool for identifying potential misinformation from AI-generated text.

Exam trap

The trap here is that candidates confuse documentation tools (Model Cards, Datasheets) or educational resources (PAIR Explorables) with active detection technologies, overlooking SynthID as the only option purpose-built for watermarking and identifying AI-generated content.

How to eliminate wrong answers

Option A is wrong because Model Cards are standardized documentation templates that disclose a model's intended use, performance, and limitations; they do not provide any detection or watermarking capability for AI-generated content. Option B is wrong because PAIR Explorables are interactive visualizations and tutorials created by Google's People + AI Research (PAIR) team to explain AI concepts, not a detection tool for AI-generated text. Option D is wrong because Datasheets for Datasets are structured documentation for datasets (e.g., provenance, bias, collection methods) and have no functionality to detect whether text was generated by an AI model.

471
MCQmedium

A company wants to use GenAI to automate customer support. They have a large knowledge base. Which approach maximizes ROI in the first 6 months?

A.Deploy a general-purpose chatbot without customization
B.Use a pre-built conversational AI platform with Retrieval-Augmented Generation (RAG)
C.Build a custom LLM from scratch using their data
D.Fine-tune a foundation model on historical support tickets
AnswerB

A pre-built conversational platform combined with Retrieval-Augmented Generation grounds responses in the existing knowledge base without training a custom model, cutting build time and cost. This satisfies the stem's six-month ROI constraint by delivering working automation quickly.

Why this answer

Maximizes ROI in the first 6 months because it leverages a pre-built conversational AI platform integrated with Retrieval-Augmented Generation (RAG). RAG allows the model to dynamically retrieve relevant information from the existing knowledge base at inference time, providing accurate, context-aware responses without the need for costly retraining or custom model development. This approach balances rapid deployment, low upfront investment, and high accuracy, making it the most cost-effective solution for automating customer support quickly.

Exam trap

Google Cloud often tests the misconception that fine-tuning is always the best way to incorporate proprietary data, but the trap here is that fine-tuning does not provide real-time access to a dynamic knowledge base and is far more resource-intensive than RAG, which is the optimal strategy for rapid, cost-effective deployment in customer support scenarios.

How to eliminate wrong answers

Option A is wrong because deploying a general-purpose chatbot without customization would rely solely on the model's pre-trained knowledge, which lacks access to the company's specific knowledge base, leading to frequent hallucinations and incorrect answers that degrade customer trust and require extensive human oversight. Option C is wrong because building a custom LLM from scratch using their data is prohibitively expensive (often millions of dollars) and time-consuming (typically 12+ months), far exceeding the 6-month ROI window and requiring massive computational resources and specialized ML teams. Option D is wrong because fine-tuning a foundation model on historical support tickets alone does not incorporate the live knowledge base; it only adapts the model to past conversation patterns, which may become stale or miss updated information, and still requires significant compute and data preparation costs without the real-time retrieval capability that RAG provides.

472
Multi-Selecthard

A research lab is planning to train a massive protein folding model similar to AlphaFold. They want to use Google Cloud infrastructure and tools. Which THREE components are most relevant?

Select 3 answers
A.Cloud TPU pods
B.Cloud Vision API
C.Vertex AI Pipeline
D.Google AI Studio
E.Google DeepMind collaboration
AnswersA, C, E

Cloud TPU pods supply the tightly coupled, high-bandwidth interconnects and matrix units needed to train massive protein folding models, satisfying the lab's requirement for Google Cloud infrastructure capable of scaling AlphaFold-class workloads across thousands of chips. Standard GPU clusters lack the same pod-level interconnect topology optimised for this scale.

Why this answer

Cloud TPU pods (A) are correct because training massive protein folding models like AlphaFold requires enormous matrix/tensor computation, and TPU pods provide the tightly-coupled, high-bandwidth interconnects and thousands of TPU chips needed for large-scale distributed training. Vertex AI Pipeline (C) is correct because it orchestrates the multi-step ML workflow — data preprocessing, training, evaluation, and deployment — as a reproducible, managed pipeline on Google Cloud, which is essential for a complex research training process. Google DeepMind collaboration (E) is correct because AlphaFold itself was developed by DeepMind, and partnering with DeepMind gives the lab access to the domain expertise, model architectures, and research guidance specific to protein folding.

Cloud Vision API (B) is not relevant because it is a pre-trained service for image classification, OCR, and object detection, not protein structure prediction. Google AI Studio (D) is not relevant because it is a lightweight developer tool for prototyping prompts with Gemini models, not a platform for large-scale scientific model training.

Exam trap

The trap here is that candidates may confuse Google's pre-built AI services (like Vision API or AI Studio) with the specialized infrastructure needed for training custom large-scale models, overlooking that TPU pods are the core compute resource for such workloads.

473
MCQeasy

A product manager wants to communicate the limitations of a new generative AI feature to stakeholders. According to Google's People + AI Guidebook, what is the BEST approach?

A.Highlight only the successful use cases to build excitement
B.Provide clear examples of the AI's capabilities and failure modes
C.Promise that future versions will overcome all limitations
D.Share a technical paper detailing the model architecture
AnswerB

The People + AI Guidebook recommends setting accurate expectations by showing both what the feature does well and where it fails. Concrete capability and failure-mode examples let stakeholders judge real-world reliability, satisfying the need to communicate limitations honestly.

Why this answer

The People + AI Guidebook emphasizes setting appropriate expectations by clearly communicating what the AI can and cannot do.

474
MCQmedium

A company wants to build a customer service chatbot that answers questions about their internal policy documents. The documents are updated monthly, and the team cannot afford to retrain a model each time. Which approach is MOST appropriate?

A.Fine-tune a base LLM on the policy documents monthly
B.Use a larger foundation model with a longer context window and paste all documents into each prompt
C.Train a custom model from scratch on the policy documents each month
D.Use Retrieval-Augmented Generation (RAG) with the policy documents indexed in a vector store
AnswerD

RAG indexes the policy documents in a vector store and retrieves relevant chunks per query, so monthly updates require only re-indexing rather than retraining. This satisfies the constraint that the team cannot afford model retraining each time documents change.

Why this answer

Retrieval-Augmented Generation (RAG) is the most appropriate approach because it allows the chatbot to answer questions by retrieving relevant chunks from the policy documents stored in a vector store, without requiring model retraining. When documents are updated monthly, only the vector index needs to be refreshed, which is far more cost-effective and faster than fine-tuning or retraining a model. This approach leverages the base LLM's language understanding while grounding responses in the latest policy content.

Exam trap

Google often tests the misconception that fine-tuning or retraining is necessary for domain-specific knowledge, when in fact RAG provides a dynamic, cost-effective solution that avoids retraining and handles frequently updated documents.

How to eliminate wrong answers

Option A is wrong because fine-tuning a base LLM monthly on policy documents is expensive, time-consuming, and introduces risk of catastrophic forgetting, where the model may lose general language capabilities. Option B is wrong because pasting all documents into each prompt is impractical due to context window limits (even large models like GPT-4 have ~128K tokens, which is insufficient for extensive policy libraries) and leads to high inference costs and latency. Option C is wrong because training a custom model from scratch each month is prohibitively expensive and computationally intensive, requiring massive datasets and GPU resources, and is unnecessary when RAG can achieve the same goal with far less overhead.

475
MCQhard

A healthcare company is deploying a GenAI-powered report generation system that processes patient summaries. They need the output to be structured JSON for downstream ingestion. The team is using Vertex AI Studio prompt design. Which approach best ensures consistent structured output?

A.Ask the model to output plain text and parse it with a regex after generation
B.Define the JSON schema in the system instruction and provide one few-shot example of the expected JSON
C.Use a lower temperature (0.0) and hope the model outputs valid JSON
D.Generate output in Markdown and convert to JSON using a secondary script
AnswerB

Embedding the JSON schema in the system instruction constrains the model's output format across every request, while the few-shot example demonstrates the exact structure and key naming expected. Together they satisfy the stem's requirement for consistent, machine-parseable JSON suitable for downstream ingestion.

Why this answer

Defining the JSON schema in the system instruction and providing a few-shot example directly constrains the model's output format at inference time, leveraging Vertex AI's instruction-following capabilities. This approach ensures the model generates valid JSON consistently without post-processing, as the schema acts as a structural template that the model learns to replicate.

Exam trap

The trap here is that candidates assume lowering temperature to 0.0 guarantees deterministic and correct structured output, but without explicit schema guidance, Google’s Vertex AI model can still produce syntactically invalid JSON due to inherent token-level variability.

How to eliminate wrong answers

Option A is wrong because relying on regex parsing of plain text is brittle and error-prone; the model may produce unstructured or malformed text that regex cannot reliably parse, especially with variable patient summary content. Option C is wrong because lowering temperature to 0.0 reduces randomness but does not guarantee valid JSON syntax; the model can still output malformed JSON (e.g., missing commas, unescaped quotes) without explicit schema guidance. Option D is wrong because generating Markdown and converting to JSON adds an unnecessary secondary processing step, introducing potential conversion errors and latency, and does not leverage the model's ability to output structured data directly.

476
MCQhard

A data scientist is using Vertex AI to generate product descriptions from a list of features. They notice that the model sometimes omits key features or invents details not present in the input. They want to reduce hallucinations and ensure all provided features are included. Which technique should they apply?

A.Enable the model's grounding feature with a public dataset.
B.Use a few-shot prompting approach with examples that include all features.
C.Set the top-p parameter to 0.1 to narrow the token selection.
D.Increase the temperature to encourage more creative outputs.
AnswerB

Few-shot prompting provides the model with examples of the desired input-output format, showing how to incorporate all features without adding extraneous details. This guides the model to follow the pattern, reducing omissions and hallucinations. It is an effective prompt engineering technique for structured generation tasks.

Why this answer

Few-shot prompting is the most direct way to guide the model to include all features and avoid hallucinations. By showing examples where every feature is incorporated accurately, the model learns the expected pattern. Other parameters like temperature or top-p affect randomness but do not provide the necessary instruction on completeness and fidelity.

Exam trap

The trap here is thinking that lowering randomness parameters like top-p will solve hallucinations, when the real issue is lack of guidance on the task format.

477
Multi-Selecthard

A company is evaluating Google Cloud's AI portfolio versus competitors. They want to leverage Gemini's unique capabilities. Which THREE differentiators should they highlight when comparing to AWS Bedrock and Azure OpenAI? (Choose 3)

Select 3 answers
A.Vertex AI as a unified ML platform
B.Custom TPU hardware for training large models
C.Grounding with Google Search for real-time, verifiable responses
D.Integration with Google Workspace (e.g., Gmail, Docs)
E.Native multimodal understanding across text, images, video, and audio
AnswersC, D, E

Google Cloud offers native grounding with Google Search, reducing hallucinations and providing citations; AWS/Azure have similar features but not with Google Search.

Why this answer

Gemini's grounding with Google Search allows it to access and cite real-time, verifiable information from the web, reducing hallucinations and improving factual accuracy. This is a unique differentiator as AWS Bedrock and Azure OpenAI do not natively integrate a live search engine for grounding responses, requiring custom RAG implementations instead.

Exam trap

This exam often tests the distinction between platform-level features (like Vertex AI or TPUs) and model-level differentiators (like grounding or multimodal understanding), causing candidates to select options that are true for Google Cloud but not unique to Gemini.

478
Multi-Selecteasy

According to Google's AI Principles, which TWO of the following are core commitments? (Select 2)

Select 2 answers
A.Maximize shareholder value
B.Always use the largest model available
C.Achieve state-of-the-art performance on all benchmarks
D.Be built and tested for safety
E.Be socially beneficial
AnswersD, E

Being built and tested for safety is one of Google's published AI Principles commitments, sitting alongside avoiding bias and upholding privacy. It satisfies the stem's requirement for a core commitment by mandating rigorous testing throughout development, ensuring systems behave safely before deployment and throughout their operational life.

Why this answer

Option D, 'Be built and tested for safety,' is correct because Google's AI Principles explicitly commit to avoiding unfair bias and ensuring AI systems are built and tested for safety, including rigorous testing for harmful outcomes and unintended behavior. Option E, 'Be socially beneficial,' is correct because the first of Google's AI Principles states that AI should be developed to benefit society broadly, considering a wide range of social and economic factors and minimizing negative impacts. These two commitments appear directly among Google's published AI Principles, which also include avoiding creating or reinforcing unfair bias, being accountable to people, incorporating privacy design principles, upholding high standards of scientific excellence, and being made available for uses that accord with these principles.

Option A, 'Maximize shareholder value,' is not part of the AI Principles, as they focus on ethical and societal outcomes rather than financial returns. Option B, 'Always use the largest model available,' is not a commitment, since the principles emphasize appropriate and beneficial use of AI rather than model size. Option C, 'Achieve state-of-the-art performance on all benchmarks,' is not a stated commitment, as the principles prioritize safety, benefit, and responsibility over benchmark leadership.

Exam trap

Google's AI Principles prioritize ethical commitments over technical excellence or business value. Candidates often mistakenly select options like maximizing performance or using the largest model, which are not part of Google's core AI Principles.

479
MCQeasy

A data scientist is using the Gemini API to generate product descriptions for an e-commerce site. The descriptions are often too verbose and include speculative claims that are not in the product specifications. The scientist wants to reduce hallucinations and control the length of the output without retraining the model. What should they do?

A.Increase the max output token count to 2048 and decrease temperature to 0.1.
B.Refine the prompt to be concise and include instructions to stick to facts and limit output to 50 words.
C.Add three few-shot examples of short, factual descriptions.
D.Set temperature to 0.0 and top_k to 1.
AnswerB

Refining the prompt directly constrains generation through the model's instruction-following behaviour, requiring no retraining. Explicit instructions to limit output to 50 words satisfy the length constraint, while directing the model to stick to supplied product specifications reduces speculative claims at inference time. This addresses both the verbosity and hallucination issues within the Gemini API workflow.

Why this answer

Refining the prompt to be concise and include explicit instructions to stick to facts and limit output to 50 words directly addresses both issues without retraining. Prompt engineering is the most effective technique for controlling output length and reducing hallucinations in the Gemini API, as it guides the model's behavior through natural language constraints rather than altering generation parameters. Note that few-shot examples (Option C) are also a form of prompt engineering and could help, but the question asks for the most direct approach: explicit natural-language instructions in the prompt are the simplest and most reliable way to enforce a strict word limit and factual grounding, whereas few-shot examples alone do not guarantee a specific output length.

Exam trap

This exam often tests the misconception that adjusting generation parameters like temperature or top_k is the primary way to control factual accuracy and length, when in fact prompt engineering is the most direct and effective method for these specific requirements without retraining. A related trap is assuming that few-shot examples are always superior to explicit instructions; for strict length control and factual grounding, clear natural-language constraints in the prompt are the most direct solution.

How to eliminate wrong answers

Option A is wrong because increasing the max output token count to 2048 would make the descriptions even more verbose, which is the opposite of what the data scientist wants; decreasing temperature to 0.1 reduces randomness but does not enforce factual adherence or length limits. Option C is wrong because adding three few-shot examples can improve style and structure but does not reliably prevent speculative claims or enforce a strict word count, especially if the examples are not perfectly aligned with the desired constraints. Option D is wrong because setting temperature to 0.0 and top_k to 1 makes the output deterministic and repetitive, which reduces creativity but does not inherently eliminate hallucinations or control verbosity; the model may still generate speculative content based on its training data.

480
MCQhard

A regional insurance company wants to launch a generative AI claims-triage assistant. The CISO requires that no claims data leave the company's existing Google Cloud project boundary and that every model call be attributable to a named employee for audit. Which combination of Google Cloud controls should the architecture team prioritize?

A.Customer-managed encryption keys in Cloud KMS plus a service account shared by the entire claims department.
B.Cloud Armor policies on the public load balancer plus Security Command Center premium findings.
C.VPC Service Controls around the project plus Cloud Audit Logs capturing the authenticated principal on each Vertex AI request.
D.Cloud CDN in front of the assistant plus Identity-Aware Proxy on the internal admin console.
AnswerC

VPC Service Controls create a service perimeter that blocks data exfiltration from the project even if credentials are misused, directly satisfying the boundary requirement. Cloud Audit Logs record the identity of the caller for Vertex AI API activity, giving auditors per-employee attribution. Together they address both the data-residency concern and the traceability requirement without extra infrastructure.

Why this answer

A service perimeter built with VPC Service Controls prevents claims data from being moved outside the approved project even by credentialed users, which is the strongest fit for the boundary mandate. Cloud Audit Logs then capture the authenticated principal behind every Vertex AI call, producing the employee-level attribution auditors demand. Encryption and edge protections are valuable but do not replace these two controls.

Exam trap

The trap here is treating encryption keys or edge security as sufficient for data-boundary compliance, when only a service perimeter actually blocks exfiltration by authorized identities.

481
Multi-Selecteasy

Which TWO of the following are valid Google Cloud AI APIs for natural language processing?

Select 2 answers
A.Speech-to-Text API
B.Translation API
C.Vision AI
D.Natural Language API
E.Document AI
AnswersB, D

Translation API performs natural language processing by detecting source language and converting text between languages, a core NLP task. It is a genuine Google Cloud AI API, so it satisfies the stem's requirement for valid NLP APIs.

Why this answer

The Translation API and Natural Language API are both valid Google Cloud AI APIs for natural language processing (NLP). The Translation API converts text between languages, which is a core NLP task, while the Natural Language API provides entity recognition, sentiment analysis, and syntax analysis, directly processing human language.

Exam trap

A common mistake is to include Speech-to-Text or Vision AI as NLP APIs when they actually belong to different AI domains (speech and vision, respectively). The correct NLP APIs are the Translation API and Natural Language API.

482
MCQmedium

A company fine-tunes a text model on internal HR policies. After deployment, the model sometimes outputs sensitive employee information. What is the most likely cause?

A.The fine-tuning dataset contained personally identifiable information that was not removed.
B.The model was not trained with reinforcement learning from human feedback (RLHF).
C.The model has insufficient parameters to generalize properly.
D.The prompt engineering was too verbose and included misleading instructions.
AnswerA

Fine-tuning embeds the training corpus into the model's weights, so any personally identifiable information left in the HR policy dataset can be reproduced verbatim at inference. The stem's constraint is sensitive employee data appearing in outputs; removing PII before fine-tuning prevents this memorisation, since no filtering layer exists afterwards.

Why this answer

The most likely cause is that the fine-tuning dataset contained personally identifiable information (PII) that was not properly scrubbed. During fine-tuning, the model learns patterns and memorizes specific sequences from the training data. If the dataset includes sensitive employee records, the model can reproduce that information verbatim when prompted, leading to data leakage.

This is a well-known risk in fine-tuning, as models can overfit to rare or unique examples in the training set.

Exam trap

Google Cloud often tests the misconception that RLHF or prompt engineering can fix data leakage issues, but the trap here is that the root cause is always the training data itself—no amount of post-hoc alignment or prompt tweaking can prevent the model from reproducing memorized sensitive content.

How to eliminate wrong answers

Option B is wrong because RLHF is a technique used to align model outputs with human preferences, not to prevent memorization of training data; it does not address the root cause of data leakage from the fine-tuning dataset. Option C is wrong because insufficient parameters would typically cause underfitting or poor generalization, not the exact reproduction of sensitive information; memorization is more likely with larger models that have higher capacity to store training examples. Option D is wrong because verbose or misleading prompt engineering might degrade output quality but cannot cause the model to output specific employee data that was not present in its training or fine-tuning data; the model can only generate information it has learned.

483
MCQeasy

A marketing agency wants to generate personalized product descriptions at scale but has no machine learning engineers on staff. They need a managed Google Cloud option that provides access to foundation models through an API with minimal infrastructure work. Which offering should they choose?

A.Vertex AI Model Garden with Gemini models accessed through the Vertex AI API.
B.Build a custom model using BigQuery ML with only SQL and no external model access.
C.Use Cloud Natural Language API sentiment analysis to generate product descriptions.
D.Deploy an open-source LLM on a self-managed Compute Engine VM with GPUs.
AnswerA

Vertex AI Model Garden provides managed access to Gemini and other foundation models through APIs, so the agency can generate content without building or hosting infrastructure. It supports enterprise controls such as IAM, quotas, and logging. For a team without ML engineers, this removes operational burden while still enabling customization through prompts and optional tuning.

Why this answer

A managed foundation model catalog accessed through an API lets teams without ML specialists generate text at scale while Google Cloud handles serving, scaling, and security. Self-managed GPU deployments demand engineering skills, BigQuery ML targets structured predictions, and the Natural Language API analyzes rather than generates text. The managed API path best matches the agency's staffing and infrastructure constraints.

Exam trap

The trap here is equating any Google Cloud AI service with generative capability, when several of these options analyze or predict rather than generate text.

484
MCQmedium

A retail company is deploying a generative AI chatbot on Vertex AI to provide product recommendations. The chatbot uses a base foundation model with no fine-tuning. Users report that the chatbot sometimes gives offensive or insensitive responses. The team must quickly implement safety controls without modifying the model. They also want to reduce irrelevant off-topic answers. Which combination of techniques should they apply?

A.Fine-tune the model on a curated dataset of safe retail conversations.
B.Set temperature to 0.0 and top_p to 0.1.
C.Enable Vertex AI Safety Filters and craft system instructions defining appropriate behavior.
D.Provide 50 few-shot examples of safe interactions.
AnswerC

Safety filters block harmful content at the API layer without retraining the model, while system instructions constrain tone and scope, reducing off-topic replies. Together they deliver fast, model-agnostic guardrails meeting both safety and relevance requirements.

Why this answer

Vertex AI Safety Filters provide out-of-the-box content moderation without modifying the model, and crafting system instructions (system-level prompts) can constrain the chatbot's behavior to stay on-topic and avoid offensive responses. This combination addresses both safety and relevance without requiring fine-tuning or altering model parameters.

Exam trap

Google often tests the distinction between parameter tuning (temperature/top_p) and safety mechanisms—candidates mistakenly think lowering randomness prevents offensive outputs, but safety requires explicit filtering or instruction-based guardrails, not just reduced creativity.

How to eliminate wrong answers

Option A is wrong because fine-tuning requires modifying the model, which contradicts the requirement to not modify the model; it also takes time and resources, not a 'quick' fix. Option B is wrong because setting temperature to 0.0 and top_p to 0.1 reduces randomness and diversity but does not prevent offensive or insensitive responses—it only makes outputs more deterministic, not safer. Option D is wrong because providing 50 few-shot examples of safe interactions can guide the model but does not guarantee safety filtering; it also requires careful curation and may not scale, and the model can still generate off-topic or offensive responses outside the examples.

485
MCQmedium

Refer to the exhibit. A team attempted to start a model tuning job but received the error 'Quota limit exceeded for tuning jobs in region us-central1'. What is the most appropriate action?

A.Request a quota increase for tuning jobs in us-central1
B.Change the region to us-west1 and retry
C.Reduce the size of the training data
D.Use a different base model
AnswerA

The error names a regional quota ceiling, not a configuration fault, so the tuning job cannot start until capacity is granted. Requesting an increase for us-central1 directly lifts that constraint, letting the job run in the required region without redesigning the pipeline or moving data.

Why this answer

The error 'Quota limit exceeded for tuning jobs in region us-central1' indicates that the project has reached its predefined resource quota for model tuning operations in that specific region. The most appropriate action is to request a quota increase from Google Cloud, as this directly resolves the capacity limitation without altering the job's configuration or data. Quotas are per-region limits enforced by the AI Platform to ensure fair resource allocation, and increasing the quota is the standard procedure when legitimate tuning needs exceed the default allowance.

Exam trap

A common pitfall is assuming quota errors can be fixed by changing job parameters (region, data size, model) instead of recognizing that quotas are administrative limits requiring a formal increase request through Google Cloud.

How to eliminate wrong answers

Option B is wrong because changing the region to us-west1 does not address the root cause of the quota limit; the tuning job may still fail if the quota in us-west1 is also insufficient or if the model or data has regional dependencies. Option C is wrong because reducing the size of the training data does not affect the quota limit for tuning jobs; quotas are based on the number of concurrent or total tuning jobs, not on data size. Option D is wrong because using a different base model does not change the quota consumption for tuning jobs; the quota applies to the tuning operation itself, regardless of which base model is selected.

486
Multi-Selectmedium

A data scientist is documenting a new dataset for a generative AI project. According to the Responsible AI toolkit, which TWO elements should they include in a Datasheet for Datasets?

Select 2 answers
A.The model architecture used to collect the data
B.The demographic composition of the data subjects
C.The hyperparameters of the model that will process the data
D.The intended use cases and limitations
E.The cost of acquiring the dataset
AnswersB, D

Datasheets for Datasets require documenting dataset composition, including demographic characteristics of data subjects, to expose potential representation gaps and bias. This transparency element satisfies the toolkit's documentation requirement, enabling downstream assessment of whether the dataset fairly represents the populations the generative AI system will serve.

Why this answer

Option B is correct because a Datasheet for Datasets, as promoted by the Responsible AI toolkit, should document the demographic composition of data subjects to surface potential representation gaps, bias risks, and fairness considerations for the generative AI project. Option D is correct because datasheets must state the intended use cases and limitations so downstream consumers understand appropriate applications and avoid misuse or overclaiming model capabilities. Options A and C are incorrect because model architecture and hyperparameters describe model training or inference configuration, not the dataset documentation itself.

Option E is incorrect because acquisition cost is a procurement or business metric, not a required element of a Datasheet for Datasets under the Responsible AI toolkit.

Exam trap

The Generative AI Leader exam often tests the distinction between dataset documentation (Datasheet for Datasets) and model documentation (Model Cards), so candidates mistakenly include model-specific details like architecture or hyperparameters instead of dataset-focused elements.

487
MCQmedium

A healthcare organization wants to use generative AI to draft patient education materials. They are concerned about the model generating incorrect medical information. Which combination of Google Cloud services should they use to ground the model's responses in trusted medical literature?

A.Use Retrieval-Augmented Generation (RAG) with Vertex AI Search and Gemini
B.Use Imagen to generate visual aids and combine with Gemini for text, then manually review
C.Fine-tune Gemini on trusted medical literature and deploy with Vertex AI Endpoints
D.Use only Gemini Pro with careful prompt engineering and system instructions
AnswerA

RAG retrieves passages from an indexed trusted corpus via Vertex AI Search, then Gemini conditions its answer on those passages, so responses cite medical literature rather than relying on parametric memory. This grounds output and reduces fabricated medical claims, satisfying the accuracy concern.

Why this answer

Retrieval-Augmented Generation (RAG) with Vertex AI Search allows the model to retrieve and cite information from a curated corpus of trusted medical literature before generating responses. This grounds the output in verified sources, reducing the risk of hallucination or incorrect medical advice, while Gemini provides the generative capabilities.

Exam trap

A common misconception tested is that fine-tuning or careful prompting alone can ensure factual accuracy, when in reality RAG provides a more reliable grounding mechanism by retrieving and citing external, up-to-date sources.

How to eliminate wrong answers

Option B is wrong because Imagen generates visual aids but does not address the core requirement of grounding text responses in trusted medical literature; manual review is not a scalable or reliable grounding mechanism. Option C is wrong because fine-tuning Gemini on trusted literature embeds knowledge into the model weights, which can still lead to outdated or hallucinated information if the training data is static, and it does not provide real-time retrieval or citation of specific sources. Option D is wrong because using only Gemini Pro with prompt engineering and system instructions does not guarantee factual accuracy; without retrieval from a trusted corpus, the model may still generate plausible but incorrect medical information.

488
MCQmedium

A company is evaluating whether to build a custom fine-tuned model for code generation or use a pre-built API like Gemini API. The code generation needs to follow the company's internal coding standards. Which consideration is MOST important in deciding to build vs buy?

A.Pre-built APIs may not generate code that adheres to internal standards without extensive prompting
B.Pre-built APIs are always cheaper than fine-tuned models
C.Fine-tuned models have lower latency than pre-built APIs
D.Data privacy is easier to achieve with pre-built APIs
AnswerA

Pre-built APIs generate code from general training data, so they cannot inherently encode your proprietary coding standards; alignment depends entirely on prompt engineering, which is brittle and inconsistent. That directly satisfies the stem's constraint that generated code must follow internal standards, making fine-tuning the more reliable route.

Why this answer

The core build-vs-buy decision hinges on whether a pre-built API can meet the specific requirement of adhering to internal coding standards. General-purpose APIs like Gemini are trained on public code and will not inherently follow a company's proprietary style guides, naming conventions, or architectural patterns without extensive prompt engineering or fine-tuning. This makes option A the most decision-relevant consideration.

Exam trap

Generative AI Leader questions often test the misconception that 'pre-built APIs are always cheaper/faster' — the trap is ignoring that customization requirements (like internal coding standards) can make fine-tuning the only viable path regardless of cost.

How to eliminate wrong answers

Option B is wrong because pre-built APIs are not 'always' cheaper — at scale, API token costs can exceed the cost of hosting a fine-tuned model, and the statement is an absolute that ignores volume, latency, and customization trade-offs. Option C is wrong because fine-tuned models do not inherently have lower latency; in fact, self-hosted fine-tuned models often have higher latency due to infrastructure overhead, while managed APIs are optimized for throughput. Option D is wrong because data privacy is generally easier to control with a self-hosted fine-tuned model (data never leaves the VPC), not with a pre-built API where prompts are sent to a third party.

489
MCQmedium

A global media company wants to add generative AI capabilities to its existing applications, including text summarization, image generation, and code assistance. They plan to use Google Cloud services and need a managed solution that provides access to multiple foundation models through a single API, with enterprise-grade security and scalability. Which Google Cloud offering should they choose?

A.Dialogflow CX
B.Document AI
C.AutoML
D.Vertex AI
AnswerD

Vertex AI is Google Cloud's unified platform for building, deploying, and scaling machine learning models, including generative AI. It provides access to foundation models like Gemini, Imagen, and Codey through a single API, along with tools for tuning, evaluation, and deployment. This matches the company's need for a managed solution supporting multiple modalities with enterprise security.

Why this answer

Vertex AI is the correct choice because it is Google Cloud's comprehensive platform for generative AI, offering access to multiple foundation models through a single API. It supports text, image, and code generation, and includes enterprise features like security, scalability, and integration with other Google Cloud services. The other options are either specialized tools or subsets of Vertex AI that do not fulfill the need for a unified, multi-modal generative AI platform.

Exam trap

The trap here is confusing AutoML as a standalone generative AI service when it is actually a component of Vertex AI focused on custom model training.

490
MCQeasy

A developer wants to generate product descriptions from a list of features using Vertex AI. Which model type is best suited for this task?

A.An embedding model (e.g., textembedding-gecko@001).
B.A chat model (e.g., chat-bison@001).
C.A text generation model (e.g., text-bison@001).
D.A code generation model (e.g., code-bison@001).
AnswerC

A text generation model such as text-bison@001 maps a feature list to fluent prose, satisfying the stem's requirement to produce product descriptions. Unlike classification or embedding models, it outputs free-form natural language, which is precisely the generative capability Vertex AI needs here.

Why this answer

Text-bison@001 is a dedicated text generation model optimized for tasks like summarization, translation, and content creation from structured inputs. It can take a list of features as a prompt and generate coherent, descriptive product descriptions without needing conversational context or code-specific outputs.

Exam trap

The trap here is that candidates may confuse 'text generation' with 'chat' or 'embedding' models, assuming any generative model can handle the task, but Vertex AI separates these by specialization, and the exam tests awareness of which model class is purpose-built for non-conversational, non-code text creation.

How to eliminate wrong answers

Option A is wrong because embedding models like textembedding-gecko@001 are designed to convert text into numerical vectors for similarity search or clustering, not for generating new text. Option B is wrong because chat models like chat-bison@001 are optimized for multi-turn conversational interactions, not for single-turn structured generation tasks like producing descriptions from a feature list. Option D is wrong because code generation models like code-bison@001 are specialized for generating programming code, not natural language product descriptions.

491
MCQeasy

What is the primary purpose of the 'top-p' (nucleus sampling) parameter in text generation?

A.To define a cumulative probability threshold for token selection
B.To set a fixed number of highest probability tokens to consider
C.To limit the maximum number of tokens in the output
D.To control the randomness of the output by scaling logits
AnswerA

Top-p sets a cumulative probability threshold, so the model samples only from the smallest set of tokens whose combined likelihood reaches that value. This directly satisfies the question's focus on nucleus sampling's purpose, dynamically truncating the candidate pool per step rather than fixing a static count as top-k does.

Why this answer

The top-p parameter, also known as nucleus sampling, selects the smallest set of tokens whose cumulative probability exceeds the threshold p (e.g., 0.9). This dynamically adapts the candidate pool based on the model's confidence, allowing more diverse tokens when the distribution is flat and fewer when it is peaked. It directly implements a cumulative probability cutoff, not a fixed count or scaling factor.

Exam trap

In Google Gen AI exams, candidates often confuse top-p (nucleus sampling) with top-k sampling, mistakenly thinking top-p selects a fixed number of tokens rather than a probability-based dynamic set.

How to eliminate wrong answers

Option B is wrong because it describes top-k sampling, which selects a fixed number of highest probability tokens regardless of their cumulative probability. Option C is wrong because it refers to the max_tokens or max_length parameter, which controls output length, not token selection diversity. Option D is wrong because it describes temperature scaling, which adjusts logit probabilities via a softmax divisor before sampling, not a cumulative probability threshold.

492
Multi-Selecthard

A healthcare startup is building a GenAI application that answers patient queries based on medical literature. They need to ensure factual accuracy and compliance with healthcare regulations. Which TWO strategies should they use? (Choose 2)

Select 2 answers
A.Rely on few-shot prompting with example Q&A pairs
B.Fine-tune the model on medical literature
C.Implement a response schema for structured JSON output
D.Use RAG Engine with a curated medical knowledge base
E.Use Grounding with Google Search to verify facts
AnswersD, E

RAG retrieves answers from a controlled set of medical documents, ensuring sources are authoritative and up-to-date.

Why this answer

Grounding with Google Search improves factual accuracy by basing answers on verified search results. A response schema for structured output is not directly about accuracy. RAG with a curated medical knowledge base ensures answers come from trusted sources.

Few-shot prompting alone is insufficient. Fine-tuning on medical data is not selected because two correct options are already chosen.

493
MCQhard

A multinational corporation uses Vertex AI to fine-tune a language model on proprietary customer data. They want to ensure that the fine-tuned model does not inadvertently memorize and regurgitate sensitive customer information. Which approach is most effective?

A.Regularly audit the model by prompting it with known sensitive phrases
B.Apply differential privacy techniques during fine-tuning
C.Use a smaller model to reduce memorization capacity
D.Train the model with a higher learning rate to reduce overfitting
AnswerB

Differential privacy adds calibrated noise during fine-tuning, bounding any single record's influence on learned parameters. This directly prevents the model from memorising and regurgitating individual customer records, satisfying the stem's requirement that sensitive proprietary data cannot be reproduced verbatim by the fine-tuned model.

Why this answer

Differential privacy during training limits memorization by adding noise, making it much harder for the model to leak sensitive data from the training set.

494
MCQhard

A company is deploying a large language model on Vertex AI for real-time inference. They observe high latency and want to optimize. They have already enabled model caching. What next step should they take to reduce latency?

A.Add more GPUs to the prediction endpoint
B.Use a larger, more accurate model variant
C.Increase the batch size for inference requests
D.Apply model quantization to reduce precision
AnswerD

Quantization stores weights in lower precision, such as INT8, shrinking memory footprint and enabling faster matrix operations, which directly cuts inference latency. Caching is already applied, so reducing per-token compute is the remaining lever for real-time serving on Vertex AI.

Why this answer

Model quantization reduces the precision of the model's weights (e.g., from FP32 to INT8), which decreases memory footprint and accelerates computation on the hardware, directly lowering inference latency. Since Vertex AI already has model caching enabled, quantization is the next logical optimization step to reduce latency without requiring additional infrastructure changes.

Exam trap

Google often tests the misconception that adding more GPUs or increasing batch size reduces latency for real-time inference, when in fact these optimizations primarily improve throughput and can increase per-request latency.

How to eliminate wrong answers

Option A is wrong because adding more GPUs increases parallelism but does not reduce per-request latency; it primarily improves throughput and can even increase overhead due to inter-GPU communication. Option B is wrong because using a larger, more accurate model variant increases computational complexity and memory requirements, which would increase latency, not reduce it. Option C is wrong because increasing batch size improves throughput for batched requests but does not reduce latency for individual real-time inference requests; it may actually increase the time to complete a single request.

495
MCQeasy

A company is deciding between building a custom fine-tuned model vs. using a pre-built API for a document summarization task. The documents contain domain-specific jargon. Which factor STRONGLY favors using a pre-built API with prompt engineering?

A.The need to handle highly specialized industry terminology
B.The need for the model to learn a unique writing style from past summaries
C.The need to keep all data on-premises for security compliance
D.The requirement for low initial development cost and fast time-to-market
AnswerD

Low initial development cost and fast time-to-market strongly favour the pre-built API, since prompt engineering requires no labelled dataset, training compute, or model hosting. Fine-tuning demands curated domain examples and GPU training cycles, delaying deployment. This directly satisfies the stem's constraint: summarising jargon-heavy documents without the overhead of custom model development.

Why this answer

Using a pre-built API with prompt engineering eliminates the need for expensive model training infrastructure and specialized ML expertise, enabling rapid deployment at low initial cost. For a document summarization task, prompt engineering can leverage the API's existing capabilities without custom fine-tuning, making it ideal when speed and budget are primary constraints.

Exam trap

Google often tests the misconception that prompt engineering can fully replace fine-tuning for domain adaptation, when in reality prompt engineering is limited by context window size and cannot permanently encode specialized knowledge or writing styles.

How to eliminate wrong answers

Option A is wrong because handling highly specialized industry terminology actually favors fine-tuning, as pre-built APIs may lack domain-specific vocabulary and require additional prompt engineering that cannot fully capture nuanced jargon. Option B is wrong because learning a unique writing style from past summaries requires the model to internalize patterns through training data, which is a core strength of fine-tuning, not prompt engineering that only provides temporary context. Option C is wrong because keeping data on-premises for security compliance strongly favors building a custom model, as pre-built APIs typically require data to be sent to external servers, violating data residency requirements.

496
MCQeasy

A startup wants to quickly prototype a generative AI application that can write marketing copy. They have limited machine learning expertise and want to avoid managing infrastructure. They prefer a fully managed, no-code or low-code solution that provides access to Google's foundation models. Which Google Cloud offering should they use?

A.Vertex AI Studio
B.BigQuery ML
C.Vertex AI Pipelines
D.Cloud Functions
AnswerA

Vertex AI Studio is a Google Cloud console tool that allows users to quickly prototype and test generative AI models, including Gemini, for tasks like text generation. It provides a no-code interface for prompt design and model tuning, making it ideal for users with limited ML expertise who want to avoid infrastructure management.

Why this answer

Vertex AI Studio is the correct choice because it offers a user-friendly, no-code environment for experimenting with and prototyping generative AI models. It allows users to design prompts, test models like Gemini, and even fine-tune them without managing infrastructure. The other options are either too technical or not focused on generative AI prototyping.

Exam trap

The trap here is assuming that any Google Cloud AI service with 'AI' in the name is suitable for quick generative AI prototyping, when many require coding or are for different purposes.

497
Multi-Selectmedium

A product team is new to generative AI and wants to understand which characteristics distinguish foundation models from traditional task-specific ML models. Which two statements accurately describe foundation models? (Choose two.)

Select 2 answers
A.They are always smaller and cheaper to run than task-specific models because they do less work.
B.They must be trained from scratch for every new use case before they can produce any useful output.
C.They can generate novel content such as text, images, or code rather than only classifying existing inputs.
D.They typically require a large labeled dataset specific to each target task before they can generate text.
E.They are pretrained on broad data and can be adapted to many downstream tasks through prompting or fine-tuning.
AnswersC, E

Generative foundation models produce new content, including prose, summaries, images, and source code, instead of merely assigning a label to an input. This generative capability underpins use cases like drafting, translation, and code assistance. It distinguishes them from discriminative models whose output is a category or score.

Why this answer

Foundation models are pretrained on broad data, adapt to many tasks through prompting or fine-tuning, and can generate new content across modalities. These traits explain their flexibility and rapid reuse. The other statements describe classical supervised learning, require impractical retraining, or incorrectly claim foundation models are small and cheap, which misrepresents their scale and cost profile.

Exam trap

The trap here is mixing classical supervised-learning assumptions, such as needing large per-task labeled datasets, into descriptions of foundation models.

498
MCQhard

A logistics company wants to build a generative AI application that answers questions over thousands of internal policy PDFs stored in Cloud Storage. They need Google Cloud to handle document ingestion, chunking, indexing, and retrieval for grounding, while they focus only on the application logic. Which Google Cloud offering should they use?

A.Vertex AI Search
B.Cloud Vision API
C.Vertex AI Feature Store
D.Document AI
AnswerA

Vertex AI Search is the managed retrieval and grounding service that ingests documents from sources such as Cloud Storage, handles chunking and indexing, and serves relevant passages to ground generative responses. It removes the need to build a custom retrieval pipeline, letting the logistics team concentrate on application logic. This directly satisfies the ingestion, indexing, and retrieval requirements.

Why this answer

Vertex AI Search is purpose-built for ingesting enterprise documents, chunking and indexing them, and retrieving relevant passages to ground generative responses. It supports Cloud Storage as a data source and abstracts away the retrieval pipeline, so the logistics team can focus on application logic. The other services handle features, document parsing, or image analysis and do not deliver managed semantic retrieval for grounding.

Exam trap

The trap here is assuming any document-processing service provides retrieval, when Document AI and Cloud Vision extract content but do not index or serve passages for grounding.

499
MCQmedium

A logistics company wants its internal assistant to answer questions using its private fleet maintenance manuals. The manuals change weekly, and the company does not want to retrain a model each time. Accuracy must be traceable to the source document. Which Google Cloud approach should they implement?

A.Use Retrieval-Augmented Generation with Vertex AI Search as the grounding data store
B.Distill the manuals into a system instruction embedded in every prompt
C.Fine-tune a Gemini model on the maintenance manuals using supervised tuning
D.Increase the model temperature so it produces more detailed answers
AnswerA

RAG with Vertex AI Search retrieves relevant passages from the indexed manuals at query time and passes them to Gemini as grounding context. Because the index can be refreshed as documents change, no retraining is needed, and responses can cite the source passages, satisfying both the weekly update and traceability requirements.

Why this answer

Retrieval-Augmented Generation keeps the knowledge in an external index rather than in model weights. Vertex AI Search can index the private manuals, refresh them as they change, retrieve relevant passages per query, and return grounding metadata that supports citations. Fine-tuning, temperature changes, and static system instructions either require retraining, do not retrieve private data, or cannot scale to frequently updated documents.

Exam trap

The trap here is assuming that fine-tuning is the standard way to teach a model private knowledge, when frequently changing documents are better served by retrieval that keeps the index current.

500
MCQmedium

A logistics company wants to build a generative AI application that answers employee questions using its internal policy documents while keeping the data inside its own Google Cloud project. The team needs enterprise-grade security, access control, and the ability to choose among multiple foundation models. Which offering should they choose?

A.Google AI Studio
B.Vertex AI with Gemini models
C.Cloud Vision API
D.Gemini for Google Workspace
AnswerB

Vertex AI lets the company call Gemini models through an enterprise platform that enforces IAM, VPC Service Controls, and customer-managed encryption keys, all within its own Google Cloud project. It also exposes multiple foundation models and grounding options, so internal policy documents can drive answers without the data leaving the company's controlled environment, matching every stated requirement.

Why this answer

Vertex AI provides the enterprise controls, model flexibility, and grounding capabilities needed to build a policy-question answering application inside the company's own project. The productivity, prototyping, and vision services either lack the required governance or address entirely different workloads.

Exam trap

The trap here is confusing an easy prototyping tool with a production platform, since both can call Gemini models but only one provides enterprise governance.

501
Multi-Selectmedium

A team is using a generative AI model to answer customer questions about a complex product. They want to improve the factual accuracy and reduce hallucinations. Which TWO techniques should they apply? (Choose two.)

Select 2 answers
A.Use a grounded prompt that instructs the model to answer only from provided context and to say 'I don't know' if unsure.
B.Ground the model with Retrieval-Augmented Generation (RAG) using an authoritative product knowledge base.
C.Fine-tune the model on a small set of generic customer service dialogues.
D.Reduce the top-k parameter to a very small value.
E.Increase the temperature to make the model more confident.
AnswersA, B

A grounded prompt explicitly constrains the model to rely on the supplied context and to admit uncertainty when the answer is not present. This reduces hallucinations by discouraging the model from inventing information. It works well with RAG and is a simple, effective prompt engineering technique for factual accuracy.

Why this answer

Grounding with RAG supplies authoritative context, and a grounded prompt instructs the model to use only that context and to express uncertainty when appropriate. Together they attack hallucinations at the source: the model no longer relies solely on memorized parameters and is discouraged from fabricating. Other techniques like temperature or top-k affect randomness, not factual grounding, and fine-tuning on generic data does not provide the needed product facts.

Exam trap

The trap here is thinking that sampling parameters or generic fine-tuning can improve factual accuracy, when only grounding techniques like RAG and grounded prompts address hallucination directly.

502
MCQhard

A company is running a GenAI proof-of-concept (PoC) for internal document Q&A. The PoC shows high latency and cost. The team suspects they are using an unnecessarily large model for the task. What is the BEST action to optimize?

A.Disable grounding and rely solely on the model's internal knowledge
B.Increase the batch size for requests
C.Switch to a smaller model from Vertex AI Model Garden
D.Fine-tune the existing model on the company's documents
AnswerC

A smaller model from Vertex AI Model Garden reduces inference compute per request, directly cutting both latency and token cost. Since the PoC's task is internal document Q&A, a lighter model likely meets accuracy needs, satisfying the stem's optimisation goal without redesigning the pipeline.

Why this answer

Switching to a smaller model from Vertex AI Model Garden directly addresses the root cause of high latency and cost: an unnecessarily large model. Smaller models have fewer parameters, requiring less compute per inference, which reduces both response time and operational expense while often being sufficient for domain-specific tasks like internal document Q&A.

Exam trap

Google often tests the misconception that fine-tuning or disabling features like grounding can solve performance issues, when the real bottleneck is model size and compute efficiency.

How to eliminate wrong answers

Option A is wrong because disabling grounding removes the ability to retrieve and cite actual document content, forcing the model to rely on its internal knowledge which may be outdated or incorrect for company-specific documents, and does not address model size or latency. Option B is wrong because increasing batch size improves throughput for bulk processing but does not reduce per-request latency or cost for interactive Q&A; it may even increase memory pressure and latency for real-time use cases. Option D is wrong because fine-tuning the existing large model on company documents can improve answer relevance but does not reduce the model's parameter count or inference cost; it may actually increase latency if the fine-tuned model is still large and requires additional serving infrastructure.

503
MCQhard

A financial services firm is designing a generative AI assistant for advisors. Compliance requires that every response be traceable to an approved source document and that no response rely on the model's pretrained knowledge alone. The team plans to use Gemini models on Vertex AI. Which design choice most directly satisfies the traceability requirement?

A.Enable grounding with citations against a curated corpus of approved documents
B.Fine-tune the model on historical advisor emails to match the firm's tone
C.Increase the model's temperature so responses vary and appear less templated
D.Set a low maximum output token limit to keep answers short and reviewable
AnswerA

Grounding with citations ties each generated statement to retrieved source passages and returns references the firm can audit. This directly satisfies the requirement that responses trace back to approved documents and avoid unsupported pretrained knowledge. Curating the corpus ensures only sanctioned material is retrievable, which is essential in a regulated advisory context.

Why this answer

Grounding with citations retrieves passages from a curated approved corpus and returns references alongside generated text, giving auditors a clear link from each claim to its source. This prevents reliance on pretrained knowledge alone. Temperature, fine-tuning, and token limits influence style, behavior, or length but cannot establish provenance, so they do not satisfy the compliance mandate.

Exam trap

The trap here is equating fine-tuning or brevity controls with traceability, when only retrieval-based grounding with citations provides auditable source references.

504
MCQeasy

A marketing team uses a Gemini model through the Vertex AI API to draft campaign copy. The drafts are creative but frequently wander off topic and include unsupported claims. The team wants a low-effort improvement before considering any model customization. Which action should they take first?

A.Rewrite the prompt to include a clear role, the target audience, explicit constraints, and a required output format.
B.Raise the temperature so the model explores more ideas and eventually lands on better copy.
C.Switch the model to a larger version with more parameters to improve instruction following.
D.Create a supervised fine-tuning job with a few hundred labeled campaign examples.
AnswerA

Prompt engineering is the lowest-effort, highest-leverage first step. Specifying a role, audience, constraints, and output format narrows the model's generation space and typically removes off-topic drift and unsupported claims without any training cost. It can be iterated in minutes and does not require new data or infrastructure.

Why this answer

Prompt engineering with an explicit role, audience, constraints, and output format is the fastest and cheapest way to align model output with a brief. Fine-tuning, larger models, and higher temperature all add cost or increase variance without addressing the underlying ambiguity in the instructions.

Exam trap

The trap here is reaching for fine-tuning or a bigger model before exhausting prompt engineering, even though the symptoms point to underspecified instructions.

505
Multi-Selectmedium

A company wants to deploy a GenAI code review assistant that integrates into their existing Git workflow. They want to use a managed Google Cloud service to minimize operational overhead. Which TWO services should they consider? (Choose two.)

Select 2 answers
A.Cloud Build
B.Compute Engine
C.Cloud Spanner
D.Vertex AI Agent Builder
E.Cloud Run
AnswersA, D

Cloud Build is a managed CI/CD service that natively triggers builds from Git repositories, satisfying the Git-workflow integration and low operational overhead constraints. It runs the review pipeline serverlessly, so the team avoids provisioning or patching build infrastructure themselves.

Why this answer

Cloud Build (A) is correct because it is a managed CI/CD service that natively integrates with Git repositories (via Cloud Source Repositories, GitHub, or Bitbucket) and can run build/review steps defined in cloudbuild.yaml, letting the code review assistant trigger automatically on commits and pull requests with minimal operational overhead. Vertex AI Agent Builder (D) is correct because it is a managed Google Cloud service for building, deploying, and orchestrating GenAI agents and applications, providing the LLM/reasoning layer needed for an AI code review assistant without managing infrastructure. Compute Engine (B) is not appropriate because it is raw IaaS VMs, which would require the company to manage OS patching, scaling, and runtime, contradicting the goal of minimizing operational overhead.

Cloud Spanner (C) is a globally distributed relational database and does not provide CI/CD or GenAI agent capabilities. Cloud Run (E) is a managed serverless container platform, but by itself it is a generic compute runtime rather than the Git-integrated build pipeline or GenAI agent framework the scenario requires.

Exam trap

Google often tests the distinction between managed CI/CD services (Cloud Build) and general-purpose compute (Cloud Run, Compute Engine), where candidates mistakenly choose Cloud Run because it is serverless, but it lacks native Git integration for automated code review triggers.

506
MCQmedium

A company has a Gemini-based application that sometimes produces factually incorrect answers. They want to improve accuracy without retraining the model. Which technique should they implement?

A.Reduce the top-k value to 10
B.Implement Retrieval-Augmented Generation (RAG) with a curated knowledge base
C.Use prompt engineering to instruct the model to 'be more accurate'
D.Increase the temperature to 1.0 for more diverse outputs
AnswerB

Retrieval-Augmented Generation grounds responses in a curated knowledge base, retrieving relevant documents at inference time so the model cites verified content rather than relying on parametric memory. This directly addresses the factual inaccuracy constraint without retraining, since the Gemini weights stay frozen while accuracy improves through external retrieval.

Why this answer

Retrieval-Augmented Generation (RAG) is the correct technique because it grounds the model's responses in a curated, external knowledge base, providing factual context that reduces hallucinations without modifying the model's weights. This directly addresses the need for improved accuracy in a Gemini-based application while avoiding costly retraining.

Exam trap

This question tests the misconception that prompt engineering alone can fix factual accuracy issues, when in reality, without external knowledge grounding, the model remains reliant on its parametric memory which is prone to hallucination.

How to eliminate wrong answers

Option A is wrong because reducing top-k to 10 limits the number of candidate tokens considered during generation, which can reduce diversity but does not inherently improve factual accuracy; it may even cause the model to miss correct but less probable tokens. Option C is wrong because instructing the model to 'be more accurate' via prompt engineering is a vague directive that does not provide new factual information; the model cannot correct its own knowledge gaps through simple instruction. Option D is wrong because increasing temperature to 1.0 increases output randomness and diversity, which typically worsens factual accuracy by encouraging the model to explore less probable, often incorrect, token sequences.

507
MCQmedium

A company has a large dataset of customer support tickets stored in BigQuery. They want to predict ticket severity (high, medium, low) using SQL queries without moving data out of BigQuery. Which service should they use?

A.Vertex AI AutoML Natural Language
B.Cloud Natural Language API
C.Gemini API
D.BigQuery ML
AnswerD

BigQuery ML trains and runs models directly inside BigQuery using SQL, so the ticket data never leaves the warehouse. This satisfies the stem's constraint of predicting three-class severity without data movement, unlike exporting to a separate ML platform.

Why this answer

BigQuery ML (D) is correct because it allows users to create and execute machine learning models directly within BigQuery using SQL, without moving data out of the warehouse. For a classification task like predicting ticket severity, BigQuery ML supports models such as logistic regression, boosted trees, and deep neural networks, all trained and deployed using standard SQL queries on data already in BigQuery.

Exam trap

The certification often tests the distinction between services that require data movement (like Vertex AI AutoML) versus services that operate directly on the data warehouse (like BigQuery ML), and the trap here is assuming that any ML service in Google Cloud must involve exporting data to a separate AI platform.

How to eliminate wrong answers

Option A is wrong because Vertex AI AutoML Natural Language requires exporting data from BigQuery to a Cloud Storage bucket and then importing it into Vertex AI, which violates the requirement of not moving data out of BigQuery. Option B is wrong because Cloud Natural Language API is a pre-trained API for sentiment analysis, entity extraction, and syntax analysis, not a custom classification model that can be trained on the company's specific ticket severity labels. Option C is wrong because Gemini API is a generative AI API for tasks like text generation and summarization, not designed for custom classification model training or SQL-based ML workflows within BigQuery.

508
MCQhard

A company is using a fine-tuned LLM for generating financial reports. They need to ensure that the output complies with regulatory standards and does not include speculative content. Which combination of techniques should they implement?

A.Increase the model's safety settings to maximum, use a low top-p value, and limit output tokens.
B.Fine-tune the model on historical compliant reports, use RAG with a regulatory database, and implement a human-in-the-loop review.
C.Use a larger model with more parameters and rely on its inherent knowledge.
D.Use a system instruction to adhere to regulations, set temperature to 0.0, and apply a keyword filter.
AnswerB

Fine-tuning on historical compliant reports instils regulatory tone and structure, while RAG grounds each generation in current regulatory text, preventing speculative output. Human-in-the-loop review provides the final compliance gate. Together these satisfy the stem's dual constraint: regulatory compliance and absence of speculative content.

Why this answer

Fine-tuning on historical compliant reports ensures the model learns from past regulatory requirements, RAG with a regulatory database provides up-to-date compliance information, and human-in-the-loop review adds a final verification layer to catch any non-compliant or speculative content.

Option A (safety settings, low top-p, limit tokens) may reduce harmful content but does not guarantee regulatory compliance. Option C (larger model) alone does not enforce specific regulations. Option D (system instruction, temperature 0.0, keyword filter) is insufficient for complex regulatory standards.

509
MCQeasy

A small marketing agency wants to experiment with Google's generative AI models without writing code or managing cloud infrastructure. They need a browser-based environment to draft campaign ideas and test prompts quickly. Which Google Cloud offering should they use?

A.Vertex AI Studio
B.BigQuery ML with remote model inference
C.Vertex AI Pipelines
D.Cloud Run functions calling the Gemini API
AnswerA

Vertex AI Studio provides a browser-based interface for designing, testing, and refining prompts against Gemini models without requiring code or infrastructure management. It is intended for rapid experimentation, letting the agency compare model outputs and tune parameters interactively, which fits the need to draft campaign ideas quickly.

Why this answer

Vertex AI Studio is the console-based workspace for prompt design and model testing, so a non-technical marketing team can draft and compare campaign ideas in a browser without code. The other choices require SQL, pipeline definitions, or deployed serverless code, none of which match a quick, no-code ideation need.

Exam trap

The trap here is confusing a code-first serving or orchestration service with a no-code prompt design workspace.

510
Multi-Selectmedium

A company is deploying a generative AI system for resume screening. They want to ensure fairness and avoid bias. Which TWO actions should they take? (Choose 2)

Select 2 answers
A.Evaluate the model's decisions across gender and ethnicity groups using a diverse test set
B.Ensure training data is representative of the candidate population
C.Apply SynthID watermarking to all generated decisions
D.Remove protected attributes like gender and race from training data
E.Use a larger model to improve accuracy
AnswersA, B

Evaluating decisions across gender and ethnicity groups directly satisfies the fairness constraint by measuring disparate impact on protected attributes. A diverse test set exposes performance gaps that aggregate accuracy hides, revealing whether the model disadvantages specific demographics. This subgroup analysis is the technical mechanism for detecting bias before deployment in resume screening.

Why this answer

Option A is correct because fairness auditing requires measuring model outcomes across demographic subgroups (e.g., gender, ethnicity) using a diverse, representative test set, so disparate impact or bias can be detected and quantified before deployment. Option B is correct because if the training data is not representative of the candidate population, the model will learn skewed patterns and produce biased predictions, so ensuring representative training data is a foundational bias-mitigation step. Option C is incorrect because SynthID is a watermarking technique for identifying AI-generated content, not a fairness or bias-mitigation control for screening decisions.

Option D is incorrect because simply removing protected attributes does not eliminate bias, since proxy variables and correlated features can still encode discriminatory patterns. Option E is incorrect because a larger model improves capacity or accuracy but does not inherently ensure fairness or reduce bias.

511
MCQhard

A data scientist is using Vertex AI RAG Engine to build a question-answering system over a large corpus of technical manuals. Users report that answers are often verbose and include irrelevant details. Which configuration change is MOST likely to improve answer conciseness?

A.Increase the chunk size of documents in the index
B.Use a different embedding model for retrieval
C.Decrease the number of retrieved documents (k) in the retrieval step
D.Increase the max output token limit
AnswerC

Lowering k reduces the volume of retrieved context passed to the generative model, directly limiting the extraneous material it can draw into the response. Since verbosity stems from irrelevant retrieved chunks, fewer documents constrain the model to the most pertinent passages, improving conciseness without altering the underlying corpus or prompt.

Why this answer

In a RAG pipeline, the number of retrieved documents (k) directly controls how much context is fed to the LLM. Reducing k limits the amount of potentially irrelevant material the model sees, which typically makes answers shorter and more focused. This is the most direct lever for conciseness.

Exam trap

The trap is confusing retrieval quality with generation length — candidates pick embedding model changes or chunk size, but the direct control on answer verbosity is the number of retrieved documents (k) fed into the prompt.

How to eliminate wrong answers

Option A is wrong because increasing chunk size gives the model larger, more diffuse passages, which usually makes answers more verbose, not less. Option B is wrong because swapping embedding models changes retrieval relevance, not the verbosity of the generated answer — a better embedding model may improve precision but does not directly control output length. Option D is wrong because increasing the max output token limit allows longer answers, which is the opposite of what is needed.

512
MCQmedium

A company is building a legal document review assistant using Gemini 1.5 Pro. They want to ensure the model can handle large documents of up to 500 pages in a single prompt. Which feature of Gemini 1.5 Pro is MOST important for this requirement?

A.Temperature
B.Top-k sampling
C.Fine-tuning
D.Large context window
AnswerD

A large context window lets Gemini 1.5 Pro ingest up to 500 pages of text within a single prompt, satisfying the stem's single-prompt constraint. Unlike retrieval-augmented approaches that chunk documents externally, the native long-context capability preserves cross-references across the entire legal document, which matters for coherent review.

Why this answer

Gemini 1.5 Pro has a large context window (up to 1 million tokens), allowing it to process large documents. Temperature, top-k, and fine-tuning do not directly address context length.

513
MCQeasy

A financial services company wants to build a generative AI application that drafts personalized emails to clients. They require the model to be accessible via a fully managed API with minimal infrastructure management. Which Google Cloud service should they use?

A.AutoML Natural Language
B.Vertex AI Gemini API
C.Dialogflow CX
D.Cloud Natural Language API
AnswerB

The Vertex AI Gemini API provides access to Google's Gemini models through a fully managed endpoint, abstracting infrastructure management. It supports text generation tasks like drafting personalized emails, and integrates with other Google Cloud services for security and scalability. This aligns with the requirement for minimal operational overhead.

Why this answer

The Vertex AI Gemini API offers a managed, serverless way to access powerful generative models. It eliminates infrastructure management, supports text generation, and integrates with Google Cloud's security and monitoring. Other options are either not generative or require custom model training.

Exam trap

The trap here is confusing pre-trained NLP APIs that analyze text with generative APIs that create new text.

514
MCQmedium

An e-commerce company wants to generate realistic product images from text descriptions using Google Cloud AI. Which service should they use?

A.Imagen on Vertex AI
B.Gemini 1.5 Pro
C.Chirp
D.Vertex AI Codey
AnswerA

Imagen on Vertex AI generates photorealistic images directly from text prompts, satisfying the requirement to create realistic product visuals from descriptions. Unlike Vertex AI Search or Vision API, which retrieve or analyse existing media, Imagen synthesises new imagery, making it the appropriate Google Cloud service for this text-to-image scenario.

Why this answer

Imagen on Vertex AI is the correct service because it is specifically designed for text-to-image generation, allowing the e-commerce company to create realistic product images from textual descriptions. It leverages Google's advanced diffusion models to produce high-fidelity visuals, making it the ideal choice for this use case.

Exam trap

The trap here is that candidates may confuse Gemini's multimodal capabilities (which can process images but not generate them from scratch) with a dedicated image generation service, leading them to select Gemini 1.5 Pro instead of Imagen.

How to eliminate wrong answers

Option B (Gemini 1.5 Pro) is wrong because it is a multimodal large language model focused on understanding and generating text, code, and images in a conversational context, not a dedicated text-to-image generation service. Option C (Chirp) is wrong because it is a speech-to-text and text-to-speech model for audio processing, not for generating images. Option D (Vertex AI Codey) is wrong because it is a code generation and completion model for software development, not for creating visual content.

515
MCQeasy

A marketing team is using an LLM on Vertex AI to generate product descriptions. They want to consistently control the creativity and randomness of the output. Which parameter should they adjust?

A.Top-P
B.Temperature
C.Max output tokens
D.Top-K
AnswerB

Temperature controls the randomness of the model's output. Lower values make the output more deterministic and focused, while higher values increase creativity and diversity. In this scenario, adjusting temperature allows the team to consistently control the creativity of the generated product descriptions, aligning with their goal.

Why this answer

Temperature is the primary parameter for controlling the randomness and creativity of generative AI output. Lower temperatures yield more predictable, focused responses, while higher temperatures yield more diverse and creative ones. For a marketing team seeking consistent control over creativity, temperature is the correct choice.

Exam trap

The trap here is confusing sampling parameters like Top-K and Top-P with temperature, which directly governs creativity.

516
MCQmedium

A company uses a Gemini 1.5 Pro model with a 1 million token context window. They want to process a large 500-page PDF for Q&A. What is the MAIN advantage of using the long context window over a RAG approach?

A.Simpler architecture with no need for chunking or retrieval systems
B.Lower cost per API call
C.Faster inference speed
D.Higher accuracy for all queries
AnswerA

Long context removes the chunking, embedding and vector-retrieval pipeline entirely, so the whole 500-page PDF is passed directly to Gemini 1.5 Pro in a single prompt. This satisfies the scenario's need to process the document for Q&A without building or maintaining separate retrieval infrastructure.

Why this answer

The primary advantage of using a 1 million token context window is architectural simplicity. By ingesting the entire 500-page PDF as a single prompt, the company eliminates the need for document chunking, embedding generation, and a retrieval system (RAG). This reduces system complexity, maintenance overhead, and potential failure points, as the model can directly attend to all content in one pass.

Exam trap

Candidates often mistakenly think that a larger context window is always cheaper, faster, or more accurate, when in reality it trades off simplicity for higher cost, slower inference, and potential accuracy degradation for mid-context information.

How to eliminate wrong answers

Option B is wrong because a 1 million token context window incurs significantly higher cost per API call due to the massive input token count, making it more expensive than a RAG approach that only sends relevant chunks. Option C is wrong because processing a 1 million token context window is slower than RAG, as the model must compute attention over all tokens, leading to higher latency. Option D is wrong because a long context window does not guarantee higher accuracy for all queries; RAG can be more accurate for specific, localized questions by retrieving only relevant information, while long context models may suffer from 'lost in the middle' effects where information in the middle of the context is less attended to.

517
Multi-Selecthard

Which THREE of the following are key considerations when deploying a generative AI model in a production environment with strict latency requirements? (Choose three.)

Select 3 answers
A.Deploy the largest model variant available to ensure highest quality.
B.Implement speculative decoding to generate candidate tokens with a smaller draft model and verify with the large model.
C.Use model quantization (e.g., int8) to reduce precision and speed up matrix multiplications.
D.Cache the key-value caches from previous decoding steps to avoid redundant computation.
E.Increase the inference batch size to maximize GPU utilization.
AnswersB, C, D

Speculative decoding attacks latency directly: a smaller draft model proposes candidate tokens cheaply, and the large model verifies them in parallel, preserving output quality while cutting sequential decoding steps. This meets the strict latency requirement without retraining or accuracy loss.

Why this answer

Option B is correct because speculative decoding uses a smaller, faster draft model to propose multiple candidate tokens that the larger target model then verifies in parallel, reducing the number of expensive forward passes and lowering end-to-end latency while preserving output quality. Option C is correct because quantizing weights and activations to lower precision such as int8 shrinks memory bandwidth demands and enables faster matrix multiplications on supported hardware, directly cutting per-token inference time. Option D is correct because caching the key-value tensors from prior decoding steps avoids recomputing attention over the entire sequence for each new token, which is essential for keeping autoregressive generation latency low.

Option A is not appropriate because the largest model variant increases compute and memory cost, worsening latency rather than meeting strict requirements. Option E is not appropriate because increasing batch size improves throughput and GPU utilization but typically raises per-request latency, which conflicts with strict latency goals.

Exam trap

Google Cloud often tests the distinction between latency and throughput, so the trap here is that candidates confuse batch size (which improves throughput) with latency reduction, or assume larger models always yield better performance without considering inference speed.

518
MCQmedium

A company is using a generative AI model to automatically screen job applications. They want to ensure the model does not discriminate based on gender or ethnicity. Which of the following actions should they take as part of responsible AI?

A.Use a more powerful model to improve accuracy, as bias decreases with better performance
B.Blindly apply a fairness algorithm without understanding the context
C.Remove all demographic data from the training set to prevent the model from learning biases
D.Evaluate the model's hiring decisions for disparate impact across gender and ethnicity
AnswerD

Evaluating hiring decisions for disparate impact directly measures whether outcomes disadvantage protected groups, satisfying the stem's fairness constraint. Unlike input-only checks, this outcome-based audit detects proxy discrimination that emerges from correlated features, revealing bias the model's training data introduced. It provides the empirical evidence needed to remediate gender and ethnicity disparities before deployment.

Why this answer

Responsible AI requires measuring outcomes, not just cleaning inputs. Evaluating hiring decisions for disparate impact across gender and ethnicity directly tests whether the model produces discriminatory outcomes, using metrics like the 4/5ths rule or statistical parity. This is the only option that actually verifies fairness in production decisions.

Exam trap

The trap is the intuitive but false belief that removing protected attributes (fairness through unawareness) eliminates bias — exam writers include this as a distractor because it sounds responsible, but proxy variables and historical bias persist.

How to eliminate wrong answers

Option A is wrong because model accuracy and fairness are independent dimensions — a more powerful model can still encode and amplify historical bias present in training data; bias does not automatically decrease with performance. Option B is wrong because blindly applying a fairness algorithm without understanding context can introduce new harms, violate legal constraints, or optimize the wrong metric; fairness interventions must be context-aware and validated. Option C is wrong because removing demographic data does not prevent bias — proxy variables (zip code, name, school) can leak protected attributes, and the model can still learn discriminatory patterns from correlated features.

519
Multi-Selectmedium

A research team wants to train a very large transformer model using Google's custom AI accelerators. They need the highest compute density and tightest interconnection for distributed training. Which THREE are true about Google's TPU infrastructure? (Choose 3)

Select 3 answers
A.TPU pods can scale to thousands of chips with low-latency interconnects
B.TPU v3 is the latest generation offering the best performance
C.TPU v4 pods offer high-bandwidth interconnects for large-scale distributed training
D.GPUs are Google's primary accelerator for large transformer training
E.TPU v5e is designed for cost-efficient inference and small-scale training
AnswersA, C, E

TPU pods are designed to scale to thousands of TPU chips with high-speed interconnects for efficient distributed training.

Why this answer

TPU pods are designed to scale to thousands of TPU chips using a custom high-speed interconnect (e.g., the 2D torus mesh in TPU v2/v3 and the 3D torus in TPU v4), which provides low-latency, high-bandwidth communication essential for distributed training of very large transformer models. This architecture allows the research team to achieve the highest compute density and tightest interconnection for their workload.

Exam trap

Candidates may assume TPU v3 is the latest generation because it is widely documented in older materials, but Google has since released v4, v5e, and v5p, with v5p being the current top performer for training.

520
MCQeasy

A media company wants to generate realistic images for a new marketing campaign. They need a Google Cloud service that can create images from text prompts and offers enterprise-grade controls for content safety and intellectual property. Which service should they use?

A.Vertex AI Vector Search
B.Vertex AI Imagen
C.Cloud Vision API
D.Vertex AI Gemini
AnswerB

Vertex AI Imagen is Google Cloud's text-to-image generation model, designed for enterprise use. It provides high-quality image generation from text prompts and includes safety filters and controls to prevent harmful or copyrighted content. It integrates with Vertex AI for management, monitoring, and compliance, making it the right choice for a media company's marketing needs.

Why this answer

Vertex AI Imagen is specifically built for generating images from text prompts with enterprise-grade safety and IP controls. It is the Google Cloud offering for text-to-image generation, making it the correct choice for a media company that needs realistic images for marketing while ensuring compliance and safety.

Exam trap

The trap here is assuming that Gemini can handle all generative tasks equally well; while Gemini can generate images, Imagen is the specialized model for high-quality text-to-image generation.

521
MCQhard

A global company deploying gen AI across multiple regions needs to minimize latency and comply with data sovereignty. What architecture should they adopt?

A.Single global deployment with CDN
B.Multi-region deployment with Vertex AI
C.Use a third-party API
D.On-premises deployment only
AnswerB

Multi-region deployment with Vertex AI places model endpoints and inference within each region, so requests are served locally rather than crossing borders. This directly satisfies the data sovereignty constraint, since data residency is preserved per region, while regional endpoints cut network round-trip time, addressing the latency requirement.

Why this answer

A multi-region deployment using a platform like Vertex AI lets the company place model endpoints and data processing in each region where users and regulatory requirements exist, minimizing network latency while keeping data resident within sovereign boundaries. Vertex AI supports regional endpoints and data residency controls, so inference and training can occur locally rather than routing everything through one jurisdiction. This directly satisfies both the latency and data sovereignty constraints simultaneously.

Exam trap

The trap here is assuming a CDN or global load balancer solves latency for AI workloads; CDNs cache static assets, not dynamic model inference, so candidates who conflate content delivery with compute placement pick the wrong answer.

How to eliminate wrong answers

Option A is wrong because a single global deployment with a CDN accelerates static content delivery but does not solve data sovereignty (data still resides in one region) and adds latency for inference calls that must traverse to the origin region. Option C is wrong because a third-party API typically routes data to the vendor's infrastructure, which usually violates data sovereignty requirements and adds unpredictable network latency. Option D is wrong because on-premises-only deployment cannot serve a global user base with low latency and sacrifices the elasticity and managed services needed for gen AI at scale.

522
MCQmedium

You are an ML engineer at a retail company. You have deployed a generative AI model on Vertex AI to generate product descriptions. The model uses a custom container and is deployed to a single endpoint. Recently, you noticed that inference latency has increased significantly during peak hours, causing timeouts. You have checked the logs and found that the CPU utilization on the deployed instances is consistently above 90% during peak hours. The model is currently deployed with a single machine type (n1-standard-4) and no scaling. You need to reduce latency without incurring excessive cost. What should you do?

A.Optimize the model using quantization and reduce the number of replicas
B.Switch to batch prediction instead of online prediction
C.Change the machine type to n1-standard-8 and enable autoscaling with min replicas=1, max replicas=5
D.Add a GPU accelerator to the existing machine
AnswerC

CPU is saturated above 90%, so the bottleneck is compute per replica. Doubling to n1-standard-8 gives each replica more vCPU headroom, and autoscaling to five replicas absorbs peak-hour concurrency while min replicas=1 keeps idle cost low.

Why this answer

Upgrading to a larger machine (n1-standard-8) provides more CPU cores to handle the increased inference workload, while enabling autoscaling (min=1, max=5) allows the deployment to dynamically add replicas during peak hours to distribute the load and reduce latency. This combination addresses the high CPU utilization without over-provisioning during off-peak times, thus controlling cost.

Exam trap

The trap here is that candidates often assume adding a GPU (Option D) is always the best way to reduce inference latency, but for CPU-bound models with high utilization, scaling out with more replicas and a larger CPU machine is more cost-effective and directly addresses the bottleneck.

How to eliminate wrong answers

Option A is wrong because quantization reduces model size and can improve latency, but reducing the number of replicas would worsen the bottleneck by decreasing capacity, not help with high CPU utilization. Option B is wrong because batch prediction is designed for asynchronous, non-real-time processing and does not solve online inference latency during peak hours; it would change the use case entirely. Option D is wrong because adding a GPU accelerator to the existing n1-standard-4 machine would not address the CPU bottleneck (inference is CPU-bound in this scenario) and would increase cost unnecessarily without guaranteed latency improvement for a CPU-bound model.

523
MCQmedium

A company is using a generative AI model to screen job applications. They want to ensure compliance with Google's AI Principle of avoiding unfair bias. Which practice is most effective in mitigating bias during the screening process?

A.Use a pre-trained model without modification
B.Remove all demographic information from the resumes before processing
C.Audit the training data for demographic representativeness and evaluate the model using fairness metrics
D.Only allow human reviewers to see the top 10% of candidates
AnswerC

Auditing training data exposes under-representation that skews outputs, while fairness metrics such as demographic parity quantify disparate impact across protected groups. This directly satisfies the principle of avoiding unfair bias, since bias originates in data and must be measured, not assumed absent.

Why this answer

Auditing training data for demographic representativeness and evaluating the model using fairness metrics directly addresses the root causes of bias in generative AI systems. This practice aligns with Google's AI Principle of avoiding unfair bias by proactively identifying and mitigating imbalances in the data and measuring the model's performance across demographic groups using metrics like demographic parity or equal opportunity.

Exam trap

Google often tests the misconception that removing demographic features is sufficient to eliminate bias, but the trap here is that proxy variables and latent correlations in the data can still cause the model to discriminate indirectly.

How to eliminate wrong answers

Option A is wrong because using a pre-trained model without modification can propagate and amplify existing biases present in the training data, as the model may have learned skewed correlations from unrepresentative or biased datasets. Option B is wrong because simply removing demographic information from resumes does not eliminate bias; the model can still infer protected attributes from proxy features such as names, zip codes, or educational institutions, leading to indirect discrimination. Option D is wrong because only allowing human reviewers to see the top 10% of candidates does not mitigate bias in the initial AI screening; it merely shifts the decision point and can still reflect the model's biased rankings, while human reviewers may also introduce their own unconscious biases.

524
MCQmedium

A company is rolling out a generative AI tool to employees. To ensure successful adoption, they plan to provide training and identify early adopters. Which change management practice is MOST critical early in the rollout?

A.Set up a dedicated help desk for AI tool issues
B.Conduct mandatory training for all employees before launch
C.Identify and train AI champions who can support their teams
D.Monitor usage metrics for the first month
AnswerC

Identifying and training AI champions directly satisfies the stem's need to build peer support during early rollout. Champions are influential early adopters who model usage, answer colleagues' questions and reduce resistance, embedding the tool within existing team workflows far faster than centralised training alone achieves.

Why this answer

Identifying AI champions who can advocate and help peers is crucial for organic adoption. Training all employees at once is less effective. Monitoring usage is important but later.

Setting up a help desk is reactive.

525
MCQhard

The exhibit shows the deployment configuration for a conversational AI model used in a finance application. Users report that responses are creative but often contain factually incorrect financial advice. Which parameter change would most improve factual accuracy?

A.Add grounding sources, such as "EnterpriseSearch" or "Web"
B.Lower temperature to 0.1
C.Increase topP to 1.0
D.Increase maxOutputTokens to 1024
AnswerA

Adding grounding sources constrains generation to retrieved enterprise or web content, so the model cites evidence rather than relying on parametric memory. This directly addresses the stem's factual-accuracy failure, where ungrounded creativity produces incorrect financial advice. Retrieval augmentation, not sampling parameters, is the mechanism that anchors responses in verifiable data.

Why this answer

Adding grounding sources like EnterpriseSearch or Web provides the model with access to authoritative, up-to-date financial data, which directly reduces hallucinations by anchoring responses in verified facts rather than relying solely on the model's parametric knowledge. This is the most effective technique for improving factual accuracy in a domain where correctness is critical.

Exam trap

The Generative AI Leader exam often tests the misconception that adjusting sampling parameters (temperature, topP) can fix factual accuracy issues, when in reality those parameters only control output randomness and diversity, not the truthfulness of the underlying knowledge.

How to eliminate wrong answers

Option B is wrong because lowering temperature to 0.1 makes the model more deterministic and less creative, but it does not introduce new factual information; it only reduces randomness in token selection, which cannot fix incorrect knowledge baked into the model. Option C is wrong because increasing topP to 1.0 includes all possible tokens in the sampling pool, which actually increases the chance of selecting less likely and potentially incorrect tokens, harming factual accuracy. Option D is wrong because increasing maxOutputTokens to 1024 allows longer responses but does not improve the correctness of the content; it may even amplify errors by generating more text based on the same flawed internal knowledge.

Page 6

Page 7 of 14

Page 8