Courseiva

Google Cloud Generative AI Leader Generative AI Leader (Generative AI Leader) — Questions 676–750

1008 questions total · 14pages · All types, answers revealed

Page 9

Page 10 of 14

Page 11
676
MCQhard

A healthcare startup fine-tunes a model to generate patient education materials. They want to ensure the model never gives medical advice, only information. They add a safety instruction, but the model sometimes still gives advice. What advanced technique should they apply?

A.Hard-code a list of prohibited phrases in a post-processing script
B.Add a secondary classifier to rewrite any detected advice into general information
C.Use semantic similarity to a 'medical advice' embedding and reject if close
D.Apply RLHF with a reward model that penalizes outputs containing medical advice
AnswerD

RLHF trains a reward model that scores outputs, then optimises the model against it; penalising medical-advice content directly shapes generation away from advice. A single safety instruction only conditions the prompt, so it cannot reliably enforce the constraint across varied inputs.

Why this answer

RLHF (Reinforcement Learning from Human Feedback) directly addresses the model's behavior by training a reward model that penalizes outputs containing medical advice. This aligns the model's generation with the safety instruction at a fundamental level, rather than relying on brittle post-hoc filters or static embeddings that can be easily circumvented by novel phrasings.

Exam trap

Candidates often mistakenly believe that simple post-processing filters or static embedding comparisons are sufficient to enforce safety. However, only advanced alignment techniques like RLHF can truly align the model's generation, as it changes the model's behavior during training rather than applying brittle surface-level checks.

How to eliminate wrong answers

Option A is wrong because hard-coding a list of prohibited phrases is brittle and fails against adversarial or paraphrased advice that doesn't match the exact phrases. Option B is wrong because adding a secondary classifier to rewrite detected advice introduces latency, potential for semantic drift, and cannot handle nuanced contexts where advice is implied rather than explicit. Option C is wrong because semantic similarity to a static 'medical advice' embedding is threshold-dependent and can produce false positives (flagging general information) or false negatives (missing advice phrased differently), and it does not train the model to avoid the behavior.

677
Multi-Selecteasy

A developer wants to use the Gemini API to generate creative text. Which TWO parameters can they adjust to influence the output?

Select 2 answers
A.Color space
B.Audio sample rate
C.Top-k
D.Image size
E.Temperature
AnswersC, E

Top-k truncates the sampling pool to the k most probable next tokens, so lowering it makes output more focused and raising it increases variety. This is a sampling parameter that directly shapes creative text generation, satisfying the stem's requirement for a parameter that influences the Gemini API's output.

Why this answer

Top-k (C) is correct because it limits sampling to the k most probable next tokens, directly shaping the randomness and creativity of the generated text. Temperature (E) is correct because it scales the logits before the softmax, controlling how deterministic or diverse the model's token choices are. Color space (A) is irrelevant since it pertains to image pixel encoding, not text generation parameters.

Audio sample rate (B) applies to audio processing, not Gemini text output. Image size (D) affects image dimensions, which has no bearing on text generation creativity.

Exam trap

Google exams often test the distinction between model parameters that affect text generation (like Temperature and Top-k) versus media-specific parameters (like color space or image size), leading candidates to confuse domain-specific settings with generative AI controls.

678
MCQeasy

A data scientist wants to quickly prototype a text generation application using Google's foundation models. Which Google Cloud service should they use?

A.Generative AI Studio
B.Cloud Natural Language API
C.Vertex AI Prediction
D.AI Platform Training
AnswerA

Generative AI Studio provides a console and API for prompt design, tuning and rapid testing of Google's foundation models, including Gemini, without infrastructure setup. This directly satisfies the stem's prototyping constraint, unlike Vertex AI Pipelines or BigQuery ML.

Why this answer

Generative AI Studio is the correct service because it provides a purpose-built environment for quickly prototyping and experimenting with Google's foundation models, including text generation models like PaLM 2 and Gemini. It offers a no-code interface and SDK access for rapid iteration, directly aligning with the data scientist's goal of fast prototyping without needing to manage infrastructure or training pipelines.

Exam trap

The trap here is that candidates confuse the purpose of Cloud Natural Language API (a non-generative analysis tool) with generative AI capabilities, or assume Vertex AI Prediction is the correct choice for prototyping when it is actually designed for serving deployed models, not interactive experimentation.

How to eliminate wrong answers

Option B is wrong because Cloud Natural Language API is a pre-trained API for analyzing text (e.g., sentiment, entity extraction) and does not support generative text generation or foundation model prototyping. Option C is wrong because Vertex AI Prediction is used for deploying and serving trained models for inference, not for rapid prototyping or interactive experimentation with foundation models. Option D is wrong because AI Platform Training (now part of Vertex AI) is designed for training custom machine learning models, not for quickly prototyping with pre-built foundation models.

679
MCQmedium

A financial services firm wants to use generative AI to assist employees in drafting emails. They are evaluating Duet AI in Gmail (now Gemini in Workspace). Which capability directly supports this use case?

A.Speaker notes generation in Slides
B.Smart Compose in Gmail
C.Meeting summaries in Meet
D.Help me write in Docs
AnswerB

Smart Compose generates context-aware sentence suggestions as the employee types in Gmail, directly supporting email drafting. It predicts phrasing from the existing message content, so staff compose replies faster without leaving the interface, matching the stated drafting-assistance use case.

Why this answer

Smart Compose in Gmail is a generative AI feature that provides real-time, context-aware suggestions to help users draft emails faster. It directly supports the use case of assisting employees in drafting emails by predicting and completing sentences as they type, reducing effort and improving efficiency.

Exam trap

Google often tests the distinction between generative AI features across different Workspace apps, and the trap here is confusing 'Help me write in Docs' (a document-focused tool) with the email-specific Smart Compose in Gmail, leading candidates to pick a feature that is not directly integrated into the email drafting workflow.

How to eliminate wrong answers

Option A is wrong because Speaker notes generation in Slides is designed to create presentation notes, not assist with email drafting. Option C is wrong because Meeting summaries in Meet provide post-meeting recaps, not real-time email composition assistance. Option D is wrong because Help me write in Docs is a generative AI feature for document creation in Google Docs, not for drafting emails within Gmail.

680
MCQmedium

A company wants to use a generative AI model to create product descriptions from a list of features. They need the model to consistently follow a specific format: a headline, followed by three bullet points, and a closing sentence. Which technique should they use to guide the model's output structure?

A.Increasing the temperature to allow more creative freedom.
B.Using a larger model with more parameters.
C.Setting top-k to 1 to force deterministic output.
D.Few-shot prompting with examples of the desired format.
AnswerD

Few-shot prompting provides the model with several examples of the input-output format, allowing it to learn the pattern and apply it to new inputs. This is effective for enforcing a consistent structure like headline, bullets, and closing. It leverages in-context learning without retraining the model.

Why this answer

Few-shot prompting is ideal for teaching the model a specific output format by providing examples. It allows the model to infer the pattern and apply it consistently. Other options like increasing temperature or changing model size do not directly address the need for structured output.

Exam trap

The trap here is assuming that a larger model or deterministic settings will automatically produce the desired format, when prompt design is the key.

681
MCQeasy

A startup's developer wants to quickly test different prompts against Gemini models, compare model outputs side by side, and export working prompt code, all from a browser interface before committing to an application architecture. Which Google Cloud offering should the developer use?

A.Vertex AI Studio
B.Vertex AI Model Garden
C.Gemini for Google Workspace
D.Vertex AI Pipelines
AnswerA

Vertex AI Studio provides an interactive console for designing and testing prompts against Gemini models, comparing responses, adjusting parameters such as temperature, and exporting the resulting code. It is exactly the rapid experimentation surface described, letting a developer validate prompt behavior in a browser before building the surrounding application.

Why this answer

Vertex AI Studio is the browser-based environment built for prompt design and rapid experimentation with Gemini models, including side-by-side comparison and code export. When the goal is to evaluate prompt behavior before committing to an application design, this console is the intended starting point rather than pipeline orchestration, end-user assistants, or a model catalog.

Exam trap

The trap here is conflating the model catalog with the prompt-design console, since both live under Vertex AI but only one offers interactive prompting and code export.

682
Multi-Selecthard

Which THREE of the following are potential risks when deploying generative AI?

Select 4 answers
A.Hallucinations
B.Memorization of sensitive training data
C.Bias and fairness issues
D.Increased model accuracy
E.Toxic or harmful content generation
AnswersA, B, C, E

Hallucinations are a genuine generative AI risk: the model generates fluent but factually incorrect or fabricated output because it predicts probable tokens rather than verifying truth. This directly satisfies the stem's requirement to identify a deployment risk, since unverified content can mislead users and damage trust.

Why this answer

Option A (Hallucinations) is correct because generative AI models can produce fluent but factually incorrect or fabricated outputs, which poses a risk when users trust them for decisions or information. Option B (Memorization of sensitive training data) is correct because models may reproduce verbatim or near-verbatim content from their training corpus, potentially leaking PII, proprietary data, or copyrighted material. Option C (Bias and fairness issues) is correct because models learn statistical patterns from training data and can amplify societal biases, leading to discriminatory or unfair outcomes in hiring, lending, or content moderation.

Option E (Toxic or harmful content generation) is correct because generative models can be prompted or inadvertently produce hate speech, harassment, self-harm instructions, or other harmful material. Option D (Increased model accuracy) is not a risk but a potential benefit, so it does not belong among the deployment risks.

Exam trap

Google Cloud often tests the distinction between risks and benefits, so the trap here is that candidates may mistakenly identify 'increased model accuracy' as a risk, when it is actually a performance improvement and not a deployment risk.

683
MCQeasy

What is the primary benefit of using foundation models (like Gemini) as opposed to training a model from scratch?

A.They are always faster at inference
B.They guarantee 100% accuracy on all tasks
C.They require less data and compute to adapt to new tasks
D.They are open-source and free to use
AnswerC

Foundation models arrive pre-trained on broad data, so adaptation via prompting, fine-tuning or retrieval needs only task-specific examples and modest compute. This directly satisfies the scenario's contrast with from-scratch training, which demands massive datasets and GPU time before any useful capability emerges.

Why this answer

Foundation models like Gemini are pre-trained on vast datasets, capturing general language understanding and patterns. Adapting them to new tasks via fine-tuning requires significantly less task-specific data and computational resources compared to training a model from scratch, which demands enormous datasets and compute for initial training. This transfer learning approach is the primary benefit, enabling efficient customization for specialized applications.

Exam trap

Google exams often test the misconception that 'pre-trained' means 'free' or 'always faster,' leading candidates to pick options that confuse inference speed or licensing with the core benefit of reduced data and compute for adaptation.

How to eliminate wrong answers

Option A is wrong because inference speed depends on model architecture, size, and optimization, not on whether the model is a foundation model or trained from scratch; a smaller custom model can be faster at inference than a large foundation model. Option B is wrong because no model, including foundation models, can guarantee 100% accuracy on all tasks due to inherent data biases, distribution shifts, and the complexity of real-world tasks. Option D is wrong because while some foundation models are open-source (e.g., Llama), many, including Gemini, are proprietary and require API access or licensing fees, so they are not universally free or open-source.

684
MCQmedium

A data scientist wants to run large-scale distributed training of a custom deep learning model using Google's custom AI accelerators. Which infrastructure should they choose to minimize cost while leveraging Google's proprietary chips?

A.Cloud TPU v5e
B.Compute Engine with NVIDIA A100 GPUs
C.Vertex AI Workbench with custom machines
D.Google Colab Pro with TPU runtime
AnswerA

Cloud TPU v5e is Google's cost-optimised Tensor Processing Unit generation, purpose-built for large-scale distributed training on Google's proprietary accelerators. It satisfies the stem's dual constraint of minimising cost while leveraging Google-designed chips, unlike GPU-based or general-purpose compute options.

Why this answer

Cloud TPU v5e is the correct choice because it is Google's proprietary custom AI accelerator designed specifically for large-scale distributed training of deep learning models, offering superior cost-efficiency compared to GPUs for many workloads. TPU v5e provides a balanced price-performance ratio for medium-to-large training tasks, and Google's TPU architecture is optimized for TensorFlow and JAX, enabling efficient scaling across multiple TPU pods. This minimizes cost while leveraging Google's custom chips, as opposed to using NVIDIA GPUs which are not Google's proprietary hardware.

Exam trap

The trap here is that candidates may confuse 'custom AI accelerators' with any high-performance hardware like GPUs, but the question specifically requires Google's proprietary chips (TPUs), and Cloud TPU v5e is the only option that directly provides cost-optimized, large-scale distributed training using Google's own accelerators.

How to eliminate wrong answers

Option B is wrong because Compute Engine with NVIDIA A100 GPUs uses third-party hardware (NVIDIA) rather than Google's proprietary chips, and while powerful, it typically incurs higher costs for large-scale distributed training compared to TPUs for suitable workloads. Option C is wrong because Vertex AI Workbench with custom machines is a development environment for building and training models, not a specific infrastructure choice for leveraging Google's custom AI accelerators; it can use TPUs or GPUs but does not inherently minimize cost with proprietary chips. Option D is wrong because Google Colab Pro with TPU runtime is designed for small-scale experimentation and prototyping, not for large-scale distributed training, and it lacks the scalability and cost efficiency of Cloud TPU v5e for production workloads.

685
MCQeasy

A retail chain's leadership wants to understand how generative AI could improve its customer service operations but is unsure where to start. They ask a cloud consultant to identify candidate use cases, estimate potential value, and flag which ones are realistically achievable with current technology. Which activity should the consultant perform first?

A.Fine-tune a foundation model on the retailer's historical customer service transcripts and measure its response quality.
B.Run a generative AI use case discovery workshop with business stakeholders to inventory candidate scenarios and prioritize them by value and feasibility.
C.Deploy a general-purpose chatbot on the public website and monitor engagement metrics for one quarter.
D.Migrate the retailer's contact center platform to a new SaaS vendor that advertises built-in generative AI features.
AnswerB

Discovery workshops bring business and technical stakeholders together to enumerate candidate scenarios, then rank them by expected business value and technical feasibility. This directly answers leadership's request for candidate use cases, value estimates, and realism checks, and it produces a prioritized shortlist that later pilots can draw from, which is the standard first move in a generative AI business strategy engagement.

Why this answer

The consultant should begin with structured use case discovery, because the organization needs an inventory of candidate scenarios ranked by business value and technical feasibility before committing to any build or procurement. This produces the prioritized shortlist that later pilots, model choices, and investment decisions can be based on, aligning generative AI work with measurable business outcomes.

Exam trap

The trap here is equating progress with building something, so the first instinct becomes fine-tuning a model or deploying a chatbot instead of first identifying and prioritizing which use cases deserve investment.

686
MCQmedium

A regional insurance company wants to let claims adjusters ask natural-language questions about 12 years of policy documents and claim histories stored in Cloud Storage. Leadership wants a working prototype in two weeks, minimal model-tuning effort, and answers that cite the exact source document. Which Google Cloud approach should the team implement?

A.Build a retrieval-augmented generation flow using Vertex AI Search to index the corpus and ground Gemini responses in retrieved passages.
B.Increase the Gemini model's context window to the maximum and paste the entire document corpus into every prompt.
C.Fine-tune a Gemini model on the historical claims corpus, then deploy the tuned endpoint behind an internal web app.
D.Train a custom text embedding model from scratch on the claims corpus, then serve similarity search from a self-managed cluster.
AnswerA

Vertex AI Search indexes the Cloud Storage corpus and returns relevant passages at query time, which Gemini then uses as grounding context. This delivers citable, source-linked answers without retraining, and the managed indexing pipeline keeps the prototype within a two-week window while honoring the citation requirement.

Why this answer

Retrieval-augmented generation pairs a managed search index with a foundation model, so answers are grounded in retrieved passages and can point back to the originating document. Vertex AI Search handles ingestion and indexing of the Cloud Storage corpus, while Gemini synthesizes the retrieved context into an answer. This avoids costly fine-tuning and oversized prompts while meeting the citation and speed requirements.

Exam trap

The trap here is assuming that fine-tuning or a larger context window can substitute for retrieval when the real requirement is traceable, up-to-date grounding in a large document corpus.

687
Multi-Selectmedium

A company wants to adopt GenAI for code generation and review. To ensure code quality and security, they plan to implement a change management program. Which THREE actions are most effective?

Select 3 answers
A.Create a code review checklist specifically for AI-generated code
B.Conduct training sessions on prompt engineering for code generation
C.Appoint a single AI champion to manage all code generation
D.Phase out senior developers to rely entirely on AI-generated code
E.Pilot the tool with a small group of developers before company-wide rollout
AnswersA, B, E

AI-generated code can contain hallucinated dependencies, insecure patterns and licence issues, so a dedicated checklist enforces consistent scrutiny of those specific failure modes. This satisfies the stem's code quality and security constraint by embedding verification into the existing review gate.

Why this answer

Option A is correct because a dedicated code review checklist for AI-generated code ensures reviewers systematically check for common LLM pitfalls such as hallucinated APIs, insecure patterns, license issues, and subtle logic errors, directly supporting code quality and security. Option B is correct because training on prompt engineering helps developers craft precise prompts, provide context, and constrain outputs, which improves the relevance, correctness, and security of generated code and reduces rework. Option E is correct because a limited pilot with a small developer group allows the organization to validate tooling, measure quality and security impacts, refine review processes, and gather feedback before a company-wide rollout, lowering risk.

Option C is not appropriate because concentrating all responsibility in a single AI champion creates a bottleneck and a single point of failure rather than a scalable change management program. Option D is incorrect because phasing out senior developers removes the human expertise needed to review, validate, and secure AI-generated code, which would undermine quality and security rather than improve them.

Exam trap

The trap is selecting options that sound efficient (single AI champion, full AI reliance) but actually undermine quality and security — the exam rewards balanced, risk-aware change management over speed-focused shortcuts.

688
MCQeasy

Which technique allows a model to incorporate real-time data from external APIs?

A.RAG with tool calling
B.Prompt engineering
C.Fine-tuning
D.Model pruning
AnswerA

RAG with tool calling lets the model invoke external APIs at inference time, retrieving live data rather than relying solely on static training weights. This satisfies the stem's real-time external API constraint, which plain retrieval-augmented generation alone cannot meet.

Why this answer

RAG with tool calling is correct because it enables a generative AI model to query external APIs in real-time, retrieve up-to-date information, and incorporate that data into its response. This technique combines retrieval-augmented generation (RAG) with function calling, where the model outputs a structured request (e.g., a JSON object) to invoke an API, receive the result, and then generate a context-aware answer. Unlike static methods, this allows dynamic data integration without retraining.

Exam trap

A common pitfall is thinking that prompt engineering alone can achieve real-time data integration with external APIs. However, in Google Cloud's generative AI context, only RAG with tool calling (function calling) provides the explicit mechanism to execute API calls and incorporate live results into model responses.

How to eliminate wrong answers

Option B is wrong because prompt engineering only modifies the input text to guide model behavior, but it cannot fetch live data from external sources—it relies solely on the model's pre-existing knowledge. Option C is wrong because fine-tuning updates the model's weights on a fixed dataset, which does not enable real-time API access; it only improves performance on static tasks. Option D is wrong because model pruning reduces model size by removing redundant weights, which has no mechanism for external data retrieval or API interaction.

689
Multi-Selectmedium

Which TWO strategies are effective for reducing latency in a generative AI chat application deployed on Vertex AI? (Select 2)

Select 2 answers
A.Deploy on TPU instead of GPU
B.Use streaming responses
C.Increase the max output tokens
D.Enable model quantization
E.Use larger batch sizes
AnswersB, D

Reduces perceived latency.

Why this answer

Streaming responses reduce perceived latency by sending tokens to the client as they are generated, rather than waiting for the full response. This leverages server-sent events (SSE) or chunked transfer encoding to deliver partial results immediately, improving user experience in chat applications.

Exam trap

Google Cloud often tests the distinction between reducing actual latency (e.g., model optimization) versus reducing perceived latency (e.g., streaming), and candidates mistakenly choose options that increase throughput (like larger batch sizes) without realizing they harm per-request latency.

690
MCQeasy

A startup is building a customer support chatbot using Vertex AI and wants to ground responses in their product documentation to reduce hallucinations. Which approach should they use?

A.Enable Vertex AI Grounding with a custom enterprise data store containing the documentation.
B.Use the Codey API for text generation.
C.Use the base model without any grounding to maximize flexibility.
D.Fine-tune the model on the documentation and deploy.
AnswerA

Grounding with a custom enterprise data store retrieves passages from the product documentation and injects them into the prompt, so responses cite actual content rather than relying on parametric memory. This directly addresses the hallucination constraint in the stem.

Why this answer

Vertex AI Grounding with a custom enterprise data store is the correct approach because it allows the chatbot to retrieve and cite specific chunks from the product documentation in real time, directly reducing hallucinations by constraining responses to verified content. This method uses the underlying grounding service to query a vector-based data store (powered by Vertex AI Search) and append source references to the model's output, ensuring factual accuracy without retraining.

Exam trap

Google Cloud often tests the misconception that fine-tuning is the best way to incorporate domain knowledge, but the trap here is that fine-tuning does not provide dynamic, verifiable grounding with citations, whereas Vertex AI Grounding with a custom data store does, making it the correct choice for reducing hallucinations in a retrieval-augmented generation use case.

How to eliminate wrong answers

Option B is wrong because the Codey API is designed for code generation tasks (e.g., code completion, chat), not for grounding responses in external documents; it lacks the retrieval-augmented generation (RAG) capabilities needed to reduce hallucinations from product documentation. Option C is wrong because using a base model without grounding maximizes flexibility but also maximizes the risk of hallucination, as the model relies solely on its training data and cannot verify facts against the documentation. Option D is wrong because fine-tuning the model on the documentation embeds the content into the model's weights, which is static, costly to update, and does not provide real-time citation or retrieval; it also risks overfitting and does not leverage Vertex AI's built-in grounding infrastructure for dynamic fact-checking.

691
Multi-Selecthard

A financial institution wants to deploy a generative AI system for automated report generation. They require that the model does NOT expose sensitive information from its training data and that outputs are factually accurate. Which THREE techniques should they combine?

Select 3 answers
A.Use Retrieval-Augmented Generation (RAG) to ground outputs in verified documents
B.Train the model with differential privacy
C.Use a larger model with more parameters
D.Apply Reinforcement Learning from Human Feedback (RLHF) to align model behavior
E.Increase the temperature to 2.0 for creativity
AnswersA, B, D

RAG grounds generation in retrieved, verified documents rather than parametric memory, directly satisfying the factual-accuracy constraint. Because the model conditions on supplied source text instead of recalling training data, sensitive information from training is less likely to surface. It does not, however, prevent leakage from the underlying model itself.

Why this answer

Option A is correct because Retrieval-Augmented Generation (RAG) grounds the model's outputs in an external, verified document store at inference time, so factual claims are anchored to authoritative sources rather than the model's parametric memory, directly improving factual accuracy. Option B is correct because training with differential privacy (e.g., DP-SGD) adds calibrated noise to gradients and bounds per-example influence, providing a formal guarantee that the model cannot memorize and expose sensitive training-data details. Option D is correct because Reinforcement Learning from Human Feedback (RLHF) fine-tunes the model against human preference and safety signals, aligning its behavior to refuse leakage and to prefer truthful, accurate responses.

Option C does not belong because simply scaling parameters increases capacity and can worsen memorization of training data without any privacy or factuality guarantee. Option E does not belong because raising temperature to 2.0 increases randomness and hallucination, degrading factual accuracy rather than improving it.

692
MCQeasy

A small marketing agency wants its staff to draft blog posts, summarize meeting notes, and brainstorm campaign ideas using a conversational assistant. The agency has no cloud engineering team and prefers a ready-to-use product with enterprise-grade data protections rather than building anything. Which Google Cloud offering best matches this need?

A.Vertex AI Model Garden
B.Vertex AI Agent Builder
C.Google Kubernetes Engine
D.Gemini for Google Workspace
AnswerD

Gemini for Google Workspace embeds generative AI assistance directly into Gmail, Docs, Sheets, Meet, and related apps, so non-technical staff can draft, summarize, and brainstorm where their content already lives. It is a ready-to-use product with enterprise controls and requires no engineering effort, matching an agency without a cloud team. This is the most direct fit for everyday productivity tasks.

Why this answer

Gemini for Google Workspace delivers generative AI assistance inside the productivity apps the agency already uses, letting staff draft, summarize, and brainstorm without any development work. The remaining offerings are developer platforms or infrastructure for building and serving models or applications, which conflict with the requirement for a ready-to-use product and the absence of an engineering team.

Exam trap

The trap here is equating 'generative AI on Google Cloud' with Vertex AI, overlooking that Gemini for Google Workspace is the packaged assistant for everyday productivity users.

693
MCQeasy

A small marketing agency wants to add a Google Cloud generative AI assistant that can summarize campaign documents stored in Google Drive and answer follow-up questions about them, with minimal development effort. Which Google Cloud offering should they choose?

A.Gemini for Google Workspace
B.Vertex AI Model Garden
C.Cloud Vision API
D.Google Cloud Speech-to-Text
AnswerA

Gemini for Google Workspace embeds generative AI directly in Docs, Gmail, Drive, and related apps, so the agency can summarize Drive documents and ask follow-up questions without building anything. It is the lowest-effort fit for teams already working inside Workspace and needing document-grounded assistance.

Why this answer

Gemini for Google Workspace is the right choice because it brings generative AI into the productivity apps where the agency already stores and edits campaign documents. It can summarize Drive files and support conversational follow-up without custom engineering. The other services are building-block APIs or model catalogs that would require substantial integration work before delivering the same outcome.

Exam trap

The trap here is assuming that any Google Cloud AI service can summarize documents, when only the Workspace-embedded Gemini offering provides that turnkey experience.

694
MCQeasy

A prompt engineer wants to improve the model's adherence to a specific output format (e.g., always start with a greeting). Which technique should they try first?

A.Use a lower temperature to make the output more deterministic.
B.Fine-tune the model on many examples of the desired format.
C.Include a system instruction at the beginning of the prompt that specifies the desired format.
D.Modify the model's tokenizer to encode the format rules.
AnswerC

System instructions are processed as high-priority context that conditions every response, making them the most direct control for enforcing a consistent output format such as an opening greeting. This precedes few-shot examples or post-processing in effort and reliability.

Why this answer

System instructions are the most direct and efficient method to enforce output formatting in large language models. By placing a clear directive at the beginning of the prompt (e.g., 'Always start your response with a greeting'), the model's attention mechanism is guided to prioritize this rule during generation, without requiring retraining or hyperparameter changes.

Exam trap

Google Cloud often tests the misconception that hyperparameter tuning (like temperature) can enforce structural output rules, when in fact it only controls randomness, not format adherence.

How to eliminate wrong answers

Option A is wrong because lowering temperature reduces randomness but does not enforce a specific structural rule like starting with a greeting; it only makes token selection more deterministic, which could still produce varied formats. Option B is wrong because fine-tuning is a resource-intensive process that requires a curated dataset and retraining, making it an overkill for a simple formatting constraint that can be achieved with a prompt instruction. Option D is wrong because modifying the tokenizer would alter how input text is split into tokens, not how the model adheres to output format rules; tokenizers have no mechanism to enforce generation constraints.

695
MCQhard

A team needs to generate photorealistic product imagery for an e-commerce catalog from text descriptions, and later needs to edit existing photos by removing unwanted objects while preserving the rest of the scene. Which Google Cloud generative media capabilities should they use for these two tasks, respectively?

A.Gemini for text-to-image generation, and BigQuery for object removal
B.Vertex AI Vision for text-to-image generation, and Imagen for object removal
C.Imagen for text-to-image generation, and a custom-trained object detection model for removal
D.Imagen for text-to-image generation, and Imagen editing capabilities for object removal and inpainting
AnswerD

Imagen generates high-quality images from text prompts, which covers creating catalog imagery from product descriptions. Its editing capabilities, including inpainting and mask-based modification, allow unwanted objects to be removed while the surrounding scene is preserved. Using one family for both generation and editing keeps the workflow consistent and satisfies both stated tasks.

Why this answer

Text-to-image generation and masked image editing are distinct capabilities, and both are provided within the same generative media family. Using it for generation from product descriptions and for inpainting-based object removal keeps quality and tooling consistent. Analytics platforms, video pipelines and detection-only models each cover adjacent ground but cannot fulfill the complete pair of tasks described.

Exam trap

The trap here is assuming that a model which can detect or analyze objects in an image can also remove them, when removal requires generative inpainting that reconstructs the background.

696
MCQeasy

Refer to the exhibit. A developer sees this error when trying to call a Vertex AI endpoint for online prediction. What permission does the requesting identity need to be granted?

A.aiplatform.prediction.predict
B.aiplatform.endpoints.predict
C.aiplatform.endpoints.use
D.aiplatform.models.predict
AnswerB

Granting `aiplatform.endpoints.predict` satisfies the online prediction constraint: it authorises the requesting identity to send prediction requests to a Vertex AI endpoint. This permission is contained in roles such as Vertex AI User, and is the minimum required for calling `predict` on a deployed endpoint.

Why this answer

The error occurs when calling a Vertex AI endpoint for online prediction, which requires the `aiplatform.endpoints.predict` permission. This permission is specifically scoped to the endpoint resource, allowing the identity to send prediction requests to a deployed model endpoint. The correct IAM role binding must include this permission for the requesting identity to successfully invoke the endpoint.

Exam trap

Google Cloud often tests the distinction between permissions scoped to endpoints versus models, and candidates mistakenly choose `aiplatform.models.predict` because they think prediction is always tied to the model, not the endpoint serving it.

How to eliminate wrong answers

Option A is wrong because `aiplatform.prediction.predict` is not a valid IAM permission in Vertex AI; the correct permission for prediction is scoped to the endpoint or model resource, not a generic 'prediction' service. Option C is wrong because `aiplatform.endpoints.use` does not exist as a permission; Vertex AI uses `aiplatform.endpoints.predict` for invoking predictions on endpoints. Option D is wrong because `aiplatform.models.predict` is a permission for calling prediction directly on a model resource, not on an endpoint, and the error specifically references an endpoint call, not a model call.

697
MCQeasy

A retail company wants to build an assistant that can answer customer questions about its product catalog, but the catalog changes daily. The team wants the model to reference the latest product information without retraining. Which approach should they use?

A.Increase the model's temperature to encourage more creative answers.
B.Fine-tune a foundation model on the product catalog each day.
C.Use retrieval-augmented generation (RAG) to fetch relevant product information at inference time.
D.Deploy the model on a larger machine type with more memory.
AnswerC

RAG retrieves up-to-date documents from an external source and provides them to the model as context, so the assistant can answer based on the latest catalog without retraining. This is ideal when data changes frequently. It also reduces hallucinations because the model grounds its response in retrieved facts.

Why this answer

Retrieval-augmented generation combines a retrieval system with a generative model, allowing the assistant to fetch the latest product details from a database or document store and include them in the prompt. This keeps answers accurate and current without the cost and delay of retraining. It also provides traceability, as the retrieved sources can be cited.

Exam trap

The trap here is assuming that fine-tuning is the only way to inject domain knowledge, when retrieval-augmented generation is specifically designed for frequently changing data.

698
MCQmedium

A healthcare company is using Vertex AI to build a generative AI assistant that helps doctors draft clinical notes. The assistant uses a fine-tuned PaLM 2 model deployed on a private endpoint. Recently, doctors have reported that the assistant takes over 30 seconds to respond, causing workflow delays. Additionally, the monthly Vertex AI costs have increased by 40% without a proportional increase in usage. The model responses are generally accurate but sometimes include irrelevant details. The company wants to improve response time and cost while maintaining acceptable quality. A review of logs shows that most requests are for similar note types (e.g., progress notes, discharge summaries) and that the same prompt is used repeatedly with minor variations. What should the company do first?

A.Switch to a larger model (e.g., Gemini 1.5 Pro) to improve response quality and reduce irrelevant details
B.Increase the Vertex AI endpoint's maximum request quota to handle concurrent requests
C.Apply model quantization (e.g., INT8) to reduce model size and inference time
D.Implement response caching for common queries and batch process similar requests
AnswerD

Caching repeated prompts and batching similar note requests cuts redundant model invocations, directly addressing the 30-second latency and 40% cost increase while preserving output quality. The logs confirm most requests share prompt patterns, making this the highest-impact first step.

Why this answer

The logs show that most requests are for similar note types with repeated prompts, making response caching ideal for reducing latency and cost. Caching stores responses for identical or near-identical queries, eliminating redundant inference calls, which directly addresses the 30-second response time and 40% cost increase without sacrificing quality.

Exam trap

The trap here is that candidates confuse performance optimization techniques (quantization, quota increases) with the root cause of redundant requests, leading them to pick options that address symptoms rather than the fundamental pattern of repeated prompts.

How to eliminate wrong answers

Option A is wrong because switching to a larger model (e.g., Gemini 1.5 Pro) would increase inference time and cost, worsening the problem, and the issue is not about quality but latency and cost. Option B is wrong because increasing the endpoint's maximum request quota does not reduce per-request latency or cost; it only allows more concurrent requests, which could further degrade performance under load. Option C is wrong because model quantization (e.g., INT8) reduces model size and inference time but requires retraining or calibration, and it may degrade accuracy for clinical notes where precision is critical; moreover, it does not address the root cause of redundant requests.

699
MCQmedium

A startup is building a generative AI tool that helps users write code. They want to launch quickly but need to ensure the generated code is secure and does not introduce vulnerabilities. They have a small team of developers with some ML experience. The tool should be cloud-hosted. Which approach balances speed, security, and cost?

A.Deploy the tool without any security checks and rely on manual review
B.Train a custom code generation model from scratch on a large dataset
C.Use a pre-trained code model (e.g., Codey) and add a security filtering layer
D.Use a smaller model and restrict outputs to only simple code patterns
AnswerC

A pre-trained code model removes the cost and delay of training from scratch, letting a small ML team launch quickly, while the added security filtering layer scans generated code for vulnerabilities. This satisfies the need to balance speed, security and cost.

Why this answer

Using a pre-trained code model such as Codey provides immediate capability without the cost and time of training from scratch, while adding a security filtering layer (SAST-style scanning, output sanitization, policy checks) addresses the security requirement. This balances speed, security, and cost for a small team.

Exam trap

The trap is assuming that a pre-trained model alone is secure, or that training from scratch yields better security — in reality, guardrails around a pre-trained model deliver the best speed/security/cost balance.

How to eliminate wrong answers

Option A is wrong because relying solely on manual review does not scale and defeats the purpose of automated generation — security vulnerabilities would slip through. Option B is wrong because training a custom code model from scratch requires massive datasets, GPU compute, and ML expertise the small team lacks, blowing the time and cost budget. Option D is wrong because restricting outputs to simple patterns severely limits usefulness and does not guarantee security — simple code can still contain injection flaws.

700
MCQhard

Refer to the exhibit. This is the IAM policy for a project containing a Vertex AI Agent Builder agent and a data store. The agent is unable to access the data store. What is the most likely cause?

A.The user needs more permissions
B.The agent needs a bigger quota
C.The agent service account needs the data store viewer role
D.The data store is not in the same region
AnswerC

The agent's service account lacks the data store viewer role, so its retrieval calls are denied. Granting that role on the data store satisfies the stem's access requirement, since the agent must read the store to ground responses.

Why this answer

The agent service account must have the Data Store Viewer role (or equivalent permissions) to read data from the data store. Without this role, the agent cannot access the indexed content, even if the user has permissions. This is a common IAM misconfiguration in Vertex AI Agent Builder.

Exam trap

A common trap in Google Cloud IAM is confusing user permissions with service account permissions. The agent uses a service account, not the user's credentials, so the data store viewer role must be granted to the service account.

How to eliminate wrong answers

Option A is wrong because the user's permissions are irrelevant; the agent operates under its own service account identity, not the user's. Option B is wrong because quota limits affect throughput or resource usage, not access control; the issue is authorization, not capacity. Option D is wrong because Vertex AI Agent Builder and data stores can be in different regions; cross-region access is supported and not a typical cause of access failures.

701
MCQmedium

A data analyst wants to build a regression model in BigQuery to predict sales from historical data without writing any Python code. Which BigQuery ML statement should they use to define the model?

A.CREATE MODEL my_model OPTIONS(model_type='linear_reg') AS SELECT ...
B.INSERT INTO model my_model VALUES ...
C.CREATE ML my_model AS (SELECT ...)
D.SELECT ML.TRAIN('linear_reg', ...)
AnswerA

`CREATE MODEL ... OPTIONS(model_type='linear_reg')` trains a linear regression directly in BigQuery using SQL, satisfying the no-Python constraint. BigQuery ML supports regression natively, so the analyst defines and fits the model without exporting data or provisioning external tooling.

Why this answer

BigQuery ML uses the standard SQL DDL statement CREATE MODEL with an OPTIONS clause specifying model_type (e.g., 'linear_reg') and a training query supplied via AS SELECT. This lets analysts train and deploy ML models entirely in SQL without exporting data or writing Python. The linear_reg model type trains a linear regression on the selected features and label column.

Exam trap

The trap here is confusing BigQuery ML's declarative CREATE MODEL DDL with procedural ML APIs like ML.TRAIN or Python-style fit/predict calls, causing candidates to pick a plausible-sounding but nonexistent statement.

How to eliminate wrong answers

Option B is wrong because INSERT INTO is a DML statement for adding rows to a table, not for defining or training a model — BigQuery ML has no 'INSERT INTO model' syntax. Option C is wrong because 'CREATE ML' is not valid BigQuery SQL; the correct keyword is CREATE MODEL, and the training data is provided with AS SELECT, not wrapped in parentheses after AS. Option D is wrong because ML.TRAIN is not a BigQuery ML function — training is done declaratively via CREATE MODEL, while ML.PREDICT and ML.EVALUATE are the functions used after training.

702
MCQmedium

A team is deploying a large language model for legal document summarization. They find the model occasionally omits critical legal clauses. Which improvement technique would be most effective?

A.Design a prompt that explicitly lists required sections
B.Increase the top_p value to 1.0
C.Fine-tune the model on legal summaries
D.Lower the temperature to 0.1
AnswerA

Explicitly listing the required sections in the prompt constrains the model's output structure, directing attention to mandatory clauses that free-form summarisation tends to drop. This prompt-engineering approach improves recall of critical content without retraining or architectural changes.

Why this answer

The most effective technique is to design a prompt that explicitly lists the required sections. This is a form of prompt engineering that directly addresses the omission issue by instructing the model to include all specified clauses, leveraging the model's instruction-following capability. It is immediate, low-cost, and does not require retraining or altering sampling parameters.

By enumerating the sections, the model is guided to cover each one, reducing the chance of missing critical content.

Exam trap

The trap here is confusing sampling parameters (temperature, top_p) with techniques that ensure completeness; candidates might think lowering temperature or increasing top_p would make the model more reliable, but they only affect randomness, not coverage of required content.

How to eliminate wrong answers

Option B is wrong because increasing top_p to 1.0 makes the sampling more diverse (nucleus sampling considers all tokens), which can increase randomness and does not ensure completeness; it may even worsen omissions. Option C is wrong because fine-tuning on legal summaries could improve domain adaptation but is expensive, time-consuming, and does not guarantee that all required sections will be included; it may also overfit to the training summaries' style rather than ensuring clause coverage. Option D is wrong because lowering temperature to 0.1 makes the output more deterministic and focused on high-probability tokens, but it does not enforce inclusion of specific sections; it can actually reduce creativity and may cause the model to stick to common patterns, potentially omitting less frequent but critical clauses.

703
MCQmedium

A developer is building a customer support chatbot using a large language model. The chatbot frequently generates plausible-sounding but incorrect answers to product questions. Which technique should be applied to improve factual accuracy?

A.Provide a few-shot example of correct answers in the prompt.
B.Use a higher temperature setting to encourage more creative responses.
C.Increase the model's context length to include more of the conversation history.
D.Enable Grounding with the company's product knowledge base.
AnswerD

Grounding retrieves authoritative product content and injects it into the prompt context, so the model conditions its response on real documentation rather than parametric memory alone. This directly addresses the hallucination constraint in the stem, replacing plausible fabrication with verifiable, source-anchored answers drawn from the company knowledge base.

Why this answer

Grounding with the company's product knowledge base (D) is the most effective technique to improve factual accuracy. Grounding involves retrieving relevant information from a trusted knowledge source and providing it to the model as context, so the model generates answers based on that data rather than relying solely on its training. This reduces hallucinations and ensures responses are accurate and up-to-date.

Exam trap

Generative AI Leader often tests the difference between prompt engineering techniques and grounding, causing candidates to choose few-shot examples or context length increases when the core issue is lack of factual knowledge.

How to eliminate wrong answers

Option A is wrong because few-shot examples can guide the model's style but do not provide the factual knowledge needed to answer product-specific questions accurately. Option B is wrong because a higher temperature increases randomness and creativity, which would likely worsen factual accuracy. Option C is wrong because increasing context length allows more conversation history but does not inject external factual knowledge; it may even introduce irrelevant information.

704
MCQmedium

A media company wants to create a custom AI assistant that can answer employee questions by referencing internal policy documents stored in Google Drive. They need a low-code solution that allows them to configure the assistant, connect data sources, and deploy it for internal use. Which Google Cloud offering should they use?

A.Vertex AI Agent Builder
B.Document AI
C.Gemini for Google Workspace
D.Vertex AI Search
AnswerA

Vertex AI Agent Builder is a low-code platform designed to help developers and business users create generative AI agents and applications. It provides tools to connect to data sources like Google Drive, configure conversational flows, and deploy agents. This matches the requirement for a low-code solution to build an internal AI assistant that references policy documents.

Why this answer

Vertex AI Agent Builder is the only Google Cloud offering that provides a low-code environment specifically for building generative AI agents and applications. It allows integration with Google Drive and other data sources, enabling the assistant to retrieve and use internal policy documents to answer employee questions. The other services are either search-focused, productivity-focused, or document-processing-focused, and lack the agent orchestration and deployment features required.

Exam trap

The trap here is confusing Vertex AI Search with Vertex AI Agent Builder, as both can work with enterprise data but only Agent Builder is designed for creating conversational agents.

705
MCQhard

A healthcare startup is exploring GenAI for clinical note summarization. They have concerns about patient data privacy. Which Google Cloud approach best addresses privacy while still using powerful models?

A.Deploy open-source models on-premises
B.Use a third-party API with anonymization of patient data
C.Use Vertex AI with model customization (fine-tuning)
D.Use Vertex AI with data residency controls and no external data sharing
AnswerD

Vertex AI keeps prompts and patient data within the selected region, and Google contractually excludes them from model training, satisfying the residency and confidentiality constraints. Data residency controls pin storage and processing to approved locations, while no external sharing prevents third-party exposure — meeting healthcare privacy requirements without sacrificing access to powerful models.

Why this answer

Vertex AI with data residency controls and no external data sharing ensures that patient data remains within specified geographic boundaries and is not used for model training or improvement, directly addressing healthcare privacy regulations like HIPAA. This approach leverages Google Cloud's powerful models while maintaining strict data governance, unlike options that risk data exposure or lack enterprise-grade controls.

Exam trap

The trap here is that candidates often assume fine-tuning (Option C) inherently provides privacy, but without explicit data residency and no-sharing policies, it fails to meet strict healthcare compliance requirements.

How to eliminate wrong answers

Option A is wrong because deploying open-source models on-premises, while offering data control, often lacks the advanced summarization capabilities and scalability of Vertex AI's foundation models, and still requires significant effort to ensure HIPAA compliance without Google's built-in privacy safeguards. Option B is wrong because using a third-party API, even with anonymization, introduces risks of data leakage or re-identification, and typically does not provide contractual guarantees against model training on patient data, violating many healthcare privacy policies. Option C is wrong because fine-tuning a model on Vertex AI without explicit data residency controls and no external data sharing may still allow Google to process data outside desired regions or use it for service improvements, failing to meet strict data privacy requirements.

706
MCQmedium

A team monitors their generative AI model on Vertex AI. They notice output quality declining. Which metric is most likely the root cause?

A.Input token count per request is increasing.
B.Output token count is decreasing.
C.Prediction latency is stable.
D.Error rate is less than 1%.
AnswerA

Rising input token counts lengthen prompts, pushing the model beyond its effective context and diluting instruction focus, which degrades output quality. This metric directly explains the decline, unlike latency or cost signals that do not affect generation fidelity.

Why this answer

A is correct because an increasing input token count per request can degrade output quality by diluting the model's attention across a longer context window. In transformer-based models like those on Vertex AI, the attention mechanism has a fixed capacity; as input tokens grow, the model may lose focus on critical information, leading to less coherent or relevant outputs. This is a common issue in production systems where users gradually add more context without trimming irrelevant tokens.

Exam trap

A common misconception is that output quality issues are always due to model errors or latency problems, rather than subtle input-side factors like token count inflation that silently degrade attention focus. This question highlights that root cause can be on the input side.

How to eliminate wrong answers

Option B is wrong because a decreasing output token count does not inherently cause quality decline; it may indicate shorter responses, but quality can remain high if the model is well-tuned. Option C is wrong because stable prediction latency suggests consistent infrastructure performance, not a root cause of output quality degradation. Option D is wrong because a low error rate (<1%) indicates the model is responding without failures, but output quality can still suffer from issues like hallucination or incoherence even when error rates are minimal.

707
Multi-Selecthard

A financial institution is deploying a generative AI solution that generates investment advice. They must ensure fairness, avoid toxic outputs, and comply with regulations like GDPR. Which TWO strategies should they implement? (Choose two.)

Select 2 answers
A.Use Vertex AI Safety Attributes to filter harmful content in both input and output.
B.Set the model temperature to 0 to eliminate creativity and reduce bias.
C.Implement a human review process for any advice above a certain risk threshold.
D.Fine-tune the model exclusively on compliant financial documents.
E.Disable request logging to avoid storing sensitive data.
AnswersA, C

Vertex AI Safety Attributes provides built-in safety filters that detect and block harmful content (e.g., hate speech, toxicity, financial misinformation) in both user prompts and model outputs, directly addressing the need to avoid toxic outputs and comply with regulations like GDPR.

Why this answer

Vertex AI Safety Attributes provides built-in safety filters that can detect and block harmful content (e.g., hate speech, toxicity, financial misinformation) in both user prompts and model outputs. This directly addresses the need to avoid toxic outputs and comply with regulations like GDPR, which require protecting users from harmful or biased advice.

Exam trap

A common misconception is that reducing model temperature or fine-tuning on compliant data alone can ensure safety and regulatory compliance, when in fact these measures do not address dynamic, context-dependent toxic outputs or logging requirements.

708
MCQmedium

A content generation model for e-commerce product descriptions repeats the same phrases across multiple descriptions (e.g., 'high-quality', 'best-in-class'). The team wants more varied and engaging output. Which parameter adjustment is most appropriate?

A.Increase the frequency penalty parameter to 1.0.
B.Decrease the max output tokens to 50.
C.Increase the temperature parameter to 1.5.
D.Set the top-p value to a very small number like 0.1.
AnswerA

Raising frequency penalty to 1.0 penalises tokens proportionally to how often they have already appeared in the generated text, directly discouraging the repeated phrases the stem describes. This satisfies the requirement for varied, engaging product descriptions without altering the model's underlying knowledge or the prompt itself.

Why this answer

Increasing the frequency penalty to 1.0 penalizes tokens that have already appeared in the generated text, directly reducing repetition of phrases like 'high-quality' and 'best-in-class'. This encourages the model to use more diverse vocabulary and sentence structures, leading to varied and engaging product descriptions.

Exam trap

The Generative AI Leader exam often tests the distinction between frequency penalty and temperature, where candidates mistakenly increase temperature to add variety, not realizing that temperature increases randomness and can break coherence, while frequency penalty directly targets repetition without sacrificing quality.

How to eliminate wrong answers

Option B is wrong because decreasing max output tokens to 50 limits the length of each description but does not address the root cause of phrase repetition; the model can still repeat phrases within the shorter output. Option C is wrong because increasing temperature to 1.5 makes the output more random and less coherent, which can lead to nonsensical descriptions rather than controlled variation. Option D is wrong because setting top-p to a very small number like 0.1 restricts the model to only the most likely tokens, which actually increases repetition and reduces diversity, the opposite of the desired outcome.

709
MCQmedium

A product team at a retail company is using a foundation model on Vertex AI to generate short marketing taglines for new products. They find the outputs are often too long and sometimes include extra commentary. They want to constrain the model to produce only a single concise tagline. Which parameter should they adjust?

A.Temperature
B.Top-p
C.Max output tokens
D.Top-k
AnswerC

Max output tokens sets a hard limit on the number of tokens the model can generate in its response. By setting this to a small value that accommodates a single tagline, the team can prevent lengthy outputs and force the model to be concise. It directly addresses the length issue described.

Why this answer

The max output tokens parameter caps the total number of tokens generated, which directly enforces a limit on response length. By setting it appropriately, the team can ensure the model produces only a short tagline and cannot continue with extra commentary. Other parameters like temperature, top-k, and top-p affect randomness and diversity, not length.

Exam trap

The trap here is confusing parameters that control randomness with those that control output length.

710
Multi-Selecteasy

A data scientist is using Vertex AI's Generative AI Studio to experiment with prompt designs. Which THREE features are available in the studio?

Select 3 answers
A.Grounding configuration
B.Model parameter adjustments (temperature, top_p, etc.)
C.Automated hyperparameter tuning
D.Prompt templates
E.A/B testing of multiple prompt versions
AnswersA, B, D

Grounding configuration lets users connect prompts to external data sources such as Vertex AI Search, improving factual accuracy. It is a native Generative AI Studio feature, satisfying the requirement to identify capabilities available when experimenting with prompt designs.

Why this answer

In Vertex AI's Generative AI Studio, grounding configuration (A) is available so you can ground model responses in specific data sources such as Vertex AI Search or your own datasets, reducing hallucinations and improving factual accuracy. Model parameter adjustments (B) are also provided, letting you tune values like temperature, top_p, top_k, and max output tokens directly in the studio to control response creativity and length. Prompt templates (D) are included as a feature, offering pre-built and customizable prompt structures that help you quickly design and reuse effective prompts.

Automated hyperparameter tuning (C) belongs to Vertex AI training services like Vertex AI Vizier or custom training jobs, not the prompt-design interface of Generative AI Studio. A/B testing of multiple prompt versions (E) is not a built-in feature of Generative AI Studio; comparing prompt variants typically requires external evaluation tooling or Vertex AI evaluation services rather than a native A/B testing option in the studio.

Exam trap

Google Cloud often tests the distinction between features available in Generative AI Studio (prompt design, model parameters, grounding, templates) versus those in other Vertex AI services (e.g., Vertex AI Training for hyperparameter tuning, Vertex AI Experiments for A/B testing). Candidates mistakenly assume all ML workflow features are present in the studio.

711
MCQhard

A company deployed a large language model on Vertex AI using the configuration shown in the exhibit. During peak usage, users report high latency. Which change is most likely to improve latency?

A.Remove the accelerator to simplify deployment.
B.Increase minReplicaCount to 3.
C.Switch to a GPU with more memory, such as NVIDIA_TESLA_A100.
D.Change machineType to n1-standard-4 to reduce cost.
AnswerB

Vertex AI scales replicas between minReplicaCount and maxReplicaCount; raising the minimum to 3 keeps additional instances warm, so peak traffic is spread across more replicas and per-request latency falls. Cold-start delays from scaling up are avoided.

Why this answer

Increasing minReplicaCount to 3 ensures that at least three instances of the model are always running and ready to serve requests. This reduces cold-start latency and distributes the load across multiple replicas, directly addressing high latency during peak usage by providing more concurrent serving capacity.

Exam trap

The Generative AI Leader exam often tests the misconception that upgrading hardware (GPU memory or type) is the primary fix for latency, when in fact scaling out replicas is the more direct solution for handling concurrent request load.

How to eliminate wrong answers

Option A is wrong because removing the accelerator (GPU/TPU) would force the model to run on CPU, drastically increasing inference latency, especially for large language models. Option C is wrong because switching to a GPU with more memory (NVIDIA_TESLA_A100) does not directly improve latency; it addresses memory capacity issues, not the throughput bottleneck caused by insufficient replicas. Option D is wrong because changing machineType to n1-standard-4 reduces CPU and memory resources, which would likely increase latency rather than improve it, and cost reduction is not the goal here.

712
MCQmedium

An e-commerce company is using a generative AI model to write product descriptions. They want to ensure that the model does not generate harmful content such as hate speech or violence. Which Google Cloud feature should they configure?

A.Cloud DLP (Data Loss Prevention)
B.Vertex AI Explainable AI
C.Vertex AI Safety Filters
D.Vertex AI Model Monitoring
AnswerC

Vertex AI Safety Filters apply configurable thresholds that block harmful categories—hate speech, violence, harassment—directly on model input and output, satisfying the requirement to prevent harmful product descriptions. Configuring these filters enforces content policy at inference time, which is precisely the constraint the stem demands.

Why this answer

Google Cloud's Vertex AI provides built-in safety filters that can be configured to block harmful content categories like hate speech, violence, sexual content, and dangerous instructions.

713
MCQmedium

A financial services firm needs to generate synthetic data for training models while ensuring that no real customer data leaks. Which technique should they use?

A.Using the Vertex AI PII redaction service
B.Using a public foundation model without fine-tuning
C.Data masking before training
D.Differential privacy during fine-tuning
AnswerD

Differential privacy adds calibrated noise during fine-tuning, bounding any single customer record's influence on the model, so synthetic outputs cannot reveal real individuals. This satisfies the stem's constraint that no actual customer data leaks while still generating usable training data.

Why this answer

Differential privacy during fine-tuning is the correct technique because it adds calibrated noise to the training process, ensuring that the synthetic data generated does not reveal information about any individual real customer record. This approach provides a formal mathematical guarantee of privacy, making it suitable for generating synthetic data that preserves statistical properties while preventing data leakage. In contrast, other methods like redaction, masking, or using a public model do not inherently prevent the model from memorizing and reproducing sensitive information.

Exam trap

The trap here is that candidates confuse data masking or redaction (which only hide data in the training set) with techniques that prevent model memorization, overlooking that models can still leak sensitive information through inference even when the input data is obfuscated.

How to eliminate wrong answers

Option A is wrong because Vertex AI PII redaction service only removes or obscures personally identifiable information from existing text, but does not generate synthetic data; the underlying real data remains and could still be leaked through model memorization. Option B is wrong because using a public foundation model without fine-tuning does not generate synthetic data specific to the firm's domain; it may produce generic outputs that lack the required statistical fidelity, and it does not provide any privacy guarantee against leaking real customer data. Option C is wrong because data masking before training only obscures fields in the training dataset, but the model can still memorize and reconstruct masked values through inference attacks, especially if the masking is deterministic or reversible.

714
Multi-Selectmedium

A financial services firm is evaluating Google Cloud generative AI offerings to build an internal assistant that answers employee policy questions using the firm's own documents. The firm requires that answers be grounded in those documents rather than the model's general knowledge, and that the assistant cite the source passages it used. Which two capabilities should the firm rely on? (Choose two.)

Select 2 answers
A.Raising the model temperature so responses vary more across repeated questions
B.Tuning a model on the firm's documents so answers are memorized in the model weights
C.Citations that return the source passages the model used for its answer
D.Grounding with Google Search to pull current public web results into the response
E.Grounding with Vertex AI Search to retrieve relevant passages from the firm's document store
AnswersC, E

Citations surface the specific document passages behind a generated answer, giving users verifiable provenance. That directly fulfills the firm's requirement that the assistant cite the sources it relied on, and it complements retrieval grounding by making the retrieved evidence visible so employees can check the policy text themselves.

Why this answer

Grounding with Vertex AI Search retrieves authoritative passages from the firm's own corpus so answers rest on internal documents, and citations expose the exact passages used, delivering the provenance the firm demands. Tuning does not guarantee traceable grounding, Google Search grounding pulls public content, and higher temperature only adds randomness, so those do not meet the requirements.

Exam trap

The trap here is assuming that fine-tuning on internal documents produces grounded, citable answers, when grounding and citations are what actually tie a response to source passages.

715
MCQeasy

A retail company with a large FAQ database wants to build a generative AI customer service chatbot that can answer questions accurately with up-to-date information. Which business strategy should they prioritize?

A.Use retrieval-augmented generation (RAG) with vector search on the FAQ database.
B.Train a new model from scratch using the FAQ data.
C.Fine-tune a foundational model on the entire FAQ dataset.
D.Use a general-purpose language model without any customization.
AnswerA

RAG with vector search retrieves relevant FAQ passages at query time and grounds the model's responses in them, keeping answers accurate and current without retraining. This directly satisfies the up-to-date information constraint while leveraging the existing FAQ database.

Why this answer

Retrieval-augmented generation (RAG) with vector search allows the chatbot to dynamically retrieve the most relevant, up-to-date FAQ entries from a large database at inference time, grounding the generative model's responses in verified content without requiring retraining. This approach combines the flexibility of a pre-trained language model with the accuracy of real-time information retrieval, ensuring answers reflect the latest FAQ updates.

Exam trap

Google Cloud often tests the misconception that fine-tuning is the best way to inject domain knowledge, but the trap here is that fine-tuning cannot efficiently handle frequently changing data, whereas RAG provides a modular, update-friendly architecture that avoids retraining costs.

How to eliminate wrong answers

Option B is wrong because training a new model from scratch on FAQ data is computationally prohibitive, requires massive datasets and resources, and still cannot guarantee up-to-date answers without frequent retraining. Option C is wrong because fine-tuning a foundational model on the entire FAQ dataset risks catastrophic forgetting of general language capabilities and does not inherently handle dynamic updates; any FAQ change would require re-fine-tuning. Option D is wrong because a general-purpose language model without customization lacks domain-specific knowledge and cannot access the company's proprietary FAQ database, leading to hallucinated or outdated answers.

716
MCQmedium

A software development team builds an internal code assistant using a generative model. The assistant writes Python functions that often contain security vulnerabilities such as SQL injection or command injection. The team wants to mitigate these vulnerabilities without adding a manual review step for every code snippet, as that would slow development. They have access to a static analysis security scanner API. Which approach best addresses the vulnerabilities while maintaining developer velocity?

A.Increase top-k sampling to generate a wider variety of code tokens.
B.After each generation, automatically run the code through the static analysis scanner, and if vulnerabilities are found, send the output back to the model for revision with the scanner's feedback.
C.Fine-tune the model on a corpus of secure code examples.
D.Add a system prompt: 'Do not generate code with security vulnerabilities.'
AnswerB

Feeding scanner findings back into the model creates an automated generate-scan-revise loop, so vulnerable patterns are corrected before the developer sees the snippet. This satisfies the no-manual-review constraint while preserving velocity, unlike prompt-only hardening, which cannot verify the emitted code.

Why this answer

It creates an automated feedback loop: the static analysis scanner detects vulnerabilities in the generated code, and the model revises the output based on that feedback. This approach directly mitigates security flaws without requiring manual review, preserving developer velocity. It leverages the scanner's precise, rule-based detection to iteratively improve the model's output, which is more reliable than relying on the model's inherent safety.

Exam trap

This exam often tests the misconception that a simple prompt or fine-tuning alone can guarantee safety, when in reality, a closed-loop validation with a dedicated security tool is required for reliable mitigation of injection vulnerabilities.

How to eliminate wrong answers

Option A is wrong because increasing top-k sampling broadens token selection, which can actually introduce more unpredictable and insecure code patterns, not reduce vulnerabilities. Option C is wrong because fine-tuning on secure code examples improves the model's baseline but does not guarantee that every generated snippet will be free of vulnerabilities, especially for novel or context-specific injection attacks. Option D is wrong because a system prompt is a weak, non-enforceable instruction; the model lacks true understanding of security and can easily generate vulnerable code despite the prompt, as it does not perform actual validation.

717
MCQmedium

A company wants to build an internal knowledge base that allows employees to ask questions about company policies in natural language. The knowledge base is stored in a Google Cloud SQL database. Which architecture should they use?

A.Use Gemini API with a prompt that includes all policies
B.Use AutoML Natural Language to classify questions
C.Export Cloud SQL to BigQuery and use BigQuery ML
D.Use Vertex AI Agent Builder with Grounding to connect to Cloud SQL
AnswerD

Vertex AI Agent Builder with Grounding queries Cloud SQL directly, retrieving policy rows as grounding context before generation. This satisfies the natural-language requirement while keeping answers anchored to the live database rather than a static index.

Why this answer

Vertex AI Agent Builder with Grounding allows connecting to the database and answering questions based on its content, providing a natural language interface.

718
MCQeasy

A developer uses a generative AI model with the system instruction shown. The response is correct but very brief. Which parameter adjustment could encourage more detail without losing accuracy?

A.Add 'Provide a detailed response' to the system instruction.
B.Set temperature to 0 to make output deterministic.
C.Set topK to 1 to focus on most likely tokens.
D.Increase temperature to 1.5 to encourage creativity.
AnswerA

Adding an explicit instruction for detail to the system instruction steers the model's generation toward longer, more thorough responses while the underlying task and data remain unchanged, preserving accuracy. This directly addresses the brevity without altering model parameters.

Why this answer

Modifying the system instruction to explicitly request a detailed response directly influences the model's output behavior without altering its underlying probability distribution. This approach preserves accuracy by keeping temperature, topK, and other sampling parameters at their default values, ensuring the model remains faithful to the training data while simply prompting for more elaboration.

Exam trap

The Google Gen AI Leader exam often tests the misconception that increasing randomness (temperature) or restricting token selection (topK) can improve detail, when in fact these parameters trade off accuracy for diversity or determinism, and the correct approach is to use prompt engineering to guide output length and style.

How to eliminate wrong answers

Option B is wrong because setting temperature to 0 makes the model deterministic by always selecting the highest-probability token, which typically results in shorter, repetitive, and less detailed responses—the opposite of what is needed. Option C is wrong because setting topK to 1 restricts token selection to only the single most likely token at each step, which similarly reduces output diversity and detail, often leading to generic or truncated answers. Option D is wrong because increasing temperature to 1.5 increases randomness in token sampling, which can introduce hallucinations, factual errors, or irrelevant content, thereby sacrificing accuracy for creativity.

719
MCQmedium

A healthcare company is building a clinical decision support system using Gemini 1.5 Pro on Vertex AI. They need responses that are highly accurate and comply with medical regulations, including traceability to source documents. They have a large corpus of curated medical guidelines stored in PDFs in Cloud Storage. Their team has experience with both fine-tuning and prompt engineering. Which approach best ensures regulatory compliance and accuracy?

A.Use a combination of grounding to the medical guidelines and prompt engineering with system instructions specifying compliance requirements.
B.Use prompt engineering with system instructions and few-shot examples, but no grounding.
C.Use grounding to the medical guidelines but rely on prompt engineering only for compliance instructions.
D.Fine-tune the model on the medical guidelines corpus to internalize the knowledge.
AnswerA

Grounding via Vertex AI Search retrieves passages from the curated PDF guidelines, so every response cites verifiable source documents — satisfying the traceability mandate. System instructions then enforce regulatory constraints at inference time, and because the corpus stays authoritative, accuracy improves without retraining risk. Fine-tuning cannot provide citation-level provenance.

Why this answer

Grounding the model to the curated medical guidelines in Cloud Storage ensures responses are directly traceable to source documents, which is critical for medical regulatory compliance. Combining this with system instructions that specify compliance requirements (e.g., HIPAA, FDA guidelines) enforces behavioral constraints without altering the model's weights, maintaining accuracy and auditability.

Exam trap

The Generative AI Leader exam often tests the misconception that fine-tuning is the best way to ensure accuracy and compliance for domain-specific tasks, but the trap here is that fine-tuning sacrifices traceability and can introduce staleness, whereas grounding with system instructions preserves source attribution and regulatory compliance.

How to eliminate wrong answers

Option B is wrong because relying solely on prompt engineering without grounding provides no mechanism to enforce traceability to specific source documents, making it impossible to meet medical regulatory requirements for evidence-based responses. Option C is wrong because while grounding provides source traceability, relying on prompt engineering alone for compliance instructions is insufficient; system instructions must be explicitly set in the model configuration to ensure consistent enforcement of regulatory constraints across all interactions. Option D is wrong because fine-tuning the model on the medical guidelines corpus internalizes knowledge into the model weights, which can lead to hallucination or outdated information over time, and critically, it breaks traceability to specific source documents since the model cannot cite exact PDF locations or versions.

720
Multi-Selecteasy

Which THREE of the following are generative AI modalities supported by Google Cloud services?

Select 3 answers
A.Image generation
B.Code generation
C.Text generation
D.Tabular data generation
E.Speech generation
AnswersA, B, C

Image generation is a core generative AI modality on Google Cloud, delivered through Vertex AI's Imagen models. It satisfies the stem's requirement by producing novel visual content from text prompts, rather than merely classifying or analysing existing images. This confirms it as one of the three supported modalities.

Why this answer

Option A (Image generation) is correct because Google Cloud's Vertex AI supports image generation through models such as Imagen, which produce images from text prompts. Option B (Code generation) is correct because Vertex AI offers Codey and Gemini models that generate, complete, and explain code. Option C (Text generation) is correct because Vertex AI's PaLM, Gemini, and other large language models generate and summarize text.

Option E (Speech generation) is also a genuine generative AI capability on Google Cloud via Vertex AI Text-to-Speech (Chirp voices), but since the question asks for THREE, the intended core modalities are text, code, and image. Option D (Tabular data generation) is not a core generative AI modality offered by Google Cloud's generative AI services.

Exam trap

Google Cloud exams often test the distinction between core generative AI modalities (text, code, image) and other AI services (like tabular data generation) that are not considered primary generative AI capabilities in the context of Vertex AI foundation models. Note that speech generation is a real generative capability on Google Cloud, so the question must clearly ask for THREE to make text/code/image the intended answer.

721
MCQmedium

A company wants to generate high-quality images from text descriptions for their marketing materials. They need the ability to edit specific regions of an image without regenerating the entire image. Which Google Cloud service should they use?

A.Gemini
B.Codey
C.Imagen
D.Veo
AnswerC

Imagen on Vertex AI supports inpainting and outpainting, letting you mask a region and regenerate only that area from a text prompt while the rest of the image stays untouched. This directly satisfies the stem's requirement to edit specific regions without regenerating the entire image.

Why this answer

Imagen offers inpainting and outpainting capabilities for editing specific regions. Gemini is multimodal but not optimized for image editing. Veo generates video, not still images.

Codey is for code.

722
MCQhard

A generative AI model for chatbot responses sometimes produces toxic language. The team wants to reduce toxicity without significantly affecting the model's helpfulness. Which approach is best?

A.Increase the temperature parameter
B.Reduce the maximum output tokens
C.Fine-tune with a dataset of non-toxic responses and use RLHF
D.Apply a toxicity classifier as a post-processing filter
AnswerC

Supervised fine-tuning on non-toxic responses teaches safer output distributions, and RLHF then optimises the policy against a reward model balancing toxicity reduction with helpfulness. This directly targets the constraint of lowering toxicity without materially degrading response quality.

Why this answer

Fine-tuning with a curated dataset of non-toxic responses directly adjusts the model's weights to reduce the likelihood of generating toxic language, while RLHF (Reinforcement Learning from Human Feedback) further aligns the model with human preferences for helpfulness and safety. This combined approach addresses the root cause of toxicity in the model's behavior without the blunt trade-offs of other methods, preserving the model's utility.

Exam trap

Google Cloud often tests the misconception that post-processing filters (like toxicity classifiers) are sufficient for safety, when in fact they fail to address the model's learned behavior and can degrade helpfulness due to false positives, making fine-tuning with RLHF the superior alignment technique.

How to eliminate wrong answers

Option A is wrong because increasing the temperature parameter increases randomness in token selection, which can actually amplify the probability of generating toxic or nonsensical outputs, not reduce them. Option B is wrong because reducing the maximum output tokens limits response length but does not influence the content or safety of the generated tokens, leaving toxicity unchanged. Option D is wrong because applying a toxicity classifier as a post-processing filter only masks toxic outputs after generation, wasting computational resources and potentially blocking helpful responses that contain false-positive flagged terms, without fixing the underlying model behavior.

723
MCQeasy

A startup with limited budget wants to quickly test a generative AI use case for personalized email marketing. Which approach minimizes time-to-market and cost?

A.Hire a team of AI researchers to build a solution.
B.Develop a custom model from scratch.
C.Fine-tune a large open-source model on internal data.
D.Use a managed API like the PaLM API with prompt engineering.
AnswerD

A managed API such as the PaLM API removes infrastructure, training, and hosting overhead, letting the startup validate personalised email marketing through prompt engineering alone. This directly minimises both time-to-market and cost, matching the limited-budget constraint.

Why this answer

Using a managed API like the PaLM API with prompt engineering eliminates the need for infrastructure setup, model training, and data preparation. This approach leverages a pre-trained model via a simple REST API call, allowing the startup to iterate on prompts and achieve personalized email content in hours rather than weeks, minimizing both time-to-market and cost.

Exam trap

Google Cloud often tests the misconception that fine-tuning (Option C) is always the fastest and cheapest path for customization, but the trap here is that fine-tuning still requires significant compute and data preparation, whereas prompt engineering on a managed API is truly zero-infrastructure and pay-per-use, making it the optimal choice for a quick, low-cost test.

How to eliminate wrong answers

Option A is wrong because hiring a team of AI researchers is expensive and time-consuming, requiring salaries, compute resources, and months of development, which contradicts the limited budget and quick testing goal. Option B is wrong because developing a custom model from scratch demands vast amounts of labeled data, significant GPU/TPU compute, and deep expertise, making it cost-prohibitive and slow for a rapid proof-of-concept. Option C is wrong because fine-tuning a large open-source model still requires substantial compute for training (e.g., GPU hours for LoRA or full fine-tuning), data curation, and deployment overhead, which exceeds the minimal cost and speed constraints of a quick test.

724
MCQmedium

A healthcare organization wants to use generative AI to draft email responses to patient inquiries. They need to ensure that the model never generates medical advice and always includes a disclaimer. Where should they enforce these constraints?

A.Select a model from Model Garden that is pre-trained on medical data
B.Fine-tune the model with a dataset that only includes disclaimers
C.Use Grounding with Google Search to restrict knowledge to approved sources
D.Configure safety settings and system instructions in the Vertex AI API call
AnswerD

System instructions define persistent behavioural guardrails, and safety settings filter prohibited content at the model layer. Enforcing both in the Vertex AI API call ensures every response omits medical advice and appends the disclaimer, regardless of caller.

Why this answer

Safety settings and system instructions are applied at the API call level to constrain model behavior. Fine-tuning would embed the rules but is heavy. Grounding doesn't enforce constraints.

Model Garden is for model selection.

725
Multi-Selectmedium

A university's IT department is evaluating Google Cloud generative AI offerings to build a course-assistant tool for students. They want to reduce engineering effort by using managed services rather than hosting models themselves. Which two Google Cloud offerings should they consider? (Choose two.)

Select 2 answers
A.A custom training pipeline built on Vertex AI Training to pretrain a domain-specific language model.
B.A self-managed open-weights model served from a Google Kubernetes Engine cluster with GPU node pools.
C.BigQuery ML with a remote model that forwards prediction requests to an external vendor's API.
D.Vertex AI Agent Builder for assembling a grounded agent with tools and a knowledge base.
E.Vertex AI's Gemini models accessed through a managed API endpoint.
AnswersD, E

Agent Builder provides a managed framework for defining instructions, connecting data stores for grounding, and adding tools such as function calls, so the department can assemble a functional assistant without building orchestration, retrieval, or conversation-state handling from scratch. That materially lowers engineering effort compared with a hand-rolled agent loop running on self-managed compute.

Why this answer

The department wants managed services to cut engineering effort. Gemini through Vertex AI supplies the model without infrastructure ownership, and Agent Builder supplies the orchestration, grounding, and tooling layer needed for a usable course assistant. Self-managing a model on GKE, pretraining a custom model, or forwarding requests to an external API all reintroduce significant engineering or governance work that the managed first-party offerings avoid.

Exam trap

The trap here is equating 'hosted on Google Cloud' with 'managed', when GKE-served open-weights models still leave the customer responsible for serving infrastructure and lifecycle management.

726
Multi-Selectmedium

A logistics company uses a generative AI model to draft incident reports from sensor logs. Reviewers find the reports are often incomplete, missing fields such as root cause and corrective action. The team wants to improve output completeness without retraining the model. Which TWO techniques should they use? (Choose two.)

Select 2 answers
A.Shorten the input sensor logs to only the most recent readings.
B.Add few-shot examples of complete incident reports that show every required field filled in.
C.Increase the model's temperature to encourage the model to explore more content.
D.Reduce the maximum output tokens to force the model to be concise.
E.Define a structured output schema listing every required field and instruct the model to populate each one.
AnswersB, E

Complete exemplars demonstrate the expected coverage pattern, teaching the model which sections must appear and how detailed they should be. Combined with a schema, examples reinforce the habit of filling all fields. This is a prompt-level change that improves completeness immediately without weight updates or new training data collection.

Why this answer

Completeness improves when the required output shape is made explicit and demonstrated. A structured schema enumerates every field so omissions become visible and correctable, while few-shot examples of complete reports show the model the expected coverage and detail. Sampling and length controls do not enforce required fields and can reduce or distort content.

Exam trap

The trap here is assuming that more creative sampling or longer/shorter outputs improve report quality, when completeness specifically requires an explicit field schema plus examples of fully populated reports.

727
MCQhard

A global bank wants to deploy a generative AI assistant for employees across multiple European countries, each with strict data residency laws. Which deployment strategy is most compliant?

A.Deploy separate model instances in each country's cloud region.
B.Use a federated learning approach where data stays on-premises.
C.Deploy a single model in a US region and use data masking.
D.Use a third-party API that processes data outside Europe.
AnswerA

Ensures data never leaves the country, meeting local compliance requirements.

Why this answer

Deploying separate model instances in each country's cloud region ensures that data never crosses national borders, directly complying with strict data residency laws like the GDPR's data localization requirements. This strategy uses regional cloud infrastructure (e.g., AWS eu-central-1, Azure westeurope) to keep both training and inference data within the specific jurisdiction, avoiding any cross-border data transfer.

Exam trap

Google Cloud often tests the misconception that data masking or anonymization alone satisfies data residency laws, but the trap here is that data residency requires the data to physically remain within the jurisdiction, not just be obfuscated.

How to eliminate wrong answers

Option B is wrong because federated learning only keeps training data on-premises, but the model parameters or gradients must still be exchanged with a central server, which can violate data residency if that server is outside the country. Option C is wrong because deploying a single model in a US region and using data masking does not prevent the underlying data from being processed or stored in the US, which violates EU data residency laws like GDPR. Option D is wrong because using a third-party API that processes data outside Europe directly violates data residency requirements, as the data physically leaves the European Economic Area (EEA) without adequate safeguards.

728
MCQmedium

A regional insurance company wants to launch a generative AI assistant that answers policyholder questions. Legal requires that every customer-facing answer be traceable to an approved policy document, and that the system never invent coverage terms. The company has about 40,000 internal policy PDFs that change quarterly. Which approach should the GenAI Leader recommend?

A.Deploy the assistant on Gemini with a system instruction telling it to be accurate and to avoid making up policy details.
B.Fine-tune a foundation model on the full set of policy PDFs and deploy the tuned model to a Vertex AI endpoint for the assistant to call.
C.Use a large context window model and paste the entire policy library into every prompt so the model always sees all approved documents.
D.Ground the assistant with retrieval-augmented generation using Vertex AI Search over the policy document corpus, and require citations in responses.
AnswerD

Retrieval-augmented generation retrieves the most relevant approved policy passages at query time and instructs the model to answer only from that retrieved context, so every answer can cite a source document. Because Vertex AI Search indexes the corpus and can be re-indexed each quarter, the assistant stays aligned with current policy language instead of relying on parametric memory that cannot be audited.

Why this answer

Grounding the assistant in the approved corpus with retrieval-augmented generation satisfies traceability because each response is generated from retrieved passages that can be cited, and it keeps answers current because the index is refreshed when policies change. Fine-tuning, whole-corpus prompting, and instruction-only approaches all leave the model without a verifiable source for coverage terms.

Exam trap

The trap here is assuming that fine-tuning on internal documents produces citations, when tuning only changes model weights and leaves no retrievable source to reference.

729
Multi-Selecthard

Which THREE factors should you consider when selecting a foundation model from Model Garden? (Choose three.)

Select 3 answers
A.Number of model versions
B.The color of the model card
C.Model size
D.Model accuracy on benchmarks
E.Model license
AnswersC, D, E

Model size determines inference latency, memory footprint and hosting cost, so it must match the deployment target's compute budget and latency requirements. Larger models generally offer greater capability but demand more resources, making size a primary selection constraint.

Why this answer

When selecting a foundation model from Model Garden, model size (C) matters because parameter count directly affects inference latency, throughput, memory footprint, and cost, so it must match your deployment and budget constraints. Model accuracy on benchmarks (D) is essential because benchmark results (e.g., MMLU, GSM8K, HELM scores) indicate how well the model performs on the tasks relevant to your use case. Model license (E) must be evaluated because it determines whether commercial use, redistribution, or fine-tuning is permitted, which directly impacts legal and business viability.

The number of model versions (A) is not a primary selection factor since version count alone says nothing about capability, cost, or fit. The color of the model card (B) is purely cosmetic and has no bearing on model selection.

Exam trap

The exam tests candidates' ability to distinguish between superficial UI elements (like card color) and substantive technical criteria (like model size, accuracy, and license) that directly affect deployment and compliance.

730
MCQhard

A cloud architect is designing a generative AI pipeline that must comply with the EU AI Act for high-risk AI systems. Which of the following is a mandatory requirement under the Act?

A.The system must be explainable using chain-of-thought reasoning
B.The system must achieve a minimum accuracy of 90% on validation data
C.The system must be trained on data that is representative of the target population
D.The system must undergo a conformity assessment before deployment
AnswerD

Under the EU AI Act, high-risk systems must complete a conformity assessment before being placed on the market, verifying compliance with requirements such as risk management, data governance and technical documentation. This is the mandatory pre-deployment gate the architect's generative AI pipeline must satisfy.

Why this answer

Under the EU AI Act, high-risk AI systems must undergo a conformity assessment before deployment to ensure compliance with requirements such as risk management, data governance, and transparency. This is a mandatory procedural step, not a performance metric or specific reasoning technique. The assessment may involve self-evaluation or third-party review depending on the system's risk category.

Exam trap

Candidates often confuse aspirational best practices (like explainability or accuracy thresholds) with actual legal mandates, leading them to pick a plausible-sounding but non-mandatory option like A or B instead of the procedural requirement in D.

How to eliminate wrong answers

Option A is wrong because the EU AI Act does not mandate a specific explainability technique like chain-of-thought reasoning; it requires general transparency and interpretability, but the method is left to the provider. Option B is wrong because the Act does not prescribe a fixed accuracy threshold like 90%; it requires appropriate levels of accuracy based on the system's intended purpose and risk, validated against representative data. Option C is wrong because while data representativeness is a key principle under the Act's data governance requirements, it is not a standalone mandatory requirement; the Act mandates that training, validation, and testing datasets be relevant, representative, and free from biases, but this is part of broader data governance obligations, not a single checkbox.

731
MCQhard

A financial services firm uses a Gemini model to generate quarterly risk summaries from internal reports. Reviewers note that summaries sometimes contradict the source tables. The team wants the model to reason step by step over the figures before writing the summary. Which technique should they use?

A.Enable a higher top-p value so the model samples more broadly and can reconcile conflicting numbers.
B.Set temperature to zero so the model always produces the same summary and cannot contradict the tables.
C.Use chain-of-thought prompting by instructing the model to compute and list intermediate values before producing the final summary.
D.Lower the maximum output tokens so the model is forced to be concise and avoid contradictory statements.
AnswerC

Chain-of-thought prompting directs the model to work through intermediate calculations and comparisons before drafting the summary, which reduces contradictions because the figures are reconciled first. It makes the reasoning inspectable, so reviewers can spot errors. This technique is well suited to tasks that require multi-step arithmetic and consistency checking.

Why this answer

Chain-of-thought prompting makes the model compute and compare intermediate values before writing, which is exactly what is needed to keep a summary consistent with source tables. Sampling changes and output limits affect randomness or length but not numerical reasoning. When a task requires multi-step reconciliation, prompting for explicit steps is the appropriate technique.

Exam trap

The trap here is treating determinism or brevity as a cure for factual inconsistency, when neither supplies the step-by-step reconciliation the task requires.

732
MCQmedium

A team is developing a mobile app that must run AI inference on-device for low latency and offline capability. Which Gemini model variant is designed specifically for on-device deployment?

A.Gemini Pro
B.Gemini Nano
C.Gemini Ultra
D.Gemini Flash
AnswerB

Gemini Nano is Google's smallest Gemini variant, built to run locally on devices such as smartphones. It delivers on-device inference without network calls, satisfying both the low-latency and offline requirements that cloud-hosted Gemini Pro or Ultra cannot meet.

Why this answer

Gemini Nano is the smallest and most efficient model in the Gemini family, specifically optimized for on-device deployment. It is designed to run directly on mobile devices (e.g., Android phones) using hardware acceleration like Google's Pixel Neural Core or Qualcomm's AI Engine, enabling low-latency inference and offline capability without requiring a cloud connection.

Exam trap

The trap here is that candidates confuse 'lightweight cloud model' (Gemini Flash) with 'on-device model' (Gemini Nano), assuming any 'fast' or 'small' variant is suitable for mobile deployment, but Flash still requires cloud connectivity and is not optimized for local hardware constraints.

How to eliminate wrong answers

Option A is wrong because Gemini Pro is a mid-size model intended for cloud-based, high-performance tasks such as complex reasoning and multimodal analysis, not for on-device deployment due to its larger memory and compute requirements. Option C is wrong because Gemini Ultra is the largest and most capable model, designed for enterprise-scale cloud workloads and advanced research, making it unsuitable for resource-constrained mobile devices. Option D is wrong because Gemini Flash is a lightweight cloud model optimized for speed and cost in cloud inference, but it is not purpose-built for on-device execution and still requires a network connection.

733
MCQeasy

A business wants to build a generative AI application but has limited data science resources. What is the recommended path?

A.Use Vertex AI's AutoML and pre-built APIs to accelerate development
B.Hire a team of ML engineers to develop an in-house solution
C.Purchase a third-party generative AI SaaS product off-the-shelf
D.Build a custom model from scratch using TensorFlow
AnswerA

Vertex AI's AutoML and pre-built APIs let teams with limited data science expertise train custom models and integrate generative capabilities without building pipelines from scratch. This accelerates development, satisfying the constraint of scarce specialist resources.

Why this answer

Vertex AI's AutoML and pre-built APIs are the recommended path because they allow the business to leverage Google's managed infrastructure and pre-trained models, significantly reducing the need for in-house data science expertise. AutoML automates model training, tuning, and deployment, while pre-built APIs (e.g., for vision, language) provide immediate access to generative capabilities without custom development. This approach accelerates time-to-market and lowers the barrier to entry for organizations with limited ML resources.

Exam trap

A common mistake is to assume that limited data science resources require outsourcing all AI work (Option C) or building from scratch (Option D), when the correct answer uses Google's managed services to reduce the need for in-house expertise while still allowing customization.

How to eliminate wrong answers

Option B is wrong because hiring a full team of ML engineers is resource-intensive and contradicts the premise of limited data science resources; it also introduces significant overhead in recruitment, management, and infrastructure. Option C is wrong because purchasing a third-party SaaS product off-the-shelf may not offer the customization, data privacy controls, or integration flexibility needed for a generative AI application, and it can lock the business into a vendor's roadmap. Option D is wrong because building a custom model from scratch using TensorFlow requires deep ML expertise, extensive training data, and computational resources, which is impractical for a team with limited data science capabilities and would delay deployment.

734
MCQmedium

A developer is using Vertex AI Studio to experiment with prompts. They want to ensure that the model's responses are grounded in factual information from a trusted knowledge base. Which feature should they enable?

A.Safety filters
B.Temperature setting reduction
C.Chain-of-thought prompting
D.Grounding with a Vertex AI Search data store
AnswerD

Grounding with a Vertex AI Search data store retrieves passages from the trusted knowledge base and supplies them to the model, so responses cite verifiable sources rather than relying on parametric memory. This satisfies the requirement that answers be grounded in factual information.

Why this answer

Vertex AI's grounding feature allows the model to cite sources from a provided knowledge base, improving factual accuracy and verifiability.

735
MCQmedium

A financial services company is deploying a generative AI model to summarize sensitive customer emails. The security team requires that no email content is used to train or improve the underlying foundation model. Which Google Cloud approach ensures this requirement is met?

A.Rely on the default data processing terms of Google Cloud, which state that customer data is never used to train foundation models.
B.Fine-tune the model on the email data using Vertex AI, then delete the fine-tuned model after use.
C.Deploy the model in a private VPC and use Cloud VPN to encrypt all traffic between the application and the model endpoint.
D.Use Vertex AI with customer-managed encryption keys (CMEK) and enable VPC Service Controls.
AnswerA

Google Cloud's terms of service for generative AI explicitly state that customer data is not used to train or improve foundation models without explicit permission. This contractual guarantee, combined with technical measures like data residency, ensures the security team's requirement is met. It is the foundational assurance that no email content will be used for training.

Why this answer

Google Cloud's generative AI services are governed by terms that prohibit using customer data to train or improve foundation models without explicit consent. This contractual commitment is the primary control that satisfies the security team's requirement. Technical measures like encryption or network isolation are important for security but do not address the specific concern about training data usage.

Exam trap

The trap here is assuming that network or encryption controls prevent data from being used for model training, when the guarantee is actually contractual and policy-based.

736
Multi-Selecthard

Which THREE are best practices for designing prompts for a generative AI model?

Select 3 answers
A.Provide few-shot examples for complex tasks
B.Include specific and clear instructions
C.Break the task into smaller steps
D.Use negative prompts to avoid undesired outputs
E.Always set temperature to 1.0 for creativity
AnswersA, B, C

Correct: Examples guide the model toward desired outputs.

Why this answer

Providing few-shot examples (e.g., 2-5 input-output pairs) helps the model infer the desired pattern, reducing ambiguity for complex tasks like classification or structured extraction. This technique leverages in-context learning, where the model uses the examples as a template without fine-tuning.

Exam trap

A common misconception tested in Google's Gen AI evaluations is that negative prompts can reliably control outputs, but they often fail due to tokenization and probability smoothing, leading to the 'forbidden token' problem where undesired content still appears.

737
Multi-Selectmedium

Which TWO techniques can help improve the factual accuracy of a language model's outputs? (Choose two.)

Select 2 answers
A.Decrease the max output tokens.
B.Increase the temperature parameter.
C.Fine-tune on a domain-specific curated dataset.
D.Implement retrieval-augmented generation (RAG).
E.Use top-k random sampling.
AnswersC, D

Fine-tuning adapts the model to domain facts.

Why this answer

Fine-tuning on a domain-specific curated dataset (C) directly adjusts the model's weights using high-quality, verified examples, teaching it to produce factually correct outputs for that domain. This reduces hallucinations by grounding the model in accurate, relevant data rather than relying solely on its pre-training distribution.

Exam trap

Google Cloud often tests the misconception that adjusting decoding parameters (like temperature, top-k, or max tokens) can improve factual accuracy, when in reality these only control output style, length, or randomness, not the correctness of the underlying information.

738
MCQmedium

A software company is using a large language model to generate code snippets from natural language descriptions. The generated code often has syntax errors and does not follow the company's coding standards. Which approach is most effective to improve the quality of the generated code?

A.Fine-tune the model on a dataset of code that adheres to the company's coding standards.
B.Add a prompt instruction to 'write clean code'.
C.Increase the temperature to generate more diverse code solutions.
D.Use a larger model with more parameters without fine-tuning.
AnswerA

Fine-tuning on a dataset of code that follows the company's standards will teach the model the specific syntax, style, and patterns used. This directly addresses both syntax errors and adherence to coding standards, making it the most effective approach for this scenario.

Why this answer

Fine-tuning on a dataset of code that adheres to the company's coding standards is the most effective way to improve the quality of generated code. It directly teaches the model the desired syntax, style, and patterns, reducing syntax errors and ensuring compliance with standards.

Exam trap

The trap here is relying on prompt engineering alone to enforce coding standards, when fine-tuning provides a more reliable solution.

739
MCQhard

A company wants to use AI to make hiring decisions. They are concerned about bias against certain demographic groups. According to Google's AI Principles, which approach is MOST aligned?

A.Pre-train the model on a dataset that is balanced across all demographics
B.Blind the model to demographic features to ensure fairness
C.Only use the model for initial resume screening, with final decisions by humans
D.Evaluate the model using diverse test sets and adjust if bias is found
AnswerD

Testing the hiring model on diverse demographic test sets and correcting any measured disparity aligns with Google's fairness principle, which requires avoiding reinforcement of unfair bias. Empirical evaluation before deployment is the most defensible approach for high-stakes hiring decisions.

Why this answer

The principle 'avoid creating or reinforcing unfair bias' requires proactive identification and mitigation. Evaluating the model on diverse test sets is a standard way to detect and address bias before deployment.

740
MCQhard

A utility company's board approves a generative AI program and asks the program lead to present a plan showing how investment will be governed and how value will be tracked from pilot through production. The lead wants a framework that ties each initiative to a business owner, defines stage gates for continued funding, and specifies which metrics justify scaling. Which approach best meets the board's expectation?

A.Adopt an AI acceptable use policy that prohibits employees from entering confidential data into public generative AI tools.
B.Standardize on a single foundation model and require every initiative to use it to simplify procurement and support.
C.Create a value realization framework that links each generative AI initiative to a business owner, defines stage-gate criteria for continued investment, and specifies scale-up metrics.
D.Build a centralized generative AI center of excellence that provides prompt libraries and engineering support to all business units.
AnswerC

The board asked for governance of investment and tracking of value across the lifecycle. A value realization framework assigns accountability, sets explicit stage gates that determine whether funding continues after each phase, and names the metrics that must be met before scaling, which is precisely the structure needed to govern a multi-initiative generative AI program from pilot to production.

Why this answer

The board is asking for investment governance and value tracking across the program lifecycle. A value realization framework supplies exactly that: a named business owner per initiative, stage gates that decide whether funding continues after each phase, and pre-agreed metrics that must be met before scaling. This makes continued investment evidence-based rather than momentum-driven.

Exam trap

The trap here is answering a funding governance question with a technical or policy control such as model standardization or an acceptable use policy, neither of which defines ownership, stage gates, or scale-up metrics.

741
MCQeasy

A developer wants to add a GenAI feature to their existing web application. They need to integrate with the app's backend using REST APIs. Which integration pattern is MOST appropriate?

A.Use Apps Script to call the Gemini API
B.Use Vertex AI Agent Builder to create an agent and embed it via iframe
C.Integrate via Vertex AI API or Gemini API
D.Build a Google Workspace add-on
AnswerC

Vertex AI and Gemini APIs expose REST endpoints, so the backend can call them directly without SDK dependencies or platform lock-in. This satisfies the stem's REST integration constraint, letting the existing web application add GenAI capability through standard HTTP requests rather than bespoke client libraries.

Why this answer

API-first integration using Vertex AI API or Gemini API is the standard way to add GenAI to existing applications. Workspace add-ons are for Google Workspace apps. Apps Script is for automating Workspace.

Agent Builder is for building conversational agents, not general API integration.

742
MCQeasy

A data scientist is using a large language model to generate product descriptions. The descriptions are often too verbose. Which parameter adjustment is most appropriate?

A.Decrease the top-k value.
B.Increase the max output tokens.
C.Decrease the temperature.
D.Increase the frequency penalty.
AnswerD

Frequency penalty reduces repetitive phrases, encouraging conciseness.

Why this answer

Increasing the frequency penalty reduces the likelihood of the model repeating the same phrases or ideas, which directly addresses verbosity by discouraging repetitive or overly detailed descriptions. This parameter penalizes tokens that have already appeared in the generated text, promoting more concise and varied output. Other adjustments like temperature or top-k affect randomness and diversity but do not specifically target repetition or length.

Exam trap

Google Cloud often tests the distinction between parameters that control randomness (temperature, top-k) versus those that control repetition (frequency penalty, presence penalty), and the trap here is that candidates confuse 'less verbose' with 'less random' and incorrectly choose temperature or top-k adjustments.

How to eliminate wrong answers

Option A is wrong because decreasing the top-k value restricts the model to a smaller set of high-probability tokens, which can actually make output more predictable and potentially more repetitive, not less verbose. Option B is wrong because increasing the max output tokens allows the model to generate longer text, which would exacerbate verbosity rather than reduce it. Option C is wrong because decreasing the temperature makes the model more deterministic and conservative, often leading to safer but not necessarily shorter or less repetitive text; it does not directly penalize repetition or length.

743
MCQhard

A model generates responses that frequently repeat phrases or words. Which parameter adjustment is most likely to fix this?

A.Increase top_k
B.Increase temperature
C.Increase repetition penalty
D.Increase max output tokens
AnswerC

Repetition penalty directly down-weights tokens already emitted, so previously generated words become less probable at each decoding step. Raising it breaks the loop of recurring phrases, satisfying the stem's requirement to stop repeated words without altering the model itself.

Why this answer

Increasing the repetition penalty directly discourages the model from selecting tokens that have already appeared in the generated sequence, thereby reducing repetitive phrases or words. This parameter works by subtracting a fixed penalty from the logits of previously generated tokens before applying the softmax function, making them less likely to be chosen again.

Exam trap

The trap here is that candidates often confuse repetition penalty with diversity-promoting parameters like temperature or top_k, mistakenly believing that increasing randomness or narrowing token selection will fix repetition, when in fact those adjustments can worsen the problem.

How to eliminate wrong answers

Option A is wrong because increasing top_k limits the sampling pool to the k most likely next tokens, which can actually increase repetition by narrowing the diversity of choices. Option B is wrong because increasing temperature flattens the probability distribution, making all tokens more equally likely, which can lead to more random and potentially more repetitive outputs, not less. Option D is wrong because increasing max output tokens only extends the length of the generated response; it does not address the underlying cause of repetition and may even exacerbate it by allowing more opportunities for the model to loop on repeated phrases.

744
MCQeasy

A junior developer is experimenting with a large language model and notices that the same prompt produces different outputs each time it is run. Which characteristic of generative AI models explains this behavior?

A.The model retrains itself after every inference, incorporating the previous output into its weights.
B.The model caches previous responses and randomly selects one from its memory.
C.The model's temperature parameter is fixed at zero, causing deterministic outputs.
D.The model uses a stochastic sampling process to select the next token from a probability distribution.
AnswerD

Generative models assign probabilities to possible next tokens and sample from that distribution. Because sampling is random, repeated runs with the same prompt can yield different outputs. This stochasticity is intentional and controlled by parameters like temperature and top-k. The developer's observation directly reflects this probabilistic, non-deterministic generation process.

Why this answer

The variability in outputs arises because generative models sample from a probability distribution over possible next tokens. This stochastic sampling, influenced by parameters like temperature, leads to different sequences even with identical prompts. The other options incorrectly attribute the behavior to retraining, caching, or a fixed temperature of zero.

Exam trap

The trap here is assuming that a model's output should be deterministic like a traditional function, overlooking the probabilistic sampling that defines generative AI.

745
MCQmedium

A startup wants to generate realistic product videos from text descriptions for social media ads. Which Google Cloud service should they use?

A.Imagen
B.Gemini Pro Vision
C.Veo
D.Codey
AnswerC

Veo is Google Cloud's generative video model, producing realistic video clips from text prompts. It directly satisfies the requirement to generate product videos from text descriptions for social media ads, unlike text- or image-only services.

Why this answer

Veo is Google Cloud's advanced video generation model that can create high-quality, realistic videos from text prompts, making it the ideal choice for generating product videos for social media ads. Unlike other services, Veo is specifically designed for video synthesis, offering capabilities like style control and cinematic effects directly from text descriptions.

Exam trap

The trap here is that candidates often confuse multimodal understanding (Gemini Pro Vision) with generative creation (Veo), or assume that image generation (Imagen) can be trivially extended to video without understanding the distinct temporal modeling required.

How to eliminate wrong answers

Option A is wrong because Imagen is a text-to-image generation model, not a video generation service; it produces static images, not dynamic video content. Option B is wrong because Gemini Pro Vision is a multimodal model that can analyze and understand images and videos, but it does not generate new video content from text descriptions. Option D is wrong because Codey is a code generation model designed for assisting with programming tasks, not for generating visual media like videos.

746
MCQmedium

Refer to the exhibit. A sudden surge of traffic reaches 15,000 requests per second, but the endpoint can only handle 1,000 req/s per replica. What will happen to new requests?

A.They will be processed, and replicas will exceed maxReplicaCount.
B.They will be redirected to a different model.
C.They will receive HTTP 429 (Too Many Requests) errors.
D.They will be queued until capacity becomes available.
AnswerC

When incoming traffic exceeds the endpoint's per-replica capacity, the front end throttles rather than queues indefinitely, returning HTTP 429 responses to excess callers. The 15,000 req/s surge against 1,000 req/s per replica therefore produces Too Many Requests errors for new requests.

Why this answer

When a surge of 15,000 requests per second hits an endpoint configured with a maxReplicaCount (e.g., 10 replicas at 1,000 req/s each = 10,000 req/s capacity), any excess requests beyond that capacity are rejected with an HTTP 429 (Too Many Requests) status code. This is standard behavior in autoscaling systems: once the replica count reaches its maximum limit, the service cannot scale further, and new requests are throttled to prevent overload.

Exam trap

The trap here is that candidates assume autoscaling can handle any traffic surge indefinitely, ignoring the hard limit of maxReplicaCount, and thus incorrectly choose Option A or D, failing to recognize that HTTP 429 is the standard throttling mechanism when capacity is exhausted.

How to eliminate wrong answers

Option A is wrong because the maxReplicaCount is a hard upper limit; replicas cannot exceed this configured value, so new requests are not processed beyond that capacity. Option B is wrong because traffic redirection to a different model is not a standard behavior for capacity overflow; it would require explicit routing rules or a load balancer configured for failover, which is not implied in the scenario. Option D is wrong because queuing is not the default behavior for HTTP-based endpoints in this context; while some systems support request queuing (e.g., with message brokers), the exhibit describes a direct endpoint handling, and HTTP 429 is the standard response for rate limiting per RFC 6585.

747
MCQmedium

A national retail chain wants to deploy a generative AI assistant that recommends products to shoppers in six countries. The company's legal team requires that customer conversations never leave the country of origin, while the engineering team wants one consistent deployment pattern across all regions. Which Google Cloud approach best satisfies both requirements?

A.Configure VPC Service Controls per country and continue serving all shoppers from one central region.
B.Use a multi-region bucket for conversation logs and set a Cloud Storage lifecycle rule that deletes objects after 30 days.
C.Deploy the assistant separately in a Google Cloud region within each country and keep conversation storage in that same region.
D.Deploy the assistant in a single global region and rely on Google's global load balancing to route users to the nearest edge point of presence.
AnswerC

Running the assistant in a region inside each country keeps both inference and stored conversation data within national borders, satisfying the residency requirement. Reusing the same architecture, prompts, and deployment tooling in every region preserves the single consistent pattern the engineering team wants, so neither team's constraint is compromised.

Why this answer

Data residency for generative AI workloads requires that both inference and any retained conversation data physically stay inside the required jurisdiction. Only a per-country regional deployment achieves that while letting the team reuse one design, one prompt library, and one deployment mechanism everywhere. Perimeter controls, global routing, and multi-region storage change access or durability, not the physical location of processing.

Exam trap

The trap here is assuming that global load balancing or VPC Service Controls change where data is processed, when they only affect traffic routing and access boundaries.

748
Multi-Selecthard

A financial services firm is using a large language model to generate quarterly investment summaries from raw market data. The summaries occasionally contain fabricated statistics and sometimes omit key risk factors. The team wants to improve factual accuracy and completeness without retraining the model. Which TWO techniques should they apply? (Choose two.)

Select 2 answers
A.Increasing the temperature setting to encourage more creative wording.
B.Fine-tuning the model on a small set of historical summaries to improve style.
C.Retrieval-augmented generation (RAG) using a curated knowledge base of verified financial reports.
D.Implementing a grounding check with a fact-verification API that cross-references generated claims against trusted sources.
E.Using a chain-of-thought prompt that asks the model to reason step by step before writing the summary.
AnswersC, D

RAG grounds the model's output in retrieved, authoritative documents. By fetching relevant passages from verified reports, the model can cite accurate statistics and include risk factors present in those sources. This reduces hallucination and improves completeness because the model conditions on real data rather than relying solely on its parametric memory.

Why this answer

Retrieval-augmented generation supplies the model with verified, up-to-date information, reducing fabrication and improving coverage of key points. A fact-verification API adds a validation layer that catches any remaining inaccuracies. Together, they address both hallucination and omission without retraining the model.

Exam trap

The trap here is assuming that prompt engineering alone, such as chain-of-thought, can eliminate factual errors without external grounding.

749
MCQhard

A team is building a multi-modal agent that needs to accept a user's image of a handwritten note, convert it to text, and then run a sentiment analysis. They want to minimize latency and cost. Which approach is best?

A.Fine-tune Gemini Pro on handwritten notes and sentiment labels
B.Use Document AI for OCR and then call Codey for sentiment analysis
C.Use Gemini 1.5 Flash with a prompt that includes the image and asks for sentiment analysis in one call
D.Use Cloud Vision API for OCR, then feed the text to a sentiment analysis model via Vertex AI
AnswerC

Gemini 1.5 Flash natively accepts image input and performs OCR plus sentiment reasoning within a single multimodal inference, eliminating the separate transcription call and its added latency and token cost. This directly satisfies the stem's minimise-latency-and-cost constraint while handling the handwritten note end to end.

Why this answer

Gemini 1.5 Flash is a multimodal model that can directly process images and perform sentiment analysis in a single API call, eliminating the need for separate OCR and NLP services. This minimizes both latency (by reducing the number of sequential calls) and cost (by using a single, efficient model instead of multiple specialized services).

Exam trap

Google often tests the candidate's ability to recognize that multimodal models like Gemini 1.5 Flash can replace multi-step pipelines (OCR + NLP) in a single call, and the trap here is that candidates default to traditional separate-service architectures (like Cloud Vision + Vertex AI) without considering the latency and cost benefits of a unified multimodal approach.

How to eliminate wrong answers

Option A is wrong because fine-tuning Gemini Pro on handwritten notes and sentiment labels is overkill for this task, incurring high training costs and latency, and Gemini Pro is a larger, more expensive model than needed for simple OCR and sentiment analysis. Option B is wrong because Document AI is designed for structured document extraction, not general handwritten note OCR, and Codey is a code generation model, not a sentiment analysis model, making this combination technically mismatched. Option D is wrong because using Cloud Vision API for OCR followed by a separate sentiment analysis model via Vertex AI introduces additional latency and cost from multiple API calls, whereas Gemini 1.5 Flash can achieve the same result in one step.

750
Multi-Selectmedium

A company is developing an AI-powered interview assistant that screens job applicants. The responsible AI team wants to ensure the model does not discriminate based on gender, race, or age. Which TWO practices should they implement?

Select 2 answers
A.Deploy the model without human oversight to ensure consistency.
B.Regularly evaluate the model's outputs for bias using intersectional test sets.
C.Use SynthID watermarking on all model outputs.
D.Remove all demographic attributes from the training data to ensure fairness.
E.Use a diverse and representative training dataset that includes candidates from various demographics.
AnswersB, E

Intersectional test sets expose compounded bias across overlapping attributes such as race and gender, which single-axis evaluations miss. This directly satisfies the stem's requirement to prevent discrimination on gender, race, and age, since a model can pass isolated checks yet still disadvantage, for example, older women.

Why this answer

Option B is correct because regularly evaluating the model's outputs for bias using intersectional test sets allows the team to detect discrimination across combinations of protected attributes such as gender, race, and age, which is essential for responsible AI screening tools. Option E is correct because a diverse and representative training dataset that includes candidates from various demographics helps reduce sampling bias and enables the model to learn patterns that generalize fairly across groups. Option A is not appropriate because deploying without human oversight removes the ability to catch discriminatory or erroneous decisions, and consistency alone does not guarantee fairness.

Option C is unrelated because SynthID watermarking is used to mark AI-generated content, not to mitigate bias in hiring models. Option D is not sufficient and can be harmful because simply removing demographic attributes does not ensure fairness and may allow proxy variables to perpetuate discrimination while also preventing the measurement of disparate impact.

Page 9

Page 10 of 14

Page 11