Courseiva

CCNA Google Cloud's Generative AI Offerings Questions

75 of 182 questions · Page 1/3 · Google Cloud's Generative AI Offerings · Answers revealed

1
MCQmedium

A company is building a generative AI chatbot for customer support using Vertex AI. They want to ground the model responses with their internal knowledge base stored in Cloud Storage and BigQuery. Which feature should they use to ensure the model only answers from the provided data and avoids hallucination?

A.Vertex AI Grounding with Vertex AI Search
B.Vertex AI Prediction
C.Vertex AI Pipelines
D.Cloud Functions
AnswerA

Vertex AI Grounding with Vertex AI Search connects the model to your Cloud Storage and BigQuery repositories, retrieving relevant passages before generation. This retrieval-augmented approach constrains responses to indexed enterprise content, satisfying the requirement that answers derive solely from the internal knowledge base and reducing hallucination.

Why this answer

Vertex AI Grounding with Vertex AI Search is the correct feature because it allows the model to retrieve and cite information from a specified data source (such as Cloud Storage and BigQuery) to generate responses. This process, known as grounding, ensures the model's output is based solely on the provided authoritative data, effectively reducing hallucinations by constraining the model to factual, retrieved content rather than relying on its internal parametric knowledge.

Exam trap

The trap here is that candidates may confuse Vertex AI Prediction (a general model serving endpoint) with the grounding feature, mistakenly thinking that simply deploying a model with Vertex AI Prediction will automatically restrict its answers to a specific knowledge base, when in fact grounding requires explicit integration with Vertex AI Search and a configured data store.

How to eliminate wrong answers

Option B is wrong because Vertex AI Prediction is a service for deploying and serving models to generate predictions or responses, but it does not inherently include grounding capabilities to restrict answers to a specific knowledge base; it would require additional integration with a retrieval system. Option C is wrong because Vertex AI Pipelines is an orchestration service for building and managing ML workflows, not a feature for grounding model responses or preventing hallucinations. Option D is wrong because Cloud Functions is a serverless compute service for running event-driven code, and while it could be used to build a custom retrieval pipeline, it is not a native Vertex AI feature for grounding and does not provide the built-in retrieval and citation mechanisms needed to ensure answers come only from the provided data.

2
Multi-Selecteasy

Which TWO features are available in Vertex AI Studio for prompt engineering? (Choose two.)

Select 2 answers
A.Side-by-side comparison of model outputs
B.One-click deployment to a Vertex AI endpoint
C.Ability to test prompts with different model parameters (temperature, top_p)
D.Fine-tuning models directly in the interface
E.Building conversational agents with drag-and-drop
AnswersA, C

Allows output comparison.

Why this answer

Vertex AI Studio provides a side-by-side comparison feature that allows prompt engineers to evaluate outputs from multiple model configurations or parameter settings simultaneously. This enables direct visual comparison of responses, helping to identify the most effective prompt phrasing or parameter combination without manual switching.

Exam trap

The trap here is that candidates may confuse Vertex AI Studio's prompt engineering features with those of Vertex AI Agent Builder or Vertex AI Model Registry, leading them to select options like one-click deployment or drag-and-drop agent building that belong to separate services.

3
MCQmedium

A media company wants its editors to summarize long internal research documents. Legal requires that no document content be used to train or improve any model, and that data stays within the company's Google Cloud project. Which capability should the company verify before adopting a Gemini-based solution on Vertex AI?

A.That prompts and responses are excluded from model training and covered by Google Cloud's data processing terms.
B.That the summarization endpoint supports a 1-million-token context window for the longest documents.
C.That the company's editors can access the Gemini web app with their personal Google accounts for convenience.
D.That the model can be fine-tuned on the research documents so summaries match the editors' preferred style.
AnswerA

Vertex AI's terms state that customer prompts and responses are not used to train or improve Google's foundation models, and customer data is processed under the Google Cloud Data Processing Addendum. Verifying this directly satisfies legal's two requirements: no training use and residency within the company's Google Cloud project boundary, making it the decisive check before adoption.

Why this answer

Legal's requirements are about how data is governed, not about model features. Confirming that prompts and responses are excluded from training and processed under enterprise data terms addresses both the no-training condition and the residency condition. Fine-tuning contradicts the no-training rule, context window size is irrelevant to governance, and personal accounts bypass the enterprise controls the company depends on.

Exam trap

The trap here is treating a technical capability such as a large context window or fine-tuning as if it answered a data-governance question, when the real issue is contractual training exclusion and project-bound processing.

4
MCQhard

An organization is using Vertex AI to fine-tune a large language model. They notice training is taking longer than expected and cost is increasing. Which action is most likely to reduce training time and cost without significantly impacting model quality?

A.Increase the number of training steps
B.Increase the batch size
C.Use a higher learning rate
D.Enable mixed-precision training (bfloat16)
AnswerD

bfloat16 mixed-precision training halves memory bandwidth and uses tensor cores for faster matrix operations, cutting training time and compute cost. It preserves numerical range well enough that model quality remains largely unaffected, unlike aggressive precision reduction.

Why this answer

Mixed-precision training with bfloat16 reduces memory usage and accelerates computation by using half the bits of standard float32, which directly decreases training time and cost on TPUs and modern GPUs. Vertex AI supports bfloat16 natively on TPU v3+ and A100 GPUs, and for many large language models, this precision preserves model quality because bfloat16 retains the same exponent range as float32, avoiding underflow issues common with float16.

Exam trap

A common pitfall is assuming increasing batch size or learning rate will speed up training, but this can destabilize training or degrade quality. Mixed-precision training (bfloat16) directly reduces computation without the risk of underflow, making it the optimal choice for Vertex AI.

How to eliminate wrong answers

Option A is wrong because increasing the number of training steps would lengthen training time and increase cost, directly opposing the goal. Option B is wrong because increasing batch size can reduce the number of weight updates per epoch, but it often requires tuning the learning rate and may degrade convergence or model quality if pushed too high, and it does not directly address the computational bottleneck of precision. Option C is wrong because using a higher learning rate risks training instability, divergence, or poor convergence, which can significantly harm model quality and may require additional tuning steps, negating any time savings.

5
MCQhard

A company is using Gemini Pro for code generation. They want to ensure that the generated code does not contain security vulnerabilities. Which approach should they implement?

A.Enable grounding with security scanning tools
B.Use the Vertex AI Codey API with safety settings
C.Implement a human-in-the-loop review with automated scanning
D.Use a custom safety attribute filter
AnswerC

Automated scanning catches known vulnerability patterns such as injection flaws, while human review judges context and logic that scanners miss. Combining both satisfies the requirement that generated code contain no security vulnerabilities, since neither control alone covers the full risk surface.

Why this answer

Combining human-in-the-loop review with automated scanning directly addresses the need to catch security vulnerabilities in AI-generated code. Human reviewers can identify logic flaws and context-specific risks that automated tools miss, while automated scanners provide consistent, rapid detection of known vulnerability patterns (e.g., OWASP Top 10). This layered approach is a best practice for production-grade code generation with Gemini Pro, as it mitigates the inherent limitations of relying solely on AI safety filters or static analysis.

Exam trap

The trap here is that candidates confuse content safety filters (which block toxic or harmful text) with code security vulnerability scanning, leading them to incorrectly select options that rely on Vertex AI safety settings or grounding, which are not designed to detect code-level security flaws.

How to eliminate wrong answers

Option A is wrong because grounding with security scanning tools typically refers to grounding model outputs against external data sources (e.g., enterprise databases) to improve factual accuracy, not to scan generated code for vulnerabilities; it does not replace dedicated vulnerability scanning. Option B is wrong because the Vertex AI Codey API's safety settings are designed to filter harmful or toxic content (e.g., hate speech, violence) in the generated text, not to detect security vulnerabilities like SQL injection or buffer overflows in code. Option D is wrong because a custom safety attribute filter in Vertex AI is used to block outputs based on categories like toxicity or harassment, not to perform code-specific security analysis; it lacks the semantic understanding needed to identify vulnerabilities in code logic.

6
MCQmedium

A software company wants to build a generative AI application that can answer questions based on its internal documentation. The documentation is stored in Google Drive and Confluence. They need a managed solution that can index these sources, provide relevant answers with citations, and integrate with their existing identity provider for access control. Which Google Cloud service should they use?

A.Vertex AI Search
B.Vertex AI Feature Store
C.Cloud SQL
D.Vertex AI Pipelines
AnswerA

Vertex AI Search is a managed service that can index data from various sources including Google Drive and Confluence, and it provides generative AI-powered answers with citations. It supports integration with identity providers for access control, ensuring users only see documents they are authorized to access. This matches all the requirements.

Why this answer

Vertex AI Search is a fully managed service that can connect to Google Drive and Confluence, index the content, and provide generative AI answers with citations. It also integrates with identity providers to enforce document-level access control, making it the ideal solution for the software company's needs.

Exam trap

The trap here is assuming that any Vertex AI service can handle search; only Vertex AI Search is designed for enterprise document indexing and question answering.

7
MCQeasy

A developer is using Vertex AI Studio to prototype a chat application. They want to provide the model with a system instruction to set the tone and style. How should they configure this in the Vertex AI Studio interface?

A.Add the instruction as part of the prompt text
B.Set the temperature parameter to a high value
C.Use the 'System Instruction' field in the model configuration
D.Add the instruction in the 'Context' parameter
AnswerC

The System Instruction field accepts a dedicated prompt that persists across turns, defining the assistant's persona, tone and style independently of user messages. Entering the desired tone there satisfies the requirement to set behaviour at the model configuration level rather than restating it in every user prompt.

Why this answer

Vertex AI Studio provides a dedicated 'System Instruction' field in the model configuration panel, which allows developers to set the tone, style, and behavioral guidelines for the model without mixing them into the user prompt. This field is specifically designed to hold system-level instructions that are prepended to the conversation context, ensuring consistent behavior across multiple turns.

Exam trap

The trap here is that candidates often confuse the 'System Instruction' field with the 'Context' parameter, mistakenly thinking both serve the same purpose, but the 'Context' parameter is designed for providing background knowledge or few-shot examples, not for setting persistent behavioral instructions.

How to eliminate wrong answers

Option A is wrong because adding the instruction as part of the prompt text would mix system-level guidance with user input, making it harder to maintain consistency and potentially causing the model to treat the instruction as part of the conversation rather than a persistent directive. Option B is wrong because the temperature parameter controls randomness in output generation, not the tone or style; a high temperature increases creativity and variability but does not enforce a specific behavioral instruction. Option D is wrong because the 'Context' parameter in Vertex AI Studio is used to provide background information or examples for grounding the model, not for setting system-level behavioral instructions like tone or style.

8
MCQmedium

A company is using Google Cloud's generative AI offerings to build a customer-facing application. They need to ensure that the AI-generated content complies with their brand guidelines and does not produce harmful or inappropriate responses. They also want to monitor and filter content in real-time. Which Google Cloud feature should they use?

A.Vertex AI safety filters and attributes
B.Cloud Armor
C.Identity and Access Management (IAM)
D.Cloud Data Loss Prevention (DLP)
AnswerA

Vertex AI provides configurable safety filters and attributes that can block or flag harmful content in real-time. These can be customized to align with brand guidelines and compliance requirements. They are integrated into the model API, allowing real-time monitoring and filtering of generated content.

Why this answer

Vertex AI safety filters and attributes are the correct choice because they are specifically designed to detect and block harmful content in generative AI outputs. They can be configured to enforce brand-specific guidelines and are applied in real-time during model inference. The other options are security or data protection services that do not address content moderation for generative AI.

Exam trap

The trap here is confusing general security services like Cloud Armor or IAM with content moderation features tailored for generative AI outputs.

9
MCQhard

A financial analytics firm is building a generative AI application that must analyze long earnings call transcripts and produce summaries. The transcripts often exceed 200,000 tokens, and the firm wants to minimize cost while maintaining high accuracy. They plan to use Gemini models on Vertex AI. Which approach should they take?

A.Use gemini-1.0-pro and rely on its built-in automatic long-document summarization feature.
B.Split the transcript into 1,000-token chunks and summarize each chunk separately, then concatenate the summaries.
C.Fine-tune a smaller Gemini model on the firm's past transcripts and use it for summarization.
D.Use gemini-1.5-pro with its long context window and pass the entire transcript in a single request.
AnswerD

Gemini 1.5 Pro supports a context window of up to 2 million tokens, which can accommodate transcripts exceeding 200,000 tokens in a single request. This avoids the complexity and potential accuracy loss of chunking and summarization pipelines, and it often reduces overall cost by eliminating multiple inference calls and preprocessing steps.

Why this answer

Gemini 1.5 Pro's extended context window allows the entire transcript to be processed in one request, preserving cross-references and reducing the need for complex chunking pipelines. Chunking loses context, gemini-1.0-pro lacks the required context length, and fine-tuning does not increase context capacity, so the long-context model is the most accurate and cost-effective choice.

Exam trap

The trap here is believing that fine-tuning can overcome a model's context window limitation, when fine-tuning only adapts behavior and does not expand the maximum input length.

10
MCQmedium

A company is building a customer support chatbot using Vertex AI Agent Builder. They want the agent to answer questions based on internal knowledge base documents stored in Cloud Storage. Which feature should they configure to ensure the agent can retrieve relevant information from these documents?

A.Deploy the agent to a Vertex AI endpoint
B.Fine-tune a Gemini model on the knowledge base
C.Enable grounding with a data store
D.Configure a safety filter to block irrelevant queries
AnswerC

Grounding with a data store indexes the Cloud Storage documents and retrieves relevant passages at query time, anchoring responses in that content. This satisfies the requirement to answer from internal knowledge base documents rather than the model's parametric memory.

Why this answer

Vertex AI Agent Builder uses grounding to connect the agent to external data sources, such as documents stored in Cloud Storage. By enabling grounding with a data store, the agent can retrieve and reference relevant information from the knowledge base documents in real time, ensuring accurate and context-aware responses without requiring model retraining.

Exam trap

Candidates often confuse fine-tuning with grounding. They mistakenly choose fine-tuning (Option B) assuming the model must be retrained on the knowledge base, but the correct approach for retrieval-based Q&A is grounding with a data store.

How to eliminate wrong answers

Option A is wrong because deploying the agent to a Vertex AI endpoint is about making the agent accessible for inference, not about connecting it to a knowledge base for retrieval. Option B is wrong because fine-tuning a Gemini model on the knowledge base would adapt the model's weights to the specific data, which is unnecessary and inefficient for retrieval-based tasks; Vertex AI Agent Builder uses retrieval-augmented generation (RAG) via grounding instead. Option D is wrong because configuring a safety filter blocks harmful or irrelevant queries but does not enable the agent to retrieve information from the knowledge base documents.

11
MCQmedium

A fintech startup is building a generative AI application that generates personalized investment advice based on user profiles and market data. They are using Vertex AI Agent Builder to create an agent that retrieves information from a BigQuery table containing user data and from a real-time market data API. The agent needs to ensure that responses comply with financial regulations, meaning the model must not give specific stock recommendations unless the user explicitly requests them after disclaimers. The team has implemented grounding with both sources. During testing, the agent sometimes spontaneously suggests buying a particular stock without being asked, which could lead to regulatory issues. The team wants to enforce strict control over the agent's behavior. What should the team do?

A.Increase the safety filter sensitivity to block any financial recommendations
B.Add more historical data to the BigQuery table to improve grounding accuracy
C.Implement a custom system instruction that explicitly prohibits unsolicited stock recommendations and requires a disclaimer before any advice
D.Fine-tune the model on a dataset of compliant conversations
AnswerC

A custom system instruction directly constrains the model's generative behaviour, prohibiting unsolicited stock recommendations and mandating a disclaimer before advice. This satisfies the stem's regulatory constraint by enforcing deterministic policy at inference time, unlike grounding, which only supplies factual context and cannot govern what the agent chooses to say.

Why this answer

System instructions in Vertex AI Agent Builder allow you to define strict behavioral rules that the agent must follow, such as prohibiting unsolicited stock recommendations and requiring a disclaimer before any advice. This directly addresses the regulatory compliance issue by enforcing a policy at the agent's instruction layer, which overrides any learned or grounded behavior. Unlike other options, this approach provides explicit, enforceable control without altering data sources or model training.

Exam trap

The trap here is that candidates often confuse grounding (data retrieval) with behavioral control (system instructions), assuming that better data or safety filters can enforce compliance, when in fact only explicit instructions in the agent's configuration can enforce such nuanced policies.

How to eliminate wrong answers

Option A is wrong because increasing safety filter sensitivity would block all financial recommendations, including compliant ones after disclaimers, which breaks the required user-requested flow and is too blunt for nuanced regulatory compliance. Option B is wrong because adding more historical data to the BigQuery table improves grounding accuracy but does not prevent the agent from spontaneously generating unsolicited stock recommendations; grounding only ensures factual retrieval, not behavioral constraints. Option D is wrong because fine-tuning the model on a dataset of compliant conversations may reduce but not eliminate unsolicited recommendations, as the model can still generalize or hallucinate; it also requires significant effort and does not provide a deterministic, enforceable rule like system instructions do.

12
MCQhard

A multinational corporation is using Vertex AI to generate multilingual customer support responses. They have fine-tuned the Gemini model on support tickets in English and now want to extend to 10 additional languages. The fine-tuning dataset for new languages is small (1000 tickets each). During evaluation, the model performs well for common languages (Spanish, French) but poorly for languages like Finnish and Thai. The team needs to improve performance for low-resource languages. They have budget constraints and cannot collect more data quickly. Which approach should they take?

A.Switch to Vertex AI Codey API for generating responses in all languages.
B.Use a multilingual foundation model and fine-tune with cross-lingual transfer learning techniques.
C.Deploy separate fine-tuned models for each language.
D.Collect more training data for low-resource languages via crowdsourcing.
AnswerB

Cross-lingual transfer learning leverages the multilingual foundation model's shared representations, so the 1,000-ticket datasets per language transfer knowledge from the English fine-tune. This lifts Finnish and Thai performance without collecting more data, satisfying the budget constraint.

Why this answer

Using a multilingual foundation model (like Gemini's multilingual variant) with cross-lingual transfer learning leverages the model's pre-trained knowledge across languages, allowing it to generalize from high-resource languages (Spanish, French) to low-resource ones (Finnish, Thai) even with small fine-tuning datasets. This approach is budget-friendly as it avoids separate models or costly data collection, and it directly addresses the performance gap by sharing linguistic patterns across languages.

Exam trap

The trap here is that candidates often assume more data (Option D) or separate models (Option C) are the only solutions, ignoring that cross-lingual transfer learning can effectively bootstrap low-resource languages from high-resource ones without additional data collection.

How to eliminate wrong answers

Option A is wrong because the Vertex AI Codey API is designed for code generation, not multilingual customer support responses, and switching to it would not improve performance for low-resource languages. Option C is wrong because deploying separate fine-tuned models for each language multiplies cost and maintenance overhead, and with only 1000 tickets per language, each model would suffer from the same data scarcity issue without cross-lingual benefits. Option D is wrong because the team has budget constraints and cannot collect more data quickly, making crowdsourcing infeasible in the short term, and it does not address the underlying need for transfer learning.

13
MCQmedium

A media company wants to automatically generate concise summaries of news articles using a generative AI model. They need a fully managed service that provides access to Google's foundation models without requiring machine learning expertise. Which Google Cloud offering should they choose?

A.Vertex AI Model Garden
B.Generative AI on Vertex AI (Gemini models)
C.AutoML Natural Language
D.Cloud Translation API
AnswerB

Generative AI on Vertex AI provides access to Gemini models through a fully managed service, allowing users to generate summaries via API calls without ML expertise. It offers pre-trained models that can be used out-of-the-box, exactly matching the media company's need for automatic summarization of news articles.

Why this answer

Generative AI on Vertex AI offers Gemini models as a fully managed service, enabling users to generate summaries via simple API calls without ML expertise. It provides pre-trained models that excel at summarization tasks. The other options are either model catalogs requiring configuration, custom ML training services, or specialized translation tools that do not perform summarization.

Exam trap

The trap here is confusing Vertex AI Model Garden, which hosts models but requires deployment, with the fully managed generative AI service that provides direct access to Gemini models.

14
MCQhard

A global e-commerce company is using Vertex AI to build a generative AI chatbot for customer support. The chatbot is powered by the Gemini 1.5 Pro model and uses a vector search index for retrieval-augmented generation (RAG) over product documentation. The company has deployed the application in four regions (us-central1, europe-west4, asia-east1, and australia-southeast1) using a multi-region deployment with a global endpoint. The application is critical and requires high availability with a target latency of under 500ms for the RAG pipeline. Recently, users in Australia are experiencing inconsistent latency spikes, with response times exceeding 2 seconds during peak hours. The team suspects that the issue is related to the vector search index's replication and serving configuration. The index has 10 million embeddings with a dimension of 768. It is stored in a single regional bucket in us-central1, and the vector search index endpoint is deployed in all four regions with the same deployed index ID. The team is using the default configuration for index updates and serving. Which action should the team take to resolve the latency issue for Australian users?

A.Move the vector search index to a multi-regional Cloud Storage bucket (e.g., 'us') to reduce latency for index updates.
B.Create a new regional bucket in australia-southeast1 and store a copy of the index there, then redeploy the vector search index endpoint to use the local bucket.
C.Deploy a separate vector search index endpoint for each region with its own index copy stored in a regional bucket in that region.
D.Increase the number of replicas for the vector search index in all regions to improve throughput and reduce latency.
AnswerB

Creating a local bucket in australia-southeast1 and storing a copy of the index allows the endpoint in that region to serve directly from local storage, significantly reducing read latency and resolving the issue.

Why this answer

The latency issue for Australian users is likely caused by the vector search index serving from a single regional bucket in us-central1. When users in Australia query the index, the nearest endpoint must read the index data from a bucket far away, leading to high latency during peak hours. Option B solves this by creating a local bucket in australia-southeast1 and storing a copy of the index there.

The endpoint in australia-southeast1 can then serve directly from the local bucket, reducing read latency. Moving to a multi-regional bucket (Option A) does not help because the bucket is still in the US, and Australia still has to fetch from a distant region. Option C would also work but is more complex and costlier than simply using a local bucket with the same endpoint.

Option D addresses throughput but not the geographic distance causing latency.

Exam trap

Google Cloud often tests the misconception that using a multi-regional bucket solves cross-region latency for serving, when in fact a multi-regional bucket does not bring data closer to all regions. The correct fix is to store a copy of the index locally in the affected region for serving.

How to eliminate wrong answers

Option B is wrong because creating a separate regional bucket in australia-southeast1 and storing a copy of the index there does not solve the update propagation issue; the index must be rebuilt and streamed from the source bucket (us-central1) to the new bucket, and the endpoint still references the original deployed index ID, so the local copy is not automatically used. Option C is wrong because deploying separate endpoints per region with local index copies increases operational complexity and cost, and does not address the root cause of cross-region latency during index updates; the global endpoint already routes to the nearest region, but the index data still originates from us-central1. Option D is wrong because increasing the number of replicas improves throughput and query handling within a region, but does not reduce the latency caused by cross-region data transfer during index updates; the bottleneck is the update propagation, not query processing capacity.

15
Multi-Selecteasy

Which TWO of the following are capabilities of Vertex AI Model Garden? (Choose 2)

Select 2 answers
A.Generate code snippets for common programming tasks.
B.Ability to generate images from text descriptions.
C.Deploy custom container images for model serving.
D.Access to a curated set of foundation models like PaLM and Gemini.
E.Ability to fine-tune and deploy foundation models.
AnswersD, E

Model Garden provides a curated catalogue giving discoverability and one-place access to foundation models such as PaLM and Gemini, including Google, open-source and third-party options. This satisfies the stem's capability requirement by covering the browse-and-select function rather than training or deployment.

Why this answer

Option D is correct because Vertex AI Model Garden provides access to a curated catalog of foundation models, including Google's PaLM and Gemini models, which users can discover and use directly. Option E is correct because Model Garden allows users to fine-tune and deploy foundation models, supporting customization and serving within Vertex AI. Option A is incorrect because generating code snippets is a generative AI application capability, not a core capability of Model Garden itself.

Option B is incorrect because text-to-image generation is a specific model use case, not a defining capability of Model Garden. Option C is incorrect because deploying custom container images is a Vertex AI Prediction feature, not a Model Garden capability.

Exam trap

The trap here is that candidates confuse the capabilities of Vertex AI Model Garden (model discovery, access, and deployment) with the capabilities of the underlying models themselves (e.g., code generation or image generation), or with other Vertex AI services like Prediction or Endpoints for custom container deployments.

16
MCQeasy

A startup wants to experiment with Gemini prompts in a browser-based console, compare responses from different models, and save promising prompts for later reuse, without provisioning infrastructure. Which Google Cloud offering best fits this need?

A.BigQuery ML
B.Vertex AI Studio
C.Cloud Vision API
D.Document AI
AnswerB

Vertex AI Studio is a browser-based environment for designing, testing, and saving prompts against Gemini and other models, letting users compare outputs and iterate quickly with no infrastructure to manage. It directly matches the desire to experiment interactively and retain useful prompts, making it the appropriate starting point for this exploration phase.

Why this answer

Vertex AI Studio provides an interactive, code-light workspace for writing prompts, comparing model outputs, and storing prompts for reuse, which aligns precisely with the startup's experimentation goal. The other services are specialized for data warehousing, document extraction, or image analysis and lack the prompt design and comparison capabilities, so they cannot deliver the described experience.

Exam trap

The trap here is confusing a prompt prototyping console with services that merely process data, when only Vertex AI Studio is built for interactive prompt design and reuse.

17
MCQeasy

A marketing agency wants to use Vertex AI to automatically generate social media posts for clients. They plan to use the Gemini API with few-shot prompting. The agency's developers have limited experience with generative AI and want the fastest way to prototype and iterate on prompts. They are already using Google Cloud for other services. Which approach should they take to quickly develop and test prompts?

A.Use a third-party platform like OpenAI Playground and migrate later.
B.Use Google Cloud Shell to invoke the model via curl commands.
C.Use Vertex AI Studio (Gen AI Studio) to design and test prompts interactively.
D.Write Python scripts using the Vertex AI SDK and run them in Airflow.
AnswerC

Vertex AI Studio provides an interactive console for designing, testing and iterating on prompts against Gemini models without writing code, directly satisfying the developers' need for the fastest prototyping route given their limited generative AI experience. It also integrates with their existing Google Cloud environment, so validated prompts can later be exported to code.

Why this answer

Vertex AI Studio (Gen AI Studio) is the correct choice because it provides a no-code, interactive environment specifically designed for rapid prompt engineering and iteration with Gemini models. It allows developers with limited generative AI experience to test few-shot prompts, adjust parameters, and see results immediately without writing code, making it the fastest path from concept to working prototype within the Google Cloud ecosystem.

Exam trap

The trap here is that candidates may confuse 'fastest to prototype' with 'most familiar tool' (like curl or Python scripts), overlooking that Vertex AI Studio is purpose-built for interactive, no-code prompt engineering within Google Cloud.

How to eliminate wrong answers

Option A is wrong because using a third-party platform like OpenAI Playground introduces unnecessary migration effort, potential API incompatibilities, and does not leverage the agency's existing Google Cloud investment or the Gemini API's specific capabilities. Option B is wrong because Google Cloud Shell with curl commands is a low-level, non-interactive approach that lacks the visual prompt design, parameter tuning, and example management features of Vertex AI Studio, making it slower and more error-prone for iterative prototyping. Option D is wrong because writing Python scripts with the Vertex AI SDK and running them in Airflow is a production-oriented, code-heavy approach that requires significant development effort and is not suitable for rapid prototyping and iteration by developers with limited generative AI experience.

18
MCQhard

An organization is using Vertex AI Gemini API for a multimodal chatbot. They notice that the model sometimes provides incorrect information with high confidence. They want to reduce hallucinations without retraining the model. What is the most effective approach?

A.Provide ground-truth context from a knowledge base using grounding
B.Increase the temperature parameter to make the model more creative
C.Reduce the maximum output tokens to force concise answers
D.Adjust safety settings to filter uncertain responses
AnswerA

Grounding retrieves authoritative content from a knowledge base and injects it as context, so responses are anchored to verified facts rather than parametric memory. This reduces hallucination without retraining, satisfying the stem's constraint of no model fine-tuning.

Why this answer

Grounding with a knowledge base is the most effective approach because it forces the model to base its responses on verified, external data rather than relying solely on its parametric knowledge. By providing ground-truth context via Vertex AI's grounding feature, the model can cross-reference its outputs with authoritative sources, significantly reducing the likelihood of hallucinations. This method directly addresses the root cause of incorrect high-confidence responses without requiring retraining.

Exam trap

A common misconception is that adjusting generation parameters (like temperature or token limits) can fix factual accuracy issues, when in reality only grounding or retrieval-augmented generation can address hallucinations without retraining.

How to eliminate wrong answers

Option B is wrong because increasing the temperature parameter makes the model more creative and random, which actually increases the risk of hallucinations by encouraging less deterministic outputs. Option C is wrong because reducing the maximum output tokens only truncates the response length; it does not improve factual accuracy and may lead to incomplete or misleading answers. Option D is wrong because safety settings filter content based on harm categories (e.g., hate speech, violence), not factual correctness; they cannot detect or prevent hallucinations about factual information.

19
Multi-Selectmedium

A healthcare organization is evaluating Google Cloud's generative AI offerings for building a patient triage assistant. They need a model that can process text and images from patient records and generate text responses. They also require the ability to fine-tune the model on their own data. Which two Google Cloud services or features should they use? (Choose two.)

Select 2 answers
A.Vertex AI Pipelines
B.Vertex AI Fine-Tuning
C.Gemini on Vertex AI
D.Vertex AI Model Garden
E.Vertex AI Feature Store
AnswersB, C

Vertex AI Fine-Tuning allows organizations to customize foundation models like Gemini on their own datasets. This is essential for the healthcare organization to tailor the model to their patient triage domain, improving accuracy and relevance for their specific use case.

Why this answer

Gemini on Vertex AI provides the multimodal capabilities to process text and images, and Vertex AI Fine-Tuning enables customization on the organization's data. Together, they form the foundation for a tailored patient triage assistant that can handle diverse patient records.

Exam trap

The trap here is selecting Vertex AI Model Garden as a model, but it is a catalog, not a model; the actual model and tuning service are needed.

20
MCQmedium

A machine learning engineer is deploying a large generative model on Vertex AI. The model requires a GPU with high memory. Which machine configuration should they choose?

A.c2-standard-16 with no GPU
B.a2-highgpu-4g with 4 A100 GPUs
C.n1-standard-4 with a single T4 GPU
D.n2-standard-8 with a single P4 GPU
AnswerB

The a2-highgpu-4g machine type pairs four NVIDIA A100 GPUs, each with 40 GB of high-bandwidth memory, satisfying the high GPU memory constraint for serving a large generative model. Smaller accelerator configurations would lack sufficient VRAM.

Why this answer

The a2-highgpu-4g machine series is specifically designed for large-scale GPU-accelerated workloads, offering 4 NVIDIA A100 GPUs with 40GB of high-bandwidth memory (HBM2e) each, totaling 160GB of GPU memory. This configuration provides the high memory capacity required for training or serving large generative models, such as LLMs or diffusion models, which often exceed the memory limits of smaller GPUs.

Exam trap

The trap here is that candidates may choose a cheaper or single-GPU option (like C or D) without calculating the total GPU memory needed, or mistakenly think a CPU-only instance (A) can handle GPU-accelerated workloads, ignoring that large generative models require both high GPU memory and parallel processing.

How to eliminate wrong answers

Option A is wrong because c2-standard-16 is a compute-optimized machine without any GPU, which cannot provide the GPU memory needed for large generative models. Option C is wrong because n1-standard-4 with a single T4 GPU offers only 16GB of GPU memory, insufficient for large models that require tens or hundreds of gigabytes. Option D is wrong because n2-standard-8 with a single P4 GPU provides only 8GB of GPU memory, far below the requirements for large generative models and lacks the parallelism of multiple GPUs.

21
MCQhard

A company is using Vertex AI Gemini API to analyze customer feedback. They notice that the model occasionally generates offensive content. They have already set safety settings to block high-probability harmful content. What additional step should they take to further reduce offensive outputs?

A.Set the temperature to 0.0
B.Adjust safety settings to block medium-probability harmful content
C.Enable context caching
D.Fine-tune the model on customer feedback data
AnswerB

Safety settings operate on probability thresholds per harm category, so lowering the block threshold from high to medium catches harmful content the model would otherwise emit. This tightens filtering beyond the existing high-probability configuration, directly reducing offensive outputs from the Gemini API.

Why this answer

The company has already blocked high-probability harmful content, but offensive outputs can still occur at lower probability thresholds. By adjusting safety settings to block medium-probability harmful content, they tighten the filter to catch more borderline cases without requiring model retraining or sacrificing output diversity. This leverages Vertex AI's configurable safety filters, which operate on likelihood categories (e.g., high, medium, low) rather than just binary blocking.

Exam trap

The trap here is that candidates assume fine-tuning (Option D) is the default fix for any output quality issue, but safety filtering is a separate, configurable layer that should be tuned before retraining, and temperature (Option A) is often mistakenly thought to control safety when it only controls randomness.

How to eliminate wrong answers

Option A is wrong because setting temperature to 0.0 makes the model deterministic and reduces creativity, but it does not filter or block offensive content; temperature controls randomness in token selection, not safety. Option C is wrong because context caching improves latency and cost for repeated prompts by storing context, but it has no effect on content safety or filtering harmful outputs. Option D is wrong because fine-tuning on customer feedback data could inadvertently reinforce biases or offensive patterns in the data, and it does not directly address safety filtering; safety settings are a separate, configurable layer that should be adjusted first.

22
MCQmedium

A media company uses Vertex AI to generate video captions. The generated captions sometimes contain factual errors about named entities (e.g., actor names). Which technique would most likely reduce these errors?

A.Enable response caching
B.Increase the temperature parameter
C.Use Vertex AI grounding with a knowledge base of verified entities
D.Decrease top_p to 0.3
AnswerC

Grounding with a verified entity knowledge base constrains generation to retrieved, authoritative facts, so named entities come from the knowledge base rather than the model's parametric memory. This directly targets the factual-error constraint in the stem, reducing hallucinated actor names without retraining or prompt engineering.

Why this answer

Vertex AI grounding connects the model to a knowledge base of verified entities, allowing it to retrieve authoritative facts during generation. This reduces hallucinations about named entities by constraining outputs to validated data rather than relying solely on the model's parametric knowledge.

Exam trap

The trap here is that candidates confuse techniques that control output randomness (temperature, top_p) with techniques that improve factual accuracy, overlooking the fundamental need for external knowledge retrieval via grounding.

How to eliminate wrong answers

Option A is wrong because response caching stores previous outputs for reuse, which does not correct factual errors—it may even propagate them. Option B is wrong because increasing the temperature parameter increases randomness in token selection, making factual errors more likely, not less. Option D is wrong because decreasing top_p to 0.3 narrows the sampling pool to only the most probable tokens, which can reduce creativity but does not address factual accuracy about named entities—it still relies on the model's internal knowledge, which may be incorrect.

23
MCQhard

An enterprise architect must let a fleet of internal applications call Gemini models on Google Cloud while enforcing VPC Service Controls perimeters, customer-managed encryption keys, and audit logging tied to the organization's existing IAM. Which Google Cloud offering should the architect standardize on?

A.Gemini API via Google AI Studio
B.Gemini Enterprise with data connectors
C.Vertex AI with Gemini models
D.Gemini for Google Workspace
AnswerC

Vertex AI exposes Gemini models through endpoints that integrate with Cloud IAM, VPC Service Controls, customer-managed encryption keys, and Cloud Audit Logs. Because the architect must apply perimeter controls, key management, and organization-scoped identity to application API calls, standardizing on Vertex AI delivers the governance surface the scenario demands.

Why this answer

Vertex AI is the Google Cloud platform where Gemini models are served to applications under the same enterprise controls as other Cloud services, including IAM, VPC Service Controls, customer-managed encryption keys, and audit logging. Because the architect needs perimeter enforcement and key management for API traffic, the Vertex AI model endpoints are the appropriate foundation rather than consumer-style APIs or assistant products.

Exam trap

The trap here is treating all Gemini access paths as equally governed, when the developer-oriented Gemini API lacks the enterprise perimeter and key controls that Vertex AI endpoints provide.

24
MCQhard

A healthcare technology company is using Vertex AI to build a generative AI assistant that answers patient questions about medications. They must ensure the assistant does not provide medical advice and adheres to safety guidelines. Which Google Cloud feature should they use to enforce these constraints?

A.Vertex AI Explainable AI
B.Vertex AI safety filters and configurable safety settings
C.Vertex AI Pipelines
D.Vertex AI Model Monitoring
AnswerB

Vertex AI provides safety filters and configurable safety settings that can block or filter harmful content, including medical advice that could be dangerous. By adjusting thresholds and categories, the company can enforce constraints to prevent the assistant from giving inappropriate medical guidance, aligning with safety guidelines.

Why this answer

Vertex AI safety filters and configurable safety settings allow developers to set thresholds and categories that block harmful or inappropriate content, including medical advice. This is the only option that provides real-time enforcement of safety guidelines. Monitoring, Explainable AI, and Pipelines serve different purposes and cannot enforce content restrictions.

Exam trap

The trap here is assuming that monitoring or explainability features can enforce safety constraints, when only safety filters and settings actively block or filter content during generation.

25
MCQeasy

A media production studio wants to create short video clips from text prompts for social media content. They need a Google Cloud generative AI model that can generate videos from text. Which offering should they use?

A.Chirp
B.Gemini
C.Imagen
D.Veo
AnswerD

Veo is Google's generative AI video model available through Vertex AI, designed to create short video clips from text prompts. It supports cinematic quality and temporal consistency, making it suitable for social media content. The studio can generate clips directly from descriptive text without manual editing. This directly matches the requirement for text-to-video generation.

Why this answer

Veo is the Google Cloud generative AI model specifically built for text-to-video generation. It enables users to create short video clips from textual descriptions, which aligns perfectly with the media studio's goal of producing social media content. Other models like Imagen and Gemini handle images or text but not video output, so Veo is the correct offering.

Exam trap

The trap here is confusing Imagen, which generates images, with Veo, which generates videos, because both are creative media models on Vertex AI.

26
Multi-Selecthard

A healthcare company wants to use generative AI to summarize patient notes while ensuring compliance with strict data privacy regulations. They plan to use Google Cloud's Vertex AI. Which two features should they implement to protect sensitive data? (Choose two.)

Select 2 answers
A.VPC Service Controls
B.Storing data in a public Cloud Storage bucket
C.Customer-managed encryption keys (CMEK)
D.Disabling audit logging
E.Public IP addressing for Vertex AI endpoints
AnswersA, C

VPC Service Controls create a security perimeter around Google Cloud services to prevent data exfiltration. For healthcare data, this ensures that Vertex AI resources cannot be accessed from outside the authorized network. It helps comply with regulations by restricting data movement and mitigating insider threats. Implementing VPC Service Controls is a recommended practice for sensitive workloads.

Why this answer

Customer-managed encryption keys (CMEK) and VPC Service Controls are both critical for protecting sensitive healthcare data in Vertex AI. CMEK gives the organization control over encryption keys, while VPC Service Controls prevent unauthorized data exfiltration. Together, they help meet regulatory requirements for data privacy and security.

Exam trap

The trap here is thinking that default Google encryption is sufficient, but for strict compliance, customer-managed keys and network perimeters are needed.

27
MCQmedium

A university research team wants to summarize long academic papers using a Google Cloud model. They need a model that supports very large context windows so an entire paper fits in a single request. Which Gemini model characteristic should they prioritize?

A.A model with the largest number of supported languages
B.A model with the lowest latency for short prompts
C.A model version with a long context window, such as Gemini 1.5 Pro
D.A model with the highest number of output tokens per response
AnswerC

Gemini 1.5 Pro offers a long context window capable of handling very large inputs, which allows an entire academic paper to be included in one request. This directly supports summarization without chunking, preserving cross-section context that improves summary coherence.

Why this answer

Summarizing an entire academic paper in one request depends on the model's input context window. Gemini 1.5 Pro's long context capability lets the team include the full document, preserving structure and cross-references. Latency, output token limits, and language coverage affect other aspects but do not remove the need for a large input window.

Exam trap

The trap here is confusing the output token limit with the input context window, leading to a model choice that still cannot ingest the full paper.

28
MCQhard

A bank is deploying a Gemini-powered assistant on Vertex AI to answer questions about loan products. Compliance requires that every answer be traceable to an approved internal source and that the assistant refuse to answer when no approved source supports the response. Which configuration should the bank implement?

A.Raise the model temperature and add a system instruction to be accurate
B.Ground the model with Vertex AI Search over approved sources and use grounding confidence thresholds
C.Enable a higher safety filter threshold so harmful content is blocked
D.Fine-tune Gemini on historical chat logs from the bank's contact center
AnswerB

Grounding with Vertex AI Search restricts answers to retrieved passages from approved sources and returns citations for traceability. Grounding confidence scores let the bank detect weak support and configure the assistant to decline rather than speculate. This satisfies both the traceability and the refusal requirements.

Why this answer

Grounding the assistant in approved sources through Vertex AI Search makes each answer retrievable from a vetted document and provides citations for audit. Using grounding confidence thresholds lets the bank define when support is too weak and instruct the assistant to refuse. Fine-tuning, temperature changes, and safety filters do not deliver source traceability or principled abstention.

Exam trap

The trap here is equating content safety filtering with factual grounding, when refusal on unsupported questions depends on retrieval quality and confidence thresholds.

29
MCQhard

A hospital wants to build a generative AI assistant that answers clinician questions using the hospital's own clinical guidelines. The guidelines change frequently, and the hospital wants the assistant to cite the exact guideline section used. Which approach best meets these requirements while minimizing model retraining?

A.Ground the assistant on the guidelines using retrieval-augmented generation with Vertex AI Search
B.Expand the model's context window and paste all guidelines into every prompt
C.Fine-tune the Gemini model every time the clinical guidelines are updated
D.Increase the model's temperature so it recalls guideline details more accurately
AnswerA

Retrieval-augmented generation retrieves relevant guideline passages at query time and grounds the model's answer in them, so citations can point to the exact section. Because the guidelines live in an index rather than the model weights, updates require re-indexing documents, not retraining, which minimizes retraining effort while keeping answers current and attributable.

Why this answer

Grounding with retrieval-augmented generation separates knowledge from the model: guideline passages are indexed and retrieved per query, so updates mean re-indexing rather than retraining, and responses can cite the retrieved section. Fine-tuning forces retraining on every change and gives no citations, temperature affects randomness not accuracy, and stuffing all guidelines into the prompt hits context limits without reliable attribution.

Exam trap

The trap here is assuming fine-tuning is the way to add domain knowledge; when content changes often and citations are required, retrieval grounding is the appropriate pattern, not weight updates.

30
Multi-Selectmedium

A media company is evaluating Google Cloud generative AI offerings to build a production application that summarizes long articles and generates headlines. They want to use Gemini models with enterprise controls and need to understand which capabilities are provided by Vertex AI. (Choose two.)

Select 2 answers
A.Automated generation of print-ready magazine layouts from article text.
B.Access to Gemini models through a managed API with enterprise security and data governance controls.
C.Guaranteed placement of generated headlines on the publication's homepage.
D.Built-in tools for tuning and evaluating Gemini models against the company's own summarization quality criteria.
E.Automatic translation of every generated article into all supported languages without configuration.
AnswersB, D

Vertex AI provides managed access to Gemini models with enterprise-grade security, identity, and data governance controls, which is essential for a production media application handling proprietary content. It supports VPC Service Controls, customer-managed encryption keys, and audit logging. This capability aligns with the requirement to use Gemini under enterprise controls rather than through a consumer interface.

Why this answer

Vertex AI delivers managed access to Gemini with enterprise security and governance, plus tuning and evaluation tools to align models with custom summarization criteria. These two capabilities directly support a production media application that needs controlled model access and measurable quality. The remaining options describe publishing outcomes or unrealistic automation that Vertex AI does not provide.

Exam trap

The trap here is confusing content generation with content publishing, assuming Vertex AI also handles layout, homepage placement, or configuration-free translation.

31
Multi-Selectmedium

A media company wants to use Google Cloud generative AI to produce short video clips from text prompts for internal storyboarding. The team needs a managed model that can generate video from text and wants to evaluate it before committing to production. Which two Google Cloud offerings or capabilities should they use? (Choose two.)

Select 2 answers
A.Cloud Translation API
B.Document AI
C.Veo on Vertex AI
D.Vertex AI Model Garden
E.Imagen on Vertex AI
AnswersC, D

Veo is Google Cloud's generative video model available through Vertex AI, and it can create short video clips from text prompts. Using it gives the media company the text-to-video capability it needs without building or hosting a video generation model itself. It fits the storyboarding use case directly.

Why this answer

Veo on Vertex AI provides the actual text-to-video generation, while Vertex AI Model Garden gives the team a managed place to discover and evaluate that model before production adoption. Together they cover both the creative capability and the evaluation step. The remaining services address images, documents, or language translation and cannot generate video.

Exam trap

The trap here is confusing image generation with video generation, or assuming any Vertex AI model catalog entry can produce motion output.

32
MCQmedium

An e-commerce company is using Vertex AI PaLM 2 for Text (via Model Garden) to generate product descriptions. They have an existing pipeline that calls the model with a prompt including product attributes. Recently, they migrated to the Gemini API. The team notices that the Gemini model sometimes outputs descriptions that are factually inconsistent with the input (e.g., wrong color or size). This was less frequent with PaLM 2. They have not changed the prompts. What is the most likely cause and solution?

A.Revert to PaLM 2 since it was more reliable for this task.
B.Add negative prompts to discourage incorrect facts.
C.Adjust the prompt to be more explicit about adhering to the input data, and reduce the temperature.
D.Increase the model's temperature to make outputs more deterministic.
AnswerC

Lowering temperature sharpens the token probability distribution, reducing creative drift away from supplied attributes, while explicit instructions reinforce grounding on the input. Together these address the factual inconsistency introduced by Gemini's changed decoding behaviour and instruction-following, without altering the existing pipeline's product-attribute prompt structure.

Why this answer

The core issue is that the prompt, originally optimized for PaLM 2, may not be sufficiently explicit for the Gemini model's different instruction-following behavior. By making the prompt more explicit about adhering strictly to the input data and reducing the temperature (e.g., to 0.2 or lower), the model's output becomes more deterministic and less prone to hallucinating incorrect attributes. This directly addresses the factual inconsistency without changing the model family, leveraging Gemini's ability to follow detailed instructions when properly guided.

Exam trap

The trap here is that candidates assume model migration is the root cause and choose to revert (Option A), when in fact the real issue is prompt adaptation and hyperparameter tuning for the new model's behavior.

How to eliminate wrong answers

Option A is wrong because reverting to PaLM 2 ignores the fact that the prompt was not optimized for Gemini; the issue is prompt engineering and hyperparameter tuning, not model reliability. Option B is wrong because negative prompts are not a standard mechanism in Gemini or PaLM 2 for text generation; they are used in image generation models (e.g., Imagen) to avoid certain concepts, not to enforce factual consistency in text outputs. Option D is wrong because increasing temperature would make outputs more random and less deterministic, worsening the factual inconsistency problem, not solving it.

33
MCQeasy

A developer wants to generate Python code using Google Cloud's generative AI. Which model should they invoke?

A.Chirp
B.Codey
C.Imagen
D.Meena
AnswerB

Codey is Google Cloud's foundation model family purpose-built for code, trained on source code and tuned for generation, completion and debugging tasks. Invoking it satisfies the stem's Python code generation requirement, whereas general-purpose text models like PaLM or Gemini lack that code-specialised tuning.

Why this answer

Codey is Google Cloud's family of models specifically designed for code generation, completion, and chat, built on the PaLM 2 architecture and fine-tuned on code-heavy datasets. For a developer needing to generate Python code, Codey is the correct choice because it is purpose-built for code-related tasks, unlike other models that specialize in different modalities.

Exam trap

The trap here is that candidates may confuse Chirp (audio) or Imagen (image) with code generation because all are Google Cloud generative AI offerings, but each is specialized for a distinct modality, and the question explicitly asks for code generation.

How to eliminate wrong answers

Option A is wrong because Chirp is Google Cloud's speech-to-text model, designed for audio transcription, not code generation. Option C is wrong because Imagen is a text-to-image generation model, focused on creating visual content from text prompts, not code. Option D is wrong because Meena is a general-purpose conversational AI model (predecessor to LaMDA) optimized for open-domain dialogue, not for generating syntactically correct Python code.

34
MCQhard

A model response deployed on Vertex AI includes safety attributes with a toxicity score of 0.9 and an insult score of 0.3. The application must reject any prediction where the toxicity score exceeds 0.8. Based on the response, what action should the application take?

A.Reject the prediction because the toxicity score exceeds 0.8.
B.Retry the request with a lower temperature.
C.Display the prediction because the insult score is below 0.8.
D.Log the prediction but still display it.
AnswerA

The safety attributes returned by Vertex AI include a toxicity score, and the application's stated threshold is 0.8. Since the exhibited score exceeds that limit, the prediction breaches the configured policy and must be rejected before reaching the user.

Why this answer

The response from the model includes a safety attribute with a toxicity score of 0.9, which exceeds the application's threshold of 0.8. Vertex AI safety attributes provide scores for categories like toxicity, insult, and sexual content, and the application must enforce its own rejection logic based on these scores. Since the toxicity score is above the defined threshold, the application should reject the prediction to comply with safety policies.

Exam trap

Google's exam often tests the distinction between different safety attribute categories (e.g., toxicity vs. insult) and the importance of applying the correct threshold to the correct score. Candidates may mistakenly focus on a lower-scoring attribute instead of the one specified in the policy.

How to eliminate wrong answers

Option B is wrong because retrying with a lower temperature would not change the underlying safety attributes; temperature controls randomness in token generation, not the safety scores, and the model's output is already generated. Option C is wrong because the application's rejection criterion is based on the toxicity score, not the insult score; even if the insult score is below 0.8, the toxicity score of 0.9 triggers rejection. Option D is wrong because the application must reject the prediction when the toxicity score exceeds 0.8, not log and display it; logging without rejection violates the stated policy.

35
MCQmedium

A software vendor is building a product that must call a Gemini model through a stable, versioned API with enterprise controls such as VPC Service Controls, and must run on Google Cloud infrastructure. Which Google Cloud offering should the vendor use?

A.Gemini for Google Workspace
B.Google AI Studio
C.The Gemini API in Vertex AI
D.Vertex AI Feature Store
AnswerC

The Gemini API in Vertex AI exposes Google's foundation models through a versioned enterprise endpoint that supports Google Cloud controls including VPC Service Controls, IAM, and audit logging. It runs on Google Cloud infrastructure and is designed for building production applications, so it satisfies the stability and governance requirements described.

Why this answer

The Gemini API in Vertex AI is the enterprise path to Google's foundation models: it offers versioned endpoints, runs on Google Cloud, and integrates with IAM, VPC Service Controls, and audit logging. Workspace is an end-user assistant, AI Studio targets prototyping without enterprise governance, and Feature Store serves ML features rather than generative model inference.

Exam trap

The trap here is confusing the convenient prototyping API with the enterprise API; only the Vertex AI endpoint provides the versioning and Google Cloud governance controls a shipped product requires.

36
Multi-Selecthard

An organization is building a generative AI application on Vertex AI. Which THREE actions should they take to ensure responsible AI practices?

Select 3 answers
A.Disable content filtering
B.Implement human review for sensitive outputs
C.Conduct fairness evaluation
D.Create a safety policy and enforce via content filtering
E.Use only Google's foundation models
AnswersB, C, D

Human review places a person in the loop to inspect sensitive outputs before they reach users, catching harmful or inaccurate content that automated filters miss. This directly satisfies the responsible AI requirement for oversight and accountability in the generative AI application built on Vertex AI.

Why this answer

Option B is correct because implementing human review for sensitive outputs ensures that high-risk or ambiguous AI-generated content is validated by a person before it reaches users, which is a core responsible AI safeguard for generative applications on Vertex AI. Option C is correct because conducting fairness evaluation (for example, using Vertex AI's model evaluation tools to assess bias across demographic groups) helps detect and mitigate discriminatory or skewed model behavior, directly supporting responsible AI. Option D is correct because creating a safety policy and enforcing it via content filtering operationalizes responsible AI by defining prohibited content and using Vertex AI safety filters (such as configurable hate speech, harassment, and dangerous content thresholds) to block violations.

Option A is not correct because disabling content filtering removes a key safety control and increases the risk of harmful outputs, which is contrary to responsible AI. Option E is not correct because using only Google's foundation models does not by itself guarantee responsible AI; responsibility depends on evaluation, policy enforcement, monitoring, and human oversight regardless of which models are used.

Exam trap

The trap here is that candidates might think disabling content filtering (Option A) improves performance, but the exam tests that responsible AI on Google Cloud's Vertex AI requires both automated filters and human oversight, as emphasized in Google's AI Principles.

37
MCQeasy

A startup wants to embed generative AI features into their mobile app but has limited ML expertise. Which Google Cloud service is best suited for rapid integration with no ML training?

A.Vertex AI Model Garden
B.Vertex AI Agent Builder
C.Gemini API
D.Cloud Run with a custom container
AnswerC

The Gemini API provides pre-trained generative capabilities callable directly from application code, requiring no model training, tuning, or ML expertise. This satisfies the stem's constraints of limited ML expertise and rapid integration into a mobile app.

Why this answer

The Gemini API provides direct, no-code access to Google's most capable generative AI models via a simple REST API, requiring zero ML training or infrastructure setup. This makes it the fastest path for a startup with limited ML expertise to embed generative AI features like text generation, summarization, or chat into a mobile app.

Exam trap

Candidates often confuse the Gemini API with Vertex AI Model Garden. They choose Model Garden because it offers a broader platform, but the requirement for 'no ML training' and 'rapid integration' indicates that the direct API is the best fit.

How to eliminate wrong answers

Option A is wrong because Vertex AI Model Garden is a curated repository of foundation models that still requires users to deploy, fine-tune, or serve models via endpoints, which demands ML expertise and operational overhead. Option B is wrong because Vertex AI Agent Builder is designed for building conversational AI agents with grounding and tool integration, which involves more complex orchestration and is overkill for simply embedding generative AI features without ML training. Option D is wrong because Cloud Run with a custom container requires the startup to containerize and deploy their own model or inference code, which necessitates ML expertise and infrastructure management, contradicting the 'no ML training' requirement.

38
Multi-Selectmedium

A financial services firm is evaluating Google Cloud generative AI offerings for an internal knowledge assistant. They need capabilities for controlling access to model endpoints and for monitoring model usage and safety signals. (Choose two.)

Select 2 answers
A.Cloud Monitoring dashboards and alerting for Vertex AI metrics
B.BigQuery ML model export to Cloud Storage
C.Cloud Storage bucket lifecycle rules on training data
D.IAM roles and policies on Vertex AI resources
E.Vertex AI Feature Store for online serving
AnswersA, D

Cloud Monitoring collects Vertex AI metrics such as request counts, latency, and error rates, and can alert on anomalies. This supports the requirement to monitor model usage and safety-related signals, enabling the firm to detect unexpected traffic or failures and respond operationally.

Why this answer

IAM roles and policies restrict and grant access to Vertex AI endpoints and related resources, satisfying the access-control requirement. Cloud Monitoring dashboards and alerts surface usage, latency, and error signals for those endpoints, satisfying the monitoring requirement. The remaining options address model export, feature serving, and storage lifecycle, none of which cover access control or operational monitoring.

Exam trap

The trap here is selecting data or ML lifecycle tools that sound adjacent to governance but do not actually control endpoint access or report usage metrics.

39
Multi-Selectmedium

A global retailer wants to build a generative AI application on Google Cloud that answers employee questions about HR policies. The team must ensure the application uses their private policy documents, keeps answers grounded in those documents, and avoids exposing confidential content to unauthorized staff. Which two Google Cloud components should they include in the architecture? (Choose two.)

Select 2 answers
A.Vertex AI Pipelines for nightly retraining of the Gemini model on HR documents
B.Vertex AI Model Registry for versioning the HR policy document set
C.Vertex AI Search configured with the HR policy documents as a data store
D.Gemini models on Vertex AI for generating grounded responses
E.Vertex AI Feature Store for storing employee tenure and department features
AnswersC, D

Vertex AI Search can index the private HR policy documents and serve relevant passages to the model at query time, which keeps answers grounded in the company's own content. Its enterprise connectors and access controls also help restrict which documents are retrievable, supporting the confidentiality requirement for unauthorized staff.

Why this answer

A grounded HR assistant needs a retrieval layer and a generation layer. Vertex AI Search indexes the private policy documents, retrieves relevant passages per question, and supports access controls, while Gemini models on Vertex AI generate answers from those passages with grounding. Feature Store, Pipelines, and Model Registry address structured features, workflow orchestration, and model versioning, none of which provide document grounding or content access control.

Exam trap

The trap here is treating model training and versioning services as substitutes for retrieval, when grounding private documents requires an index that is queried at runtime rather than weights updated by retraining.

40
MCQhard

A financial services firm is using Vertex AI to generate investment reports. They need to ensure that the model outputs are explainable and comply with regulatory requirements. Which Vertex AI feature should they use?

A.Vertex AI Model Registry
B.Vertex Explainable AI
C.Vertex AI Safety Settings
D.Vertex AI AutoML
AnswerB

Vertex Explainable AI provides feature attributions that show which input features influenced each prediction, satisfying the regulatory requirement for explainable outputs. It integrates directly with Vertex AI models, delivering explanations alongside predictions without altering the model architecture, which meets the financial firm's compliance constraint for investment reports.

Why this answer

Vertex Explainable AI provides feature attributions and explanations for model predictions, which is essential for financial services firms that must comply with regulatory requirements like the EU's GDPR or the US SEC's model risk management guidelines. It helps auditors and stakeholders understand why a model generated a specific investment report output, ensuring transparency and accountability in AI-driven decisions.

Exam trap

Google's certification exam often tests the distinction between safety/security features and explainability features, leading candidates to confuse Vertex AI Safety Settings (which block harmful content) with the need for regulatory compliance explanations.

How to eliminate wrong answers

Option A is wrong because Vertex AI Model Registry is a centralized repository for managing and deploying models, not a tool for generating explanations or ensuring regulatory compliance. Option C is wrong because Vertex AI Safety Settings are designed to filter harmful content and enforce safety policies, not to provide explainability for model outputs. Option D is wrong because Vertex AI AutoML automates model training and deployment but does not inherently provide explainability features; you would need to integrate Explainable AI separately for that purpose.

41
MCQhard

A company deploys a Gemini model on Vertex AI for a healthcare application. They need to ensure that the model does not generate medical advice and that responses are grounded in trusted medical sources. Which combination of safety measures should they implement?

A.Enable safety filters and use Vertex AI Grounding with a labeled medical dataset
B.Use Vertex AI Grounding with a public dataset and disable safety filters
C.Enable safety filters only, without grounding
D.Fine-tune the model on a curated medical dataset and disable safety filters for faster responses
AnswerA

Safety filters block the model from producing medical advice, while Vertex AI Grounding against a labelled medical dataset constrains responses to trusted sources, jointly satisfying both the no-advice and grounded-in-trusted-sources constraints in the healthcare scenario.

Why this answer

It combines two essential safety layers: safety filters block harmful content (including medical advice), and Vertex AI Grounding anchors responses to a labeled medical dataset, ensuring factual accuracy and compliance with healthcare regulations. This dual approach prevents the model from generating unverified or dangerous medical information while maintaining relevance to trusted sources.

Exam trap

The trap here is that candidates assume fine-tuning alone is sufficient for domain-specific safety, but without grounding and safety filters, the model can still hallucinate or generate unverified medical advice, which is a key distinction Google Cloud tests in the Generative AI Leader exam.

How to eliminate wrong answers

Option B is wrong because using a public dataset for grounding introduces unverified or non-authoritative medical information, and disabling safety filters removes the critical barrier against generating harmful or unlicensed medical advice. Option C is wrong because safety filters alone cannot ensure responses are grounded in trusted medical sources; they only block explicit content but do not prevent the model from fabricating medical facts. Option D is wrong because fine-tuning on a curated dataset does not guarantee real-time grounding in trusted sources, and disabling safety filters exposes the application to generating unverified medical advice, which is unacceptable in healthcare.

42
Multi-Selecthard

A machine learning engineer is tuning a large language model on Vertex AI for question answering. They want to evaluate the model's performance before deployment. Which THREE metrics should they consider?

Select 3 answers
A.Cost per training epoch
B.F1 score
C.Exact match (EM)
D.Training time per epoch
E.ROUGE-L score
AnswersB, C, E

F1 score is the harmonic mean of precision and recall, capturing the balance between correctly retrieved answers and missed ones. For question answering it exposes whether the tuned model trades precision for coverage, which accuracy alone would hide.

Why this answer

For a question-answering evaluation on Vertex AI, F1 score (B) is correct because it measures token-level overlap between the predicted answer and the ground-truth answer, balancing precision and recall and handling partial matches. Exact match (C) is correct because it reports the percentage of predictions that exactly equal the reference answer, a standard metric in QA benchmarks such as SQuAD. ROUGE-L score (E) is correct because it measures the longest common subsequence between generated and reference text, capturing fluency and recall-oriented overlap useful for generative QA outputs.

Cost per training epoch (A) and training time per epoch (D) are operational/training-efficiency measures, not model quality metrics, so they do not evaluate performance before deployment.

Exam trap

The trap here is that candidates confuse operational metrics (like cost or training time) with evaluation metrics that directly measure model output quality, leading them to select options that are irrelevant to performance assessment.

43
MCQmedium

A media company wants its editorial staff to draft blog posts inside a web-based workspace where Gemini can summarize uploaded research PDFs, generate outlines, and cite files from the team's shared drive, all without writing code or managing any Google Cloud infrastructure. Which Google Cloud generative AI offering best fits this requirement?

A.Google Cloud Vertex AI Agent Builder with a custom tool
B.Gemini for Google Workspace
C.Gemini Enterprise with a connected data store
D.Vertex AI Studio with a tuned Gemini model endpoint
AnswerB

Gemini for Google Workspace embeds Gemini directly in Gmail, Docs, Drive, and related apps, letting non-technical staff summarize PDFs, draft content, and ground responses in files they can already access. It requires no infrastructure work, which matches the editorial team's need for a no-code workspace experience with shared-drive content available in context.

Why this answer

Gemini for Google Workspace is the offering designed to bring Gemini into the productivity apps employees already use, with enterprise-grade privacy protections and grounding in content the user can access. Because the scenario calls for drafting and summarizing inside a workspace with no coding or infrastructure, the embedded Workspace experience is the natural fit rather than developer consoles or agent platforms.

Exam trap

The trap here is assuming any Gemini-branded managed service delivers in-app Workspace drafting, when most Gemini offerings are developer or standalone-assistant surfaces rather than features embedded in Docs and Drive.

44
MCQeasy

A developer wants to integrate Gemini multimodal capabilities (text + image) into a mobile app using Python. Which Google Cloud client library should they use?

A.Dialogflow CX
B.Vertex AI client library (google-cloud-aiplatform)
C.Cloud Vision API
D.Natural Language API
AnswerB

The google-cloud-aiplatform library exposes Vertex AI's Gemini multimodal endpoints, handling text and image inputs natively from Python. It satisfies the stem's requirement to integrate Gemini text-plus-image capabilities into a mobile app backend, unlike single-modality or non-Vertex client libraries.

Why this answer

The Vertex AI client library (google-cloud-aiplatform) provides the Generative AI SDK that supports multimodal capabilities, including the ability to send both text and image inputs to Gemini models. This library directly exposes the `GenerativeModel` class with methods like `generate_content()` that accept `Part` objects containing image data (e.g., `Part.from_image()` or `Part.from_uri()`), making it the correct choice for integrating Gemini multimodal features into a Python mobile app backend.

Exam trap

The trap here is that candidates confuse specialized single-modality APIs (Vision, Natural Language) with the unified multimodal API provided by Vertex AI, assuming that combining separate services is equivalent to Gemini's native multimodal reasoning.

How to eliminate wrong answers

Option A is wrong because Dialogflow CX is a conversational AI platform for building chatbots and virtual agents, not a library for directly accessing Gemini multimodal models; it lacks the low-level API to construct multimodal requests with image parts. Option C is wrong because Cloud Vision API is a specialized service for image analysis (e.g., object detection, OCR) and does not provide access to Gemini's generative multimodal capabilities or its text+image reasoning. Option D is wrong because Natural Language API is designed for text-only analysis (e.g., sentiment, entity extraction) and cannot process image inputs or generate multimodal responses.

45
Multi-Selectmedium

A company is building a generative AI application that must adhere to strict data residency regulations. Which TWO Google Cloud features can help ensure that data does not leave a specific geographic region?

Select 2 answers
A.Using the regional endpoint for Vertex AI
B.Vertex AI Model Caching
C.Global load balancer with Cloud Armor
D.Cloud CDN for content delivery
E.Deploying models on dedicated VMs in a specific region
AnswersA, E

Regional endpoints ensure API calls stay within the region.

Why this answer

To ensure data residency, you can use the regional endpoint for Vertex AI (A) which ensures that API calls and data processing stay within the specified region. Additionally, deploying models on dedicated VMs in a specific region (E) ensures that compute resources and data do not leave that region. Option B (Vertex AI Model Caching) does not enforce residency as caching may use regional resources but does not guarantee data stays within a region.

Option C (Global load balancer with Cloud Armor) is a network security and load balancing service that operates globally. Option D (Cloud CDN) caches content at edge locations globally, which would move data outside the region.

46
MCQeasy

A startup wants to generate images from text descriptions for their marketing materials. They prefer a managed service that requires minimal coding. Which Google Cloud generative AI offering should they use?

A.Vertex AI Imagen
B.Natural Language API
C.Document AI
D.Cloud Speech-to-Text
AnswerA

Vertex AI Imagen provides text-to-image generation through a managed Google Cloud service, requiring minimal coding via API calls or the console. It directly satisfies the startup's need to create marketing visuals from text descriptions without building or hosting models themselves, unlike custom training approaches.

Why this answer

Vertex AI Imagen is Google Cloud's managed generative AI service specifically designed for text-to-image generation. It requires minimal coding, as users can interact with it via the Cloud Console, API calls with simple prompts, or through Vertex AI's built-in tools, making it ideal for a startup needing to generate marketing images from text descriptions without extensive development effort.

Exam trap

The trap here is that candidates may confuse general-purpose AI services (like Natural Language API or Document AI) with generative AI offerings, overlooking that only Vertex AI Imagen is purpose-built for text-to-image generation as a managed service with minimal coding.

How to eliminate wrong answers

Option B (Natural Language API) is wrong because it is designed for analyzing and extracting insights from text (e.g., sentiment, entity recognition), not for generating images from text. Option C (Document AI) is wrong because it focuses on processing and extracting data from documents (e.g., OCR, form parsing), not on generative image creation. Option D (Cloud Speech-to-Text) is wrong because it converts audio speech into text, which is the opposite direction of generating images from text descriptions.

47
MCQmedium

A data scientist runs the above command to upload a model to Vertex AI Model Registry. The model is a TensorFlow 2.6 model trained on tabular data. After deployment to an endpoint, the prediction latency is higher than expected. What is the most likely cause?

A.The artifact URI points to a single file instead of a directory
B.The model should be uploaded with a different display name
C.The container image used is CPU-only, but a GPU-accelerated image would improve latency
D.The model is uploaded to the wrong region
AnswerA

The URI likely points to a directory with SavedModel, which is correct.

Why this answer

When uploading a model to Vertex AI Model Registry, the artifact URI must point to a directory containing the SavedModel or model artifacts, not a single file. If the URI points to a single file, the model may not load properly or may serve with higher latency. For a TensorFlow 2.6 tabular model, GPU acceleration is generally not the primary latency bottleneck; tabular models are typically small and run efficiently on CPU.

The container image choice is less likely to be the root cause than an incorrect artifact URI.

Exam trap

Candidates often assume GPU acceleration is always the fix for latency issues, but for tabular models, correct model packaging and artifact URI structure are more common pitfalls. Always verify that the artifact URI points to a directory, not a single file.

How to eliminate wrong answers

Option A is wrong because the artifact URI pointing to a single file (e.g., a SavedModel.pb) is the standard and correct way to upload a TensorFlow model; Vertex AI expects a directory containing the saved_model.pb and variables folder, but a single file URI would cause an upload error, not higher latency. Option B is wrong because the display name is a human-readable label for the model in the registry and has no effect on prediction latency; it only affects model organization and retrieval. Option D is wrong because uploading to the wrong region would result in deployment failures or cross-region network latency, but the question states the model is deployed and latency is higher than expected, not that deployment failed or that there is network latency; the most direct cause is the container image type.

48
MCQmedium

A developer is using Vertex AI's Generative AI Studio to prototype a text summarization model. The initial results are too verbose. What is the most efficient way to adjust the output length without retraining?

A.Switch to a smaller base model like BERT
B.Use a separate classifier to filter long responses
C.Fine-tune the model with a dataset of concise summaries
D.Modify the prompt with specific length instructions and adjust model parameters
AnswerD

Modifying the prompt with explicit length instructions directly constrains generation at inference time, while adjusting parameters such as max output tokens and temperature tightens verbosity further. Both act without retraining, satisfying the stem's efficiency constraint, since no fine-tuning or dataset preparation is required to shorten summaries in Generative AI Studio.

Why this answer

The most efficient way to adjust output length without retraining is to modify the prompt with specific length instructions (e.g., 'Summarize in one sentence') and adjust model parameters like max output tokens, temperature, and top_p in Vertex AI's Generative AI Studio. Option A is wrong because switching to a smaller base model like BERT would require retraining or different model architecture, and BERT is not primarily a text generation model for summarization. Option B is wrong because adding a separate classifier increases complexity and does not directly control the model's output length.

Option C is wrong because fine-tuning requires additional training data and computational resources, which is not the most efficient approach compared to prompt and parameter adjustments.

49
Multi-Selecthard

A bank is building a Gemini-powered assistant on Vertex AI that must answer questions about internal policy documents and must not fabricate policy details. The architects want to reduce hallucination and provide auditable sourcing. Which two Google Cloud capabilities should they combine? (Choose two.)

Select 2 answers
A.Enabling citations that map generated statements to retrieved source chunks
B.Fine-tuning the model on the bank's public marketing brochures
C.Disabling safety filters to allow unrestricted policy answers
D.Setting the model temperature to its maximum value
E.Grounding with Vertex AI Search over an indexed policy data store
AnswersA, E

Citation support ties each generated claim back to the specific retrieved chunk it came from, giving reviewers a verifiable trail. Combined with grounding, this satisfies the bank's need for auditable sourcing and lets compliance staff confirm that a stated policy actually exists in the corpus, which is central to the stated requirement.

Why this answer

Grounding retrieves authoritative passages so responses are conditioned on real policy text, and citations expose exactly which passages supported each statement. Together they reduce fabrication and create an auditable chain from answer to source. Generation settings and safety controls do not supply factual anchoring, and fine-tuning on unrelated marketing material cannot substitute for retrieval over the actual policy corpus.

Exam trap

The trap here is treating fine-tuning as a substitute for grounding, when fine-tuning shapes style and behavior but does not retrieve or cite the current authoritative documents.

50
MCQmedium

A media company wants its internal knowledge assistant to answer employee questions using the company's own policy documents and past project reports, while keeping the Gemini model's general reasoning ability intact. The team has a large corpus stored in Cloud Storage and does not want to retrain or fine-tune the model. Which Google Cloud approach should they use?

A.Use Vertex AI Search with a data store pointing to the Cloud Storage corpus, and ground Gemini responses on the retrieved documents.
B.Deploy Gemini on a larger machine type with more GPU memory so the model can ingest the full document corpus at runtime.
C.Increase the Gemini model's temperature setting so it can explore the policy documents more broadly at inference time.
D.Fine-tune Gemini on the policy documents and past project reports using Vertex AI supervised tuning.
AnswerA

Vertex AI Search ingests and indexes the Cloud Storage corpus into a data store, then retrieves relevant passages at query time to ground Gemini. This supplies proprietary policy and project content without altering model weights, preserving the model's general reasoning. It is the intended Google Cloud pattern for enterprise retrieval-augmented generation over existing documents.

Why this answer

Grounding Gemini with Vertex AI Search connects the model to the company's own documents at query time, so responses reflect current policy and project material while the base model's general capabilities remain untouched. Because retrieval is separate from model weights, updating the corpus requires no retraining, which matches the team's stated constraint.

Exam trap

The trap here is assuming that feeding proprietary documents to Gemini requires fine-tuning, when retrieval-based grounding is the correct pattern for document question answering.

51
MCQeasy

A startup wants to generate product descriptions from a few keywords using a large language model. They have no prior ML experience and need the fastest time-to-market. Which Google Cloud service should they use?

A.Vertex AI Studio
B.Vertex AI Workbench with custom training
C.Vertex AI Agent Builder
D.Vertex AI Model Garden
AnswerA

Vertex AI Studio provides a no-code console for prompting and testing Gemini models directly, letting non-specialists generate product descriptions from keywords without building pipelines. This satisfies the startup's lack of ML experience and fastest time-to-market constraint, since no model training or infrastructure setup is required.

Why this answer

Vertex AI Studio provides a no-code/low-code environment with pre-trained foundation models and prompt templates, enabling rapid generation of product descriptions from keywords without any ML expertise. It offers the fastest time-to-market because it eliminates the need for custom model training, infrastructure setup, or coding, directly leveraging Google's generative AI capabilities through a simple interface.

Exam trap

The trap here is that candidates might confuse Vertex AI Studio with Vertex AI Model Garden, thinking Model Garden offers a faster path because it lists models, but Model Garden still requires deployment and configuration steps, whereas Studio provides immediate generation capabilities.

How to eliminate wrong answers

Option B is wrong because Vertex AI Workbench with custom training requires writing code, selecting models, and managing training jobs, which demands ML experience and significantly increases time-to-market compared to using a pre-built solution. Option C is wrong because Vertex AI Agent Builder is designed for creating conversational agents and chatbots, not for generating product descriptions from keywords; it adds unnecessary complexity and overhead for this simple text generation task. Option D is wrong because Vertex AI Model Garden is a repository of pre-trained models that still requires users to select, deploy, and potentially fine-tune models, which involves ML knowledge and setup time, not offering the fastest path for a non-ML team.

52
MCQmedium

A marketing team wants to use Vertex AI to generate ad copy. They need the model to follow a specific tone and style. What is the best approach?

A.Use Vertex AI Grounding to retrieve style guides
B.Provide few-shot examples in the prompt and adjust temperature
C.Fine-tune the model on a dataset of past ad copy
D.Enable safety filters to enforce brand guidelines
AnswerB

Few-shot examples embed the desired tone and style directly in the prompt, conditioning the model on concrete patterns rather than abstract instructions. Adjusting temperature controls output variability, keeping ad copy consistent with the brand voice. Together they satisfy the requirement to follow a specific tone and style.

Why this answer

Few-shot prompting with adjusted temperature is the best approach because it directly controls the model's output style and tone without requiring additional infrastructure or training. By providing a few examples of desired ad copy in the prompt, the model learns the specific tone and style through in-context learning, while temperature tuning (e.g., 0.2 for deterministic output) ensures consistency. This is the most efficient and cost-effective method for immediate, controllable generation.

Exam trap

The Google Gen AI Leader exam often tests the distinction between grounding (factual retrieval) and style control (prompt engineering), leading candidates to mistakenly choose grounding for stylistic tasks when it is only for factual accuracy.

How to eliminate wrong answers

Option A is wrong because Vertex AI Grounding is designed to ground model outputs in real-world data (e.g., from Google Search or enterprise datasets) to reduce hallucinations, not to enforce a specific tone or style from style guides—it retrieves factual context, not stylistic rules. Option C is wrong because fine-tuning requires a large, labeled dataset of past ad copy and significant compute resources, which is overkill for simply enforcing tone and style; it is better suited for domain-specific tasks like learning a new format or vocabulary. Option D is wrong because safety filters in Vertex AI block harmful content (e.g., hate speech, toxicity) but do not enforce brand-specific tone or style guidelines—they are a safety mechanism, not a stylistic control.

53
MCQhard

A software company wants to use a Google Cloud generative AI model to automatically generate code snippets from natural language descriptions. The developers need the model to produce syntactically correct code in multiple programming languages. Which Google Cloud offering is specifically designed for this use case?

A.Dialogflow CX
B.Vertex AI Matching Engine
C.Gemini Pro Vision
D.Codey APIs on Vertex AI
AnswerD

Codey APIs on Vertex AI are a family of foundation models specifically fine-tuned for code generation, code completion, and code chat. They support multiple programming languages and are designed to produce syntactically correct code from natural language. This directly matches the software company's need to generate code snippets from descriptions.

Why this answer

Codey APIs on Vertex AI are purpose-built for code generation tasks, including translating natural language to code in multiple languages. They are trained on a large corpus of code and are optimized for syntax and correctness. The other options are either general multimodal models, search services, or conversational agents, none of which are specialized for code generation.

Exam trap

The trap here is assuming that any large language model can generate code equally well, but specialized models like Codey are tuned for code accuracy and multi-language support.

54
MCQmedium

A healthcare company is building a chatbot to answer patient queries based on their medical documents stored in Cloud Storage. They want to minimize latency and ensure data residency in the EU. Which Vertex AI service should they use?

A.Vertex AI Model Garden with fine-tuning
B.Vertex AI Search with document grounding
C.Vertex AI Agent Builder with web search
D.Vertex AI Codey APIs
AnswerB

Vertex AI Search with document grounding indexes the Cloud Storage documents and serves grounded answers from an EU multi-region endpoint, keeping data resident in the EU while avoiding cross-region retrieval latency. This satisfies both the data residency and latency constraints.

Why this answer

Vertex AI Search with document grounding is correct because it allows the chatbot to ground responses in the customer's own medical documents stored in Cloud Storage, ensuring low latency through optimized indexing and retrieval, while supporting data residency controls to keep data within the EU. This service is specifically designed for enterprise search and Q&A over private document repositories, making it ideal for healthcare use cases requiring compliance and fast responses.

Exam trap

The trap here is that candidates may confuse Vertex AI Search (which grounds in private documents) with Vertex AI Agent Builder (which defaults to web search), or assume fine-tuning is necessary for domain-specific Q&A when retrieval-augmented generation (RAG) with document grounding is the correct approach for minimizing latency and ensuring data residency.

How to eliminate wrong answers

Option A is wrong because Vertex AI Model Garden with fine-tuning is intended for selecting and customizing foundation models, not for directly grounding answers in specific documents; it would require additional retrieval infrastructure and does not natively enforce data residency. Option C is wrong because Vertex AI Agent Builder with web search grounds responses in public web data, not private medical documents, and cannot guarantee data residency in the EU. Option D is wrong because Vertex AI Codey APIs are specialized for code generation and completion, not for answering queries based on document content.

55
MCQhard

A healthcare startup is using Vertex AI Imagen to generate synthetic medical images for training a diagnostic model. The images must comply with HIPAA regulations and cannot contain any real patient data. The team fine-tuned Imagen on a dataset of de-identified medical scans. However, during testing, they notice that some generated images closely resemble specific patients from the original dataset, even though the dataset was de-identified. They suspect that the model memorized some training examples. The team needs to address this issue without losing image quality. They have access to the original training data and Vertex AI tools. What action should they take?

A.Use a post-processing step to blur or distort generated images.
B.Re-tune the model using differential privacy (DP-SGD) to prevent memorization of individual examples.
C.Increase the size of the training dataset by adding more synthetic images.
D.Apply stricter output safety filters to block images that look like any known patient.
AnswerB

DP-SGD adds calibrated noise during fine-tuning, bounding any single training example's influence so the model cannot memorise identifiable scans. This removes the resemblance to specific patients while preserving overall image quality, meeting the HIPAA constraint without discarding the de-identified dataset.

Why this answer

Differential privacy (DP-SGD) during fine-tuning adds calibrated noise to the gradient updates, which mathematically bounds the model's ability to memorize any single training example. This directly addresses the memorization issue while preserving the utility of the generated images, as the noise is carefully controlled to maintain overall image quality. Vertex AI supports DP-SGD through its custom training infrastructure, making it a practical choice for HIPAA-compliant medical imaging.

Exam trap

Google often tests the misconception that post-processing or filtering can solve memorization, when in fact the root cause is in the training algorithm itself, requiring a privacy-preserving technique like differential privacy.

How to eliminate wrong answers

Option A is wrong because post-processing blur or distortion degrades image quality and does not prevent the underlying memorization; the model can still reproduce patient-identifying features before any filter is applied. Option C is wrong because simply adding more synthetic images does not stop the model from memorizing specific real examples; memorization is a function of the training algorithm, not dataset size alone. Option D is wrong because output safety filters are reactive and cannot block all variations of memorized patient features; they rely on predefined rules or classifiers, which are insufficient for the nuanced, high-dimensional nature of medical images.

56
Multi-Selectmedium

Which TWO features are available in Vertex AI Agent Builder to enhance the conversational abilities of an agent? (Choose TWO.)

Select 2 answers
A.Slot filling
B.Sentiment analysis
C.Code execution
D.Intent matching
E.Knowledge base integration
AnswersA, D

Slot filling collects required parameters from user input.

Why this answer

Slot filling is correct because it allows the agent to collect required parameters (slots) from the user during a conversation, enabling multi-turn interactions to fulfill complex requests. In Vertex AI Agent Builder, slot filling is a core feature for conversational agents, as it systematically prompts for missing information (e.g., date, location) until all necessary slots are filled, enhancing the agent's ability to handle dynamic user inputs.

Exam trap

The trap here is that candidates often confuse 'knowledge base integration' as a core conversational feature, but it is actually a retrieval-augmented generation (RAG) capability for grounding, not a direct mechanism for managing dialogue flow like slot filling or intent matching.

57
MCQmedium

A software company wants to embed an AI coding assistant into its internal developer portal. The assistant must complete code in the developer's current editor context and also answer natural-language questions about the repository. Which Google Cloud offering is designed for this use case?

A.Vertex AI Pipelines, which orchestrates machine learning workflow steps such as training and evaluation.
B.Gemini Code Assist, which provides code completion and conversational assistance integrated with developer tooling.
C.Cloud Data Loss Prevention, which discovers and de-identifies sensitive data across storage systems.
D.Gemini in BigQuery, which generates SQL from natural-language prompts against datasets in BigQuery.
AnswerB

Gemini Code Assist is purpose-built for developer workflows, offering inline code completion in supported IDEs and natural-language chat about code, including repository-aware assistance where configured. That matches both requirements in this scenario, unlike general-purpose model endpoints or data analytics services that would require the company to build editor integrations and completion logic itself.

Why this answer

The scenario needs two developer-centric capabilities: inline completion in the editor and conversational answers about the repository. Gemini Code Assist is the Google Cloud offering built for both, integrating with supported IDEs and developer tooling. The other services serve data analytics, ML workflow orchestration, or data security, and none provides editor-integrated code assistance for a developer portal.

Exam trap

The trap here is selecting a general Gemini or Vertex AI capability and assuming the team will build the editor integration themselves, when a purpose-built developer assistant already covers both required behaviors.

58
MCQmedium

A retail company is building a chatbot for customer service. They need the model to generate product descriptions based on a catalog but also answer questions about store policies. The team wants to minimize latency and cost while maintaining high accuracy. Which Google Cloud generative AI offering should they use?

A.Vertex AI Model Garden with PaLM 2
B.Vertex AI Imagen
C.Vertex AI Codey APIs
D.Vertex AI Gemini API
AnswerD

Vertex AI Gemini API provides a single multimodal endpoint handling both catalogue-grounded description generation and policy question answering, so no separate model hosting is needed. Consolidating on one managed API minimises latency and cost while retaining Gemini's accuracy.

Why this answer

The Vertex AI Gemini API provides a single, unified multimodal model capable of both generating product descriptions from a catalog and answering policy questions, while offering optimized latency and cost through features like context caching and streaming. Unlike specialized models, Gemini's architecture handles diverse natural language tasks without requiring multiple endpoints, reducing operational overhead and inference time.

Exam trap

The trap here is that candidates may assume PaLM 2 (Option A) is the best general-purpose text model, but Google has since replaced PaLM 2 with Gemini as the newer, more cost-efficient, and recommended model in Vertex AI for multimodal and text generation tasks.

How to eliminate wrong answers

Option A is wrong because Vertex AI Model Garden with PaLM 2, while capable of text generation, is a separate model that does not offer the same multimodal efficiency or cost-optimized infrastructure as Gemini, and PaLM 2 is being phased out in favor of Gemini for general-purpose tasks. Option B is wrong because Vertex AI Imagen is specifically designed for image generation and editing, not for generating text-based product descriptions or answering policy questions, making it unsuitable for this use case. Option C is wrong because Vertex AI Codey APIs are specialized for code generation and code-related tasks, such as writing functions or debugging, and are not designed for general customer service text generation or policy Q&A.

59
MCQhard

A healthcare organization wants to build a generative AI application that can answer patient questions based on their own medical knowledge base. They need the model to cite sources and avoid generating unsupported information. Which Google Cloud feature should they use to ground the model's responses in their proprietary data?

A.Vertex AI Pipelines
B.Vertex AI Search and Conversation
C.Vertex AI Model Evaluation
D.Vertex AI Feature Store
AnswerB

Vertex AI Search and Conversation allows grounding generative AI responses in enterprise data by integrating with a knowledge base. It provides citations and reduces hallucinations by retrieving relevant documents before generating answers, directly meeting the healthcare organization's need for source-backed responses from their medical knowledge base.

Why this answer

Vertex AI Search and Conversation enables grounding of generative AI outputs by retrieving relevant documents from an enterprise knowledge base and providing citations. This reduces hallucinations and ensures answers are based on proprietary data. The other options are for model evaluation, pipeline orchestration, or feature management, none of which provide grounding for generative AI responses.

Exam trap

The trap here is assuming that any Vertex AI service can ground generative AI, when only specific services like Vertex AI Search and Conversation are designed for retrieval-augmented generation.

60
MCQeasy

A small startup wants to add image generation to its design tool. The developers want to call a Google Cloud generative AI model through a simple API without provisioning infrastructure, managing model servers, or handling GPU capacity planning. Which Google Cloud offering best fits this requirement?

A.Cloud Run functions that call a locally hosted diffusion model packaged in the function's container image.
B.A custom container running Stable Diffusion on Google Kubernetes Engine with node auto-provisioning enabled.
C.The Imagen model on Vertex AI accessed through the Vertex AI API, using the managed model without deploying a custom endpoint.
D.Vertex AI Model Garden, where they deploy Imagen to a dedicated endpoint with autoscaling replicas.
AnswerC

Imagen on Vertex AI is available as a fully managed model that can be invoked through the Vertex AI API without deploying or scaling any infrastructure. This matches the startup's need for simple API access, no GPU capacity planning, and no model server management, while still keeping the workload inside Google Cloud's governed platform.

Why this answer

Imagen on Vertex AI provides managed, API-based image generation where Google handles the serving infrastructure and scaling. The startup only needs to authenticate and send prompts, which satisfies the no-infrastructure and no-capacity-planning constraints. Deploying models to endpoints or self-hosting on Kubernetes or functions adds operational burden the team explicitly wants to avoid.

Exam trap

The trap here is conflating Model Garden deployment with model consumption, when managed models can be called directly through the API without any endpoint deployment.

61
MCQhard

An organization is deploying a summarization model on Vertex AI and needs to ensure that the model's responses are consistent and avoid hallucinations. They have a labeled dataset of source documents and human-written summaries. Which approach would best align the model with their quality requirements?

A.Deploy the model with a larger max_output_tokens
B.Use prompt engineering with few-shot examples
C.Increase the temperature to 0.9
D.Perform supervised fine-tuning using their labeled dataset
AnswerD

Supervised fine-tuning adjusts the model's weights using paired source documents and human-written summaries, directly teaching the desired output distribution. This satisfies the consistency and hallucination-reduction constraint, unlike prompt engineering or retrieval, which cannot reliably shape response style.

Why this answer

Supervised fine-tuning (D) directly optimizes the model's weights using the labeled dataset of source documents and human-written summaries, which teaches the model to produce consistent, factual outputs and reduces hallucinations. This approach aligns the model's behavior with the specific quality requirements by learning from ground-truth examples, unlike prompt engineering or parameter adjustments that do not modify the underlying model.

Exam trap

The trap here is that candidates often overestimate the power of prompt engineering or parameter tweaks, mistakenly believing they can achieve the same level of reliability as fine-tuning, when in fact only supervised fine-tuning directly modifies the model to internalize the labeled data's quality standards.

How to eliminate wrong answers

Option A is wrong because increasing max_output_tokens only extends the length limit of the response, which does not improve consistency or reduce hallucinations—it may actually allow the model to generate more unverified content. Option B is wrong because prompt engineering with few-shot examples provides in-context guidance but does not permanently adjust the model's weights; it is less reliable for ensuring consistent, hallucination-free outputs across diverse inputs compared to fine-tuning. Option C is wrong because increasing temperature to 0.9 increases randomness in token selection, which amplifies variability and the risk of hallucinations, directly opposing the goal of consistency.

62
MCQmedium

A logistics company wants an AI agent that answers driver questions about delivery procedures, can look up a shipment's live status through an internal REST API, and can escalate to a dispatcher when it cannot resolve an issue. Which Google Cloud offering is purpose-built for assembling this kind of agent?

A.Cloud Dataflow with a streaming pipeline for shipment events
B.Vertex AI Feature Store for serving driver profile attributes
C.Vertex AI Agent Builder with tools and a connected data store
D.Cloud Scheduler jobs that poll the shipment API and email updates
AnswerC

Vertex AI Agent Builder is designed to compose conversational agents that combine grounded knowledge from data stores with tools that call external systems such as the internal REST API, and it supports escalation flows. It directly matches the requirement for procedure answers, live shipment lookups and handoff to a dispatcher within one managed agent framework.

Why this answer

An agent framework is needed when a solution must converse, ground answers in internal knowledge and invoke external systems through tools. Vertex AI Agent Builder provides those capabilities natively, including data store grounding and tool calls, plus escalation handling. Data pipelines, feature serving and scheduled jobs each solve adjacent problems but cannot assemble an interactive, tool-using agent on their own.

Exam trap

The trap here is confusing a data or feature platform that supplies information to an agent with the agent framework itself that manages dialogue, grounding and tool calls.

63
MCQmedium

A company needs to fine-tune a foundation model on Vertex AI for a custom text classification task with only 500 labeled examples. They want to minimize cost while achieving high accuracy. What is the MOST cost-effective approach?

A.Fine-tune the foundation model using full fine-tuning on the entire dataset.
B.Use model distillation to train a smaller student model.
C.Use Vertex AI LLM-based evaluation to compare multiple large models and select the best one.
D.Design prompts with few-shot examples and test it with the available data.
AnswerD

Few-shot prompting exploits the foundation model's in-context learning, requiring no weight updates or training compute. With only 500 labelled examples, fine-tuning risks overfitting and incurs Vertex AI training costs, so prompt design satisfies the cost-minimisation constraint while testing accuracy against held-out data.

Why this answer

The most cost-effective because it leverages prompt engineering with few-shot examples, which requires no training or infrastructure costs. With only 500 labeled examples, a well-designed prompt can often achieve high accuracy for custom text classification without the expense of fine-tuning or model distillation on Vertex AI.

Exam trap

The trap here is that candidates often assume fine-tuning or distillation is always necessary for custom tasks, overlooking that prompt engineering with few-shot examples can be highly effective and cost-efficient for small datasets.

How to eliminate wrong answers

Option A is wrong because full fine-tuning on only 500 examples is computationally expensive and may lead to overfitting, especially for a large foundation model, making it cost-inefficient. Option B is wrong because model distillation requires training a smaller student model using a larger teacher model, which involves significant compute and data costs, and is not justified for a small dataset. Option C is wrong because using Vertex AI LLM-based evaluation to compare multiple large models incurs high inference costs and does not directly solve the classification task; it is an evaluation step, not a deployment approach.

64
MCQmedium

A media company needs to generate short video clips from text prompts for a marketing campaign. They want to use a Google Cloud service that can create videos up to 8 seconds long, 720p resolution, and can be accessed via Vertex AI. Which Google Cloud generative AI offering should they use?

A.Imagen on Vertex AI
B.Gemini on Vertex AI
C.Chirp on Vertex AI
D.Veo on Vertex AI
AnswerD

Veo on Vertex AI is Google Cloud's generative AI model for video generation from text prompts. It supports creating videos up to 8 seconds at 720p resolution, matching the scenario's requirements. It is accessible via Vertex AI, making it the correct choice for generating short video clips for marketing.

Why this answer

Veo on Vertex AI is specifically built for video generation from text, supporting up to 8-second 720p clips, which aligns with the media company's need for short marketing videos. Other models like Imagen, Gemini, or Chirp serve different modalities and cannot produce video outputs.

Exam trap

The trap here is assuming that Gemini, as a multimodal model, can generate video, but it only processes video inputs and does not output video.

65
MCQeasy

An engineer is testing a generative AI application using the Gemini API. They receive a 400 error with message 'INVALID_ARGUMENT: text has been blocked.' What is the most likely cause?

A.The specified Gemini model version does not exist
B.The input text was flagged by a safety filter
C.The API key is invalid
D.The API quota has been exceeded
AnswerB

A 400 INVALID_ARGUMENT with 'text has been blocked' indicates the request content tripped Gemini's safety filters, which screen prompts and responses for harmful categories. The input was rejected before generation, so no model output was produced.

Why this answer

The 400 error with 'INVALID_ARGUMENT: text has been blocked' indicates that the input text was rejected by Gemini's safety filters before any model processing occurred. These filters evaluate content against categories like hate speech, harassment, or sexually explicit material, and when triggered, they block the request at the API gateway level, returning this specific error message.

Exam trap

In Google Cloud's Gemini API, a 400 'INVALID_ARGUMENT: text has been blocked' error indicates that the input was rejected by safety filters before processing. Candidates often confuse this with authentication (invalid API key) or quota errors, which return different error codes (e.g., 403 or 429). Understanding the specific error message and its association with Gemini's safety settings is key.

How to eliminate wrong answers

Option A is wrong because a non-existent model version would return a 404 'NOT_FOUND' error, not a 400 'INVALID_ARGUMENT' error. Option C is wrong because an invalid API key would result in a 401 'UNAUTHENTICATED' or 403 'PERMISSION_DENIED' error, not a 400 error. Option D is wrong because exceeding API quota would return a 429 'RESOURCE_EXHAUSTED' error, not a 400 'INVALID_ARGUMENT' error.

66
MCQmedium

A logistics company stores delivery exception reports as PDFs in a Cloud Storage bucket. An analyst needs to ask natural-language questions across all the reports and receive answers with citations to the source pages, with minimal development work. Which Google Cloud capability should the analyst use?

A.BigQuery ML with a remote Gemini model
B.Vertex AI Pipelines orchestrating a custom document parser
C.Cloud Vision API batch annotation of the PDFs
D.Vertex AI Search with a data store grounded on the Cloud Storage documents
AnswerD

Vertex AI Search can ingest documents from Cloud Storage into a data store, index them, and answer natural-language queries with grounding and citations back to the source content. It is a managed retrieval-augmented generation capability, so the analyst gets cited answers across many PDFs with minimal development work, exactly matching the requirement.

Why this answer

Vertex AI Search ingests documents directly from Cloud Storage, builds a managed index, and returns natural-language answers grounded in those documents with citations to source pages. The alternatives all require substantial custom development: Pipelines orchestrates custom code, BigQuery ML needs the data loaded and parsed, and Cloud Vision only extracts raw text and features without answering questions.

Exam trap

The trap here is equating document text extraction with question answering; OCR or annotation services produce raw content, but only a grounded search capability delivers cited answers over the corpus.

67
Multi-Selecthard

A software company is using Vertex AI to build a generative AI application that creates code snippets from natural language descriptions. They want to improve the model's performance on their specific coding style and libraries. Which two techniques should they use? (Choose two.)

Select 2 answers
A.Deploy the model on a larger GPU to increase generation speed.
B.Use Vertex AI Model Evaluation to compare different models and select the best one.
C.Increase the model's temperature setting to encourage more creative code.
D.Use prompt engineering with few-shot examples of the desired code style.
E.Fine-tune the Gemini model on a dataset of code snippets and descriptions.
AnswersD, E

Prompt engineering with few-shot examples provides the model with concrete instances of the desired output format and style. This guides the model to generate code that matches the company's conventions without retraining, and it is a correct, efficient technique for customization.

Why this answer

Fine-tuning on a custom dataset and using few-shot prompt engineering are both effective ways to adapt a generative model to a specific coding style. Fine-tuning provides deep customization, while few-shot prompting offers a lightweight alternative. Together, they can significantly improve the relevance and accuracy of generated code.

Exam trap

The trap here is assuming that infrastructure changes like larger GPUs or evaluation tools can customize model behavior, when in fact they do not affect the model's learned patterns.

68
MCQeasy

A small marketing agency wants to quickly experiment with prompts for Gemini to draft social media posts, without writing code or provisioning infrastructure. They need a Google Cloud console experience for iterating on prompts and comparing model outputs. Which offering should they use?

A.BigQuery ML
B.Gemini for Google Workspace
C.Vertex AI Pipelines
D.Vertex AI Studio
AnswerD

Vertex AI Studio provides a console-based environment for designing, testing, and comparing prompts across Gemini models without writing code. It lets the agency iterate quickly on social post drafts and view outputs side by side, which matches the need for rapid experimentation with no infrastructure provisioning. This is the intended low-friction entry point for prompt work in Google Cloud.

Why this answer

Vertex AI Studio is the console-based workspace for prompt design, model comparison, and rapid prototyping with Gemini. It requires no coding or infrastructure setup, making it ideal for a marketing team that wants to experiment with social post drafts. The other choices are pipeline orchestration, SQL-based modeling, or office productivity tools that do not offer interactive prompt experimentation.

Exam trap

The trap here is conflating any Gemini access with prompt engineering capability, when only Vertex AI Studio provides the console playground for iterating and comparing prompts.

69
Multi-Selectmedium

A company is evaluating Google Cloud's generative AI offerings for enterprise use. Which TWO considerations are most important when selecting the right model deployment option?

Select 2 answers
A.Data residency
B.Developer preference
C.Model size
D.Latency requirements
E.Training time
AnswersA, D

Data residency determines where inference and stored prompts physically reside, satisfying legal and regulatory constraints on cross-border processing. For enterprises in regulated sectors, this governs whether a deployment option is permissible at all, independent of model quality or cost.

Why this answer

Data residency (A) is a critical consideration because enterprise deployments must comply with regional and regulatory requirements, and Google Cloud's Vertex AI lets you pin model endpoints and data processing to specific locations such as us-central1 or europe-west4 to satisfy sovereignty and compliance mandates. Latency requirements (D) are equally important because the deployment option determines network proximity and serving characteristics — for example, a regional endpoint close to users or a provisioned throughput configuration reduces round-trip time, which directly affects real-time generative AI application performance. Together, residency and latency drive the choice between options like global vs. regional endpoints, on-demand vs. provisioned throughput, and Vertex AI vs.

Gemini API. Developer preference (B) is subjective and does not determine the correct deployment architecture for enterprise compliance or performance needs. Model size (C) is a model characteristic rather than a deployment-selection criterion, since Vertex AI abstracts the underlying infrastructure.

Training time (E) is irrelevant to deployment selection because it concerns the model-building phase, not how the model is served in production.

Exam trap

A common misconception is that model size or training time are deployment considerations, when in fact they are model development concerns, not factors for selecting a deployment option.

70
MCQmedium

A healthcare company is building a chatbot to answer patient queries using Vertex AI Agent Builder. They want to ensure the chatbot only uses approved medical references and does not generate unverified advice. How should they configure the agent?

A.Set up grounding with a private data store containing verified medical documents
B.Enable strict safety filters to block any medical advice
C.Increase the temperature parameter to get more diverse responses
D.Use Vertex AI Model Monitoring to track answer accuracy
AnswerA

Grounding the agent with a private data store restricts responses to verified medical documents, preventing unverified advice. This satisfies the requirement that the chatbot only use approved references, since grounding anchors generation to retrieved source content rather than the model's parametric knowledge.

Why this answer

Vertex AI Agent Builder supports grounding with private data stores, which allows the agent to restrict its responses to only the information contained in the verified medical documents. This ensures the chatbot does not generate unverified advice by grounding its outputs in a trusted, curated dataset rather than relying on the model's general training data.

Exam trap

Google often tests the misconception that safety filters or monitoring tools can replace the need for explicit grounding in a private data store, when in fact only grounding ensures the model's output is strictly limited to approved content.

How to eliminate wrong answers

Option B is wrong because strict safety filters block all medical advice, which would prevent the chatbot from answering any queries at all, rather than controlling the source of the advice. Option C is wrong because increasing the temperature parameter makes responses more random and diverse, which is the opposite of what is needed for a controlled, fact-based medical chatbot. Option D is wrong because Vertex AI Model Monitoring tracks answer accuracy over time but does not prevent the model from generating unverified advice at inference time; it is a monitoring tool, not a grounding mechanism.

71
MCQeasy

A developer is configuring a Vertex AI Agent Builder agent to use grounding. They receive the following error when calling the API: `404 Not Found: Data store resource not found.` What is the most likely cause?

A.The data store was just created and is not yet propagated
B.The agent is not authenticated to access the data store
C.The data store has not been created in the specified project
D.The grounding configuration is missing required permissions
AnswerC

Grounding requires an existing Vertex AI Search data store to retrieve from. If the API call references a data store ID that does not exist in the specified project, the request fails immediately, since there is no index to query for grounding content.

Why this answer

The error message indicates that the data store resource cannot be found in the specified Google Cloud project. Vertex AI Agent Builder requires the data store to exist and be fully provisioned in the same project as the agent before grounding can be configured. If the data store ID or project ID is incorrect, or the data store has not been created, the API will return a 'not found' error.

Exam trap

This exam often tests the distinction between resource existence errors and permission errors; the trap here is that candidates confuse a 'not found' error (404) with an authentication or permission error (403/401), leading them to select options B or D instead of recognizing the missing resource.

How to eliminate wrong answers

Option A is wrong because a newly created data store propagates within seconds to minutes, and the API would return a 'resource not ready' or 'propagation in progress' error, not a 'not found' error. Option B is wrong because authentication errors (e.g., missing IAM permissions) produce a 403 Forbidden or 401 Unauthorized response, not a 'not found' error. Option D is wrong because missing permissions for grounding configuration would result in a permissions error (403), not a resource-not-found error.

72
MCQhard

A bank is piloting a Gemini-powered assistant that summarizes internal audit reports. Compliance requires that prompts and responses never leave the company's chosen Google Cloud region, that customer-managed encryption keys protect data at rest, and that no data is used to improve Google's models. Which combination of Vertex AI capabilities should the team configure to meet these requirements?

A.Vertex AI with CMEK enabled and a VPC Service Controls perimeter, relying on the perimeter to guarantee regional data processing.
B.Vertex AI with a global endpoint, default Google-managed encryption keys, and a prompt instruction telling the model not to retain any input.
C.Gemini accessed through AI Studio with billing enabled, plus a Cloud Armor policy restricting access to the bank's IP ranges.
D.Vertex AI with data residency controls, CMEK through Cloud KMS, and the documented data governance guarantee that customer data is not used to train Google's foundation models.
AnswerD

Vertex AI supports selecting a region for processing and storage, integrates with Cloud KMS so customer-managed encryption keys protect data at rest, and its terms state that customer data is not used to improve Google's foundation models. Together these three controls directly satisfy the residency, encryption, and no-training requirements the bank's compliance team imposed.

Why this answer

Meeting all three compliance constraints requires regional processing for residency, Cloud KMS customer-managed keys for encryption at rest, and reliance on Vertex AI's data governance commitments that customer data is not used to train Google's foundation models. Network or prompt-level controls cannot substitute for these platform-level guarantees.

Exam trap

The trap here is treating VPC Service Controls or prompt instructions as substitutes for data residency configuration and model training guarantees, which they are not.

73
MCQmedium

A media company wants to generate short video clips from text prompts for social media ads. The creative team needs a managed Google Cloud service that produces video from descriptive prompts and offers controls for aspect ratio and duration, without managing GPU infrastructure. Which offering should they use?

A.Vertex AI Chirp
B.Vertex AI Veo
C.Vertex AI Imagen
D.Cloud Vision API
AnswerB

Veo is Google's generative video model available through Vertex AI, capable of producing short video clips from text prompts with controls over characteristics such as aspect ratio and duration. Because it is offered as a managed model on Vertex AI, the creative team avoids provisioning or managing GPUs. This directly matches the requirement to generate video from descriptive prompts in a managed fashion.

Why this answer

The team needs prompt-driven video generation delivered as a managed service, which is what Veo on Vertex AI provides, including controls over clip attributes and no infrastructure management. Imagen produces still images, Chirp targets speech, and Cloud Vision API analyzes existing images, so none of the alternatives can generate video content for the social ad campaign.

Exam trap

The trap here is assuming any Google generative media model produces video, when Imagen is image-only and Veo is the model purpose-built for text-to-video generation.

74
MCQmedium

A security team wants to prevent prompt injection attacks on their generative AI application hosted on Vertex AI. Which best practice should they implement?

A.Use a custom model instead of a foundation model
B.Disable all logging
C.Use a private endpoint
D.Implement input validation and output filtering
AnswerD

Input validation strips or neutralises injected instructions before they reach the model, while output filtering catches unsafe or manipulated responses. Together they directly counter prompt injection, satisfying the security team's requirement to prevent attacks on the Vertex AI application.

Why this answer

Prompt injection attacks exploit the model's inability to distinguish between user instructions and untrusted input. Implementing input validation (e.g., sanitizing special characters or known injection patterns) and output filtering (e.g., using a classifier to detect and block malicious responses) directly mitigates this risk by controlling what the model processes and returns. On Vertex AI, this can be enforced via custom safety attributes or integration with services like Cloud DLP for data loss prevention.

Exam trap

The trap here is that candidates confuse network-level security controls (like private endpoints) with application-layer security controls, assuming that restricting network access alone can prevent content-based attacks like prompt injection.

How to eliminate wrong answers

Option A is wrong because using a custom model does not inherently prevent prompt injection; the vulnerability exists in any model that processes untrusted input, regardless of whether it is a foundation model or a custom model. Option B is wrong because disabling logging removes visibility into attack attempts and compliance auditing, but does not prevent the injection itself; logging is a detection mechanism, not a prevention control. Option C is wrong because a private endpoint (e.g., Private Service Connect) secures network traffic by keeping it within a VPC, but it does not inspect or sanitize the content of prompts or outputs, leaving the application vulnerable to injection attacks at the application layer.

75
MCQmedium

A logistics company wants to build a generative AI application that answers employee questions about internal HR policies. The company's policy documents are already stored in Google Drive and must stay synchronized automatically as they are edited. The team has limited ML engineering resources and prefers a managed, low-code path. Which Google Cloud approach should they choose?

A.Use the Gemini API with a long system prompt that pastes the full text of every HR policy document into each request.
B.Create a Vertex AI Search application with a data store connected to the Google Drive folder, then ground Gemini responses on that data store.
C.Deploy Gemini on a Compute Engine VM with an attached GPU and write a custom retrieval pipeline that polls the Drive API for changes.
D.Export the Drive documents to Cloud Storage, then train a custom Gemini model from scratch on those files using Vertex AI Training.
AnswerB

Vertex AI Search supports connecting a data store directly to Google Drive as a federated or indexed source, so document edits are picked up without manual re-uploading. Grounding Gemini on that data store gives the HR assistant accurate, cited answers with minimal ML engineering effort, matching both the synchronization and low-code requirements.

Why this answer

Grounding Gemini on a Vertex AI Search data store connected to Google Drive satisfies both constraints: it keeps answers tied to the latest policy documents automatically, and it delivers a managed, low-code experience. The other approaches either require custom ML work, break synchronization, or scale poorly as the document set grows.

Exam trap

The trap here is assuming that generative answers require fine-tuning or custom model training, when grounding on a managed search data store is the intended low-effort pattern.

Page 1 of 3 · 182 questions totalNext →

Ready to test yourself?

Try a timed practice session using only Google Cloud's Generative AI Offerings questions.