Courseiva

CCNA Google Cloud Gen Ai Offerings Questions

32 of 182 questions · Page 3/3 · Google Cloud Gen Ai Offerings topic · Answers revealed

151
MCQhard

A research team wants to fine-tune a Gemini model on Vertex AI using a dataset of proprietary scientific abstracts. They need to adjust the model's behavior with supervised fine-tuning while keeping the base model's general knowledge. Which Vertex AI capability should they use?

A.Vertex AI Pipelines
B.Vertex AI Vision
C.Supervised fine-tuning of Gemini on Vertex AI
D.Vertex AI Matching Engine
AnswerC

Supervised fine-tuning updates a Gemini model's weights using labeled input-output pairs, teaching it domain-specific patterns while retaining the base model's pretrained knowledge. This is exactly what the research team needs to adapt the model to scientific abstracts without losing general capabilities, and it is supported directly in Vertex AI.

Why this answer

Supervised fine-tuning on Vertex AI lets teams adapt Gemini models with their own labeled examples, adjusting behavior for specialized domains such as scientific literature. The tuned model retains the base model's broad knowledge while learning task-specific patterns. This is the managed capability designed for customizing Gemini, unlike vision, search, or orchestration services that serve different purposes.

Exam trap

The trap here is selecting an orchestration or retrieval service such as Vertex AI Pipelines or Matching Engine when the requirement is specifically to change model weights through supervised fine-tuning.

152
MCQmedium

A developer is using the Vertex AI PaLM API and receives a 429 Resource Exhausted error. What is the most likely cause?

A.The request payload is too large
B.The user has exceeded the allowed number of requests per minute
C.The model is not available in the current region
D.The API key is invalid
AnswerB

The 429 Resource Exhausted status signals quota exhaustion on the Vertex AI PaLM endpoint. Per-minute request quotas cap how many calls a project may issue; exceeding that rate triggers this error, so the cause is too many requests within the minute window.

Why this answer

HTTP 429 Resource Exhausted is the standard error returned when a client exceeds the rate limits or quota assigned to the API — most commonly the requests-per-minute (RPM) limit. In the Vertex AI PaLM API, each project has per-minute quotas for prediction requests, and exceeding them triggers this error. The correct remediation is to implement exponential backoff and/or request a quota increase.

Exam trap

Generative AI Leader often tests the mapping between HTTP status codes and their causes, and candidates frequently confuse 429 (rate/quota) with 400 (bad request) or 403 (permission), leading them to pick payload-size or API-key options.

How to eliminate wrong answers

Option A is wrong because an oversized payload typically returns a 400 Bad Request or 413 Payload Too Large error, not 429 Resource Exhausted — payload size is a request validity issue, not a rate/quota issue. Option C is wrong because a model unavailable in the current region returns a 404 Not Found or a region-specific error indicating the model is not supported, not a 429. Option D is wrong because an invalid API key returns 401 Unauthorized or 403 Forbidden, which are authentication/authorization errors, not resource exhaustion.

153
MCQmedium

A retailer is building a product recommendation chatbot using Vertex AI Agent Builder. They want the agent to answer questions about product availability, prices, and promotions, but also to escalate to a human agent when the query is complex. What should they configure in Agent Builder?

A.Create a playbook with a step that transfers to a human via a webhook
B.Define an agent with a 'handoff to human' intent and configure the corresponding flow
C.Integrate a tool that calls a human support API when confidence is low
D.Use Vertex AI Agent Builder's generative fallback to automatically escalate
AnswerB

Agent Builder supports handoff to a human agent through intent and flow configuration.

Why this answer

In Vertex AI Agent Builder, you define intents to handle specific user requests. A 'handoff to human' intent, when triggered, initiates a configured flow that transfers the conversation to a human agent. This is the standard method for escalation.

Option A is incorrect because playbooks define conversation steps but not directly the escalation trigger; they can be part of a flow but not the primary escalation mechanism. Option C is incorrect: while you could integrate a tool to call a human support API, Agent Builder provides a built-in handoff intent and flow, making custom tool integration unnecessary. Option D is incorrect: generative fallback only generates responses when the agent cannot match a query, it does not provide a structured escalation to a human.

154
MCQmedium

A company is building a customer service chatbot using Vertex AI Agent Builder. The chatbot needs to answer questions based on a large internal knowledge base stored in a Cloud Storage bucket. The team wants to ensure the model can reference the latest documents without fine-tuning. Which configuration should they use?

A.Fine-tune a model on the knowledge base documents
B.Use a pre-built model with no additional configuration
C.Store the documents in BigQuery and use a BigQuery connector
D.Ground the model with a Vertex AI Search data store connected to the Cloud Storage bucket
AnswerD

A Vertex AI Search data store indexes the Cloud Storage documents and grounds responses at query time, so the chatbot cites current content without retraining. This satisfies the no-fine-tuning and latest-documents constraints that embedding-only or static prompt approaches fail.

Why this answer

Grounding with a Vertex AI Search data store is the correct pattern for retrieval-augmented generation (RAG) in Vertex AI Agent Builder. The data store indexes the Cloud Storage documents and the agent retrieves relevant chunks at query time, so the model can cite the latest content without any fine-tuning. This keeps the knowledge base fresh — updating a document in the bucket automatically flows into subsequent answers.

Exam trap

Generative AI Leader often tests the distinction between fine-tuning (bakes knowledge into weights, static) and grounding/RAG (retrieves live external knowledge) — candidates pick fine-tuning when the scenario explicitly says 'without fine-tuning' or 'latest documents'.

How to eliminate wrong answers

Option A is wrong because fine-tuning bakes knowledge into model weights, requires retraining whenever documents change, and is expensive and slow — it does not satisfy the 'latest documents without fine-tuning' requirement. Option B is wrong because a pre-built model with no configuration has no access to the internal knowledge base and will hallucinate or refuse to answer domain-specific questions. Option C is wrong because while BigQuery can be a data source, the question specifies documents already stored in Cloud Storage; adding a BigQuery hop is unnecessary and does not by itself ground the model unless wired through a Vertex AI Search data store.

155
MCQeasy

A developer needs to generate embeddings for text data to be used in a semantic search application. Which Google Cloud service should they use?

A.Document AI
B.Cloud Translation API
C.Cloud Speech-to-Text
D.Vertex AI Embeddings API
AnswerD

Vertex AI Embeddings API generates dense vector representations from text, which semantic search requires for similarity comparison. It provides managed embedding generation without building custom infrastructure, matching the developer's need for text embeddings on Google Cloud.

Why this answer

Vertex AI Embeddings API is the correct choice because it provides a managed service to generate vector embeddings from text data, which are essential for semantic search applications that rely on understanding meaning rather than exact keyword matches. This API leverages large language models to convert text into high-dimensional vectors, enabling efficient similarity search using vector databases or nearest neighbor algorithms.

Exam trap

The trap here is that candidates may confuse Document AI's ability to extract text from documents with the need to generate embeddings from that text, overlooking that embedding generation is a separate, specialized step required for semantic search.

How to eliminate wrong answers

Option A is wrong because Document AI is designed for document processing tasks like OCR, parsing, and extraction of structured data from documents, not for generating text embeddings. Option B is wrong because Cloud Translation API is used for translating text between languages, not for creating vector representations of text for semantic search. Option C is wrong because Cloud Speech-to-Text converts audio to text, but does not generate embeddings or support semantic search directly.

156
MCQhard

A gaming company is using Vertex AI Imagen to create concept art. They have a stable pipeline that generates images based on text prompts. Recently, they introduced a new feature: using a reference image to guide the style (image-to-image generation). However, when using a reference image, the generated images often have unnatural color shifts and artifacts. The team suspects that the reference image is being resized to a resolution that the model wasn't trained on. They are using the default Imagen settings. What is the most likely cause and the best solution?

A.Increase the number of inference steps to improve detail.
B.The reference image is being resized to a non-standard aspect ratio; preprocess the image to the recommended resolution and aspect ratio.
C.Reduce the style weight in the image-to-image prompt.
D.Switch to a different image generation model.
AnswerB

Imagen's default preprocessing resizes reference images, and mismatched aspect ratios distort content, producing colour shifts and artifacts. Preprocessing to the recommended resolution and aspect ratio preserves the model's expected input geometry, eliminating the distortion at its source.

Why this answer

The default Imagen settings expect input images at specific resolutions (e.g., 256x256, 512x512, or 1024x1024) and a 1:1 aspect ratio. When a reference image is resized to a non-standard resolution or aspect ratio, the model's internal processing can introduce artifacts and unnatural color shifts due to misalignment with its training distribution. Preprocessing the image to the recommended resolution and aspect ratio ensures the model operates within its optimal input space, eliminating these issues.

Exam trap

The trap here is that candidates may confuse image quality issues with model hyperparameters (like inference steps or style weight) rather than recognizing that the fundamental input preprocessing—specifically resolution and aspect ratio—is the most common cause of artifacts in image-to-image generation with Imagen.

How to eliminate wrong answers

Option A is wrong because increasing inference steps primarily refines noise reduction and detail, but it does not address the root cause of resolution mismatch; it may even amplify artifacts from a poorly resized input. Option C is wrong because reducing style weight controls how strongly the reference image influences the output, but it does not fix the fundamental problem of the reference image being at a non-standard resolution or aspect ratio, which causes artifacts regardless of style influence. Option D is wrong because switching to a different model like Stable Diffusion would not resolve the resolution preprocessing issue; the team would still need to properly resize the reference image for that model, and the problem is with their pipeline, not the model itself.

157
MCQeasy

A small marketing agency wants to add an AI assistant that summarizes campaign briefs and drafts social posts. The developers have limited cloud experience and prefer a fully managed, serverless way to call Google's Gemini models without provisioning infrastructure. Which approach best meets this requirement?

A.Use the Gemini API through Google AI Studio for rapid prototyping and API-key access
B.Deploy an open model on a Compute Engine VM with an attached GPU
C.Create a Vertex AI custom training job for a bespoke summarization model
D.Build a Cloud Run service that hosts a self-managed transformer container
AnswerA

Google AI Studio and the Gemini API provide a fully managed, serverless way to call Gemini models using an API key, with no infrastructure to provision. This suits a small team with limited cloud experience that wants quick access for summarization and drafting. It is the lowest-friction entry point for experimenting with and shipping lightweight Gemini-powered features.

Why this answer

When the goal is simply to consume Gemini capabilities with minimal setup, the managed Gemini API accessed through Google AI Studio removes infrastructure concerns entirely. It supports summarization and drafting through straightforward API calls, letting a small team focus on prompts and application logic rather than serving stacks. Self-hosted or custom-trained alternatives introduce operational or data-science overhead the scenario does not require.

Exam trap

The trap here is assuming that any Google Cloud compute service, such as Cloud Run, is automatically the serverless answer for calling Gemini, when the managed Gemini API already removes the need to host anything.

158
MCQmedium

A financial services company wants to use generative AI to summarize large volumes of internal documents while ensuring that sensitive data never leaves their virtual private cloud (VPC). They need a solution that provides Gemini models with enterprise-grade security and data residency controls. Which Google Cloud service should they use?

A.Gemini API in Google AI Studio
B.Document AI
C.Vertex AI with Gemini models
D.Cloud Natural Language API
AnswerC

Vertex AI provides Gemini models with enterprise security features such as VPC Service Controls, Customer-Managed Encryption Keys (CMEK), and data residency. It allows the company to keep sensitive data within their VPC and meet compliance requirements, directly addressing the need for secure generative AI on internal documents.

Why this answer

Vertex AI offers Gemini models with enterprise-grade security, including VPC Service Controls and data residency, which are critical for handling sensitive financial data. It enables the company to use generative AI while maintaining strict control over data location and access. The other services either lack generative capabilities or do not provide the necessary security and compliance features for this scenario.

Exam trap

The trap here is assuming that the Gemini API in Google AI Studio offers the same enterprise security controls as Vertex AI, when it is primarily designed for developer ease of use.

159
MCQeasy

A media company wants to produce short video clips from text prompts for social media campaigns. They need a Google Cloud service that can generate video from text and edit existing video clips, with enterprise-grade controls. Which Google Cloud offering should they choose?

A.Vertex AI Gemini
B.Vertex AI Veo
C.Vertex AI Imagen
D.Vertex AI Chirp
AnswerB

Veo on Vertex AI is Google Cloud's generative video model that creates high-quality video from text or image prompts and supports video editing capabilities. It is integrated with enterprise controls such as IAM, VPC Service Controls, and data governance, making it the right fit for a media company needing controlled video generation and editing.

Why this answer

Veo on Vertex AI is purpose-built for generative video, producing clips from text or image prompts and supporting video editing. It is managed within Vertex AI, so organizations get enterprise security, access control, and compliance features. Other models in Vertex AI handle images, text, or speech, but not video generation and editing in one service.

Exam trap

The trap here is assuming any multimodal model such as Gemini can generate video, when video synthesis specifically requires a dedicated video generation model like Veo.

160
MCQmedium

A company is building a document summarization tool using Vertex AI Gemini API. They notice that the model sometimes returns incomplete summaries that miss key points. Which approach is most likely to improve summary quality without increasing token usage significantly?

A.Refine the system instruction to specify the desired summary format and key elements to include
B.Increase the context window to include more of the document
C.Switch to a larger Gemini model (e.g., from 1.0 Pro to 1.5 Pro)
D.Increase the max output token limit to allow longer summaries
AnswerA

Refining the system instruction specifies the required summary format and key elements, steering the model to cover them without enlarging the input. This improves completeness at negligible token cost, unlike approaches that add lengthy examples or extra context.

Why this answer

Refining the system instruction directly addresses the root cause of incomplete summaries by providing explicit guidance on the desired output format and key elements to include. This approach improves the model's adherence to the task without increasing the number of input or output tokens, as it only modifies the instruction text, not the document length or generation limits.

Exam trap

In Google Cloud exams, a common misconception is that increasing model size or token limits directly improves output quality, when in fact prompt engineering—specifically system instructions—is a more efficient and cost-effective lever for controlling model behavior.

How to eliminate wrong answers

Option B is wrong because increasing the context window adds more tokens from the document, which does not guarantee the model will focus on key points and can actually dilute attention, increasing token usage without improving summary completeness. Option C is wrong because switching to a larger Gemini model (e.g., from 1.0 Pro to 1.5 Pro) increases computational cost and token usage (due to larger model overhead) but does not inherently fix the instruction quality; the issue is prompt design, not model capacity. Option D is wrong because increasing the max output token limit allows longer summaries but does not ensure the model includes missing key points; it may simply produce more verbose text without addressing the root cause of omission.

161
MCQeasy

A small marketing analytics team wants to build a generative AI assistant that can answer questions about their proprietary campaign performance data. They have very limited machine learning engineering resources and want the fastest possible path to a working prototype on Google Cloud. Which approach should they take?

A.Deploy an open-source large language model on a self-managed Google Kubernetes Engine cluster with autoscaling GPU node pools.
B.Train a custom transformer model from scratch on their campaign data using Vertex AI Training with A3 machine types.
C.Create a BigQuery ML remote model that calls a third-party public API and join it to campaign tables in scheduled queries.
D.Use the Gemini models through Vertex AI with a managed RAG Engine backed by their campaign data stored in BigQuery.
AnswerD

The Gemini models on Vertex AI deliver strong general reasoning and generation out of the box, while the managed RAG Engine handles chunking, embedding, retrieval, and grounding against BigQuery data without the team operating any serving or indexing infrastructure. This directly matches the requirement for a fast prototype with minimal ML engineering effort, since the team only supplies data and prompts.

Why this answer

The key requirement is speed to a working prototype with minimal ML engineering, and managed Gemini endpoints plus a managed RAG pipeline satisfy that without infrastructure ownership. The team supplies campaign data and prompts while Google Cloud handles retrieval, grounding, and serving. Options that require training from scratch, self-managing GPU clusters, or calling external APIs all add engineering burden or governance risk that the scenario rules out.

Exam trap

The trap here is assuming that any generative AI prototype must involve training or hosting a model yourself, when managed Gemini endpoints with grounded retrieval remove that burden entirely.

162
MCQeasy

A startup wants to build a conversational AI assistant that can understand and generate text, images, and code. They need a single model that can handle multimodal inputs and outputs. Which Google Cloud generative AI model should they choose?

A.Gemini
B.Imagen
C.Chirp
D.Codey
AnswerA

Gemini is Google's multimodal generative AI model that natively understands and generates text, images, and code. It can process inputs across modalities and produce outputs in various formats. For a conversational assistant requiring multimodal capabilities, Gemini is the appropriate choice.

Why this answer

Gemini is the only Google Cloud generative AI model designed to be natively multimodal, capable of understanding and generating text, images, and code. This makes it ideal for a conversational AI assistant that needs to handle diverse content types within a single model.

Exam trap

The trap here is thinking that specialized models like Imagen or Codey can be combined to achieve multimodality, but the scenario asks for a single model.

163
MCQmedium

A developer is using the Vertex AI Gemini API to generate product descriptions. They get a 400 error 'INVALID_ARGUMENT: The model's maximum input token limit is 8192.' What is the most likely issue?

A.The prompt is too long
B.The API key is invalid
C.The output tokens are too high
D.The model is not available in the region
AnswerA

The 400 INVALID_ARGUMENT error explicitly names the 8192-token maximum input limit, so the submitted prompt exceeds that ceiling. Truncating or chunking the input resolves it; the constraint is input length, not output length or quota.

Why this answer

The 400 error 'INVALID_ARGUMENT: The model's maximum input token limit is 8192' explicitly indicates that the combined token count of the prompt (system instructions, user input, and any conversation history) exceeds the 8192-token context window of the Gemini model being used. This is a hard limit enforced by the Vertex AI Gemini API, and the error is triggered before any generation begins. Therefore, the most likely issue is that the prompt is too long.

Exam trap

The trap here is that candidates confuse input token limits with output token limits or general API authentication errors, but the specific error message 'maximum input token limit' directly points to prompt length as the root cause.

How to eliminate wrong answers

Option B is wrong because an invalid API key would result in a 401 Unauthorized or 403 Forbidden error, not a 400 INVALID_ARGUMENT error related to token limits. Option C is wrong because the error message specifically mentions 'input token limit', not output tokens; output token limits are enforced separately (e.g., via max_output_tokens parameter) and would produce a different error. Option D is wrong because model availability in a region would cause a 404 or 403 error (e.g., 'Model not found' or 'Permission denied'), not a token-limit-related INVALID_ARGUMENT error.

164
MCQeasy

A startup wants to quickly add a conversational AI assistant to its mobile app without managing any infrastructure. They need a managed API that provides access to Gemini models for chat and content generation. Which Google Cloud offering should they use?

A.Gemini API in Google AI Studio
B.Vertex AI Model Garden
C.Vertex AI Studio
D.Dialogflow CX
AnswerA

The Gemini API, accessible via Google AI Studio, provides a managed API to integrate Gemini models into applications with minimal setup. It is ideal for developers who want to add conversational AI without managing infrastructure, exactly matching the startup's requirement for a quick, managed solution for their mobile app.

Why this answer

The Gemini API in Google AI Studio offers a managed, serverless way to integrate Gemini models into applications, eliminating infrastructure management. It is designed for developers who need quick access to conversational AI capabilities. The other options are either prototyping tools, model catalogs, or specialized agent-building platforms that require more setup and do not provide the same level of managed API access.

Exam trap

The trap here is confusing Vertex AI Studio's prototyping interface with a production-ready managed API for application integration.

165
MCQmedium

A media company is building an application that must summarize 400-page legal contracts. Their current model has a context window of only 32,000 tokens, and truncating the documents loses critical clauses. They want to use a Gemini model on Vertex AI that can ingest the entire contract in a single request. Which Gemini model capability should they select for this workload?

A.A Gemini model fine-tuned on legal contracts with supervised tuning
B.A Gemini model variant with a 1 million token context window
C.A Gemini model variant with a 32,000 token context window plus a larger output token limit
D.A Gemini model with grounding enabled through Vertex AI Search
AnswerB

Gemini 1.5 Pro and Gemini 1.5 Flash on Vertex AI support a context window of up to 1 million tokens, which is sufficient to hold a 400-page contract and its instructions without chunking or truncation. Selecting this variant directly addresses the requirement of processing the whole document in one request.

Why this answer

The blocker is input capacity: a 400-page contract far exceeds 32,000 tokens. Gemini 1.5 Pro and Gemini 1.5 Flash on Vertex AI offer context windows up to 1 million tokens, letting the application pass the full document plus instructions in one call. Output limits, grounding, and fine-tuning change generation length, factual grounding, or behavior respectively, but none of them enlarges the input window.

Exam trap

The trap here is assuming that a larger output token limit or fine-tuning expands how much input the model can read, when context window size is a separate input constraint.

166
MCQhard

A hospital network wants to transcribe clinician dictation and then have a generative model produce structured discharge summaries from those transcripts, all within Google Cloud. Which combination of offerings matches this workflow?

A.Cloud Speech-to-Text followed by Vertex AI Gemini
B.Document AI followed by Cloud Vision API
C.Cloud Text-to-Speech followed by Vertex AI Gemini
D.Vertex AI Gemini followed by Cloud Speech-to-Text
AnswerA

Speech-to-Text converts the clinician's spoken dictation into written transcripts, and Gemini on Vertex AI can then transform those transcripts into structured discharge summaries. This sequence matches the workflow exactly, moving from audio to text and then from text to a generated structured document within Google Cloud.

Why this answer

The workflow requires converting spoken dictation into text and then generating a structured summary from that text. Speech-to-Text handles the audio transcription, and Gemini on Vertex AI performs the generative summarization, so this pairing correctly sequences the two capabilities. The alternatives invert the audio direction or substitute document and image services that cannot transcribe dictation.

Exam trap

The trap here is mixing up Text-to-Speech with Speech-to-Text, since the two names are easily reversed and only one converts dictation into written text.

167
Multi-Selecthard

A bank is evaluating Google Cloud generative AI offerings to build an internal document-processing application. Leadership requires that the solution support grounding responses in the bank's own document repository and provide enterprise controls such as IAM-based access and audit logging. Which two Google Cloud offerings should the bank consider to meet these requirements? (Choose two.)

Select 2 answers
A.Gemini for Google Workspace
B.Google AI Studio with the Gemini API
C.Vertex AI Search with enterprise data stores for grounding
D.Vertex AI with Gemini models and grounding capabilities
E.Gemini Enterprise as a standalone assistant for analysts
AnswersC, D

Vertex AI Search supports enterprise search and grounding over an organization's own content using data stores, and it operates under Cloud IAM and audit logging. That makes it a legitimate candidate for grounding document-processing responses in the bank's repository while preserving the governance controls leadership demanded.

Why this answer

Vertex AI Search with enterprise data stores and Vertex AI with Gemini grounding both let an organization ground model responses in its own documents while operating under Cloud IAM and audit logging. Those two offerings align with the bank's dual requirement of private-repository grounding and enterprise governance for a custom application, whereas prototyping tools and end-user assistant products do not.

Exam trap

The trap here is assuming any Gemini-branded service can ground on private documents under enterprise controls, when prototyping APIs and end-user assistants lack the application-level IAM and audit surface required.

168
MCQhard

A media company is using Vertex AI Imagen to generate marketing images. The output frequently contains unrealistic artifacts, especially in human faces. The team has fine-tuned the model using their brand assets. What is the most likely cause and recommended fix?

A.Safety filters are too aggressive; reduce them.
B.Negative prompts are missing; always include 'unrealistic'.
C.The fine-tuning dataset is too small or too homogeneous; augment and diversify the training data.
D.Inference steps are too low; increase to 100.
AnswerC

Fine-tuning on a narrow, homogeneous brand dataset biases the model toward those patterns, degrading general facial structure and producing artefacts. Augmenting with diverse, larger image sets restores the visual distribution the model needs, correcting the unrealistic faces while retaining brand style.

Why this answer

Unrealistic artifacts in fine-tuned generative models, especially in human faces, typically stem from a training dataset that is too small or lacks diversity. When the dataset is homogeneous, the model overfits to limited patterns and fails to generalize, leading to distorted outputs. Augmenting and diversifying the training data with varied poses, lighting, and ethnicities helps the model learn robust facial features.

Exam trap

The trap here is that candidates confuse inference parameters (like steps or safety filters) with data quality issues, assuming artifacts are due to model settings rather than the fundamental cause of insufficient or non-diverse training data.

How to eliminate wrong answers

Option A is wrong because safety filters in Vertex AI Imagen block harmful content (e.g., violence, hate speech) and do not cause unrealistic artifacts; reducing them would not fix facial distortions and could introduce policy violations. Option B is wrong because negative prompts guide the model to avoid certain concepts, but simply including the word 'unrealistic' is not a technical fix—the model needs diverse training data, not a prompt hack. Option D is wrong because inference steps control the denoising process and image quality, but increasing them to 100 would not address overfitting from a poor dataset; the default steps (typically 50) are sufficient for high-quality outputs.

169
MCQhard

A developer receives the above JSON response from a Vertex AI language model. The output content is correct, but the developer expected the model to not answer geography questions. What should the developer do to prevent the model from responding to geography queries?

A.Adjust the safety filter thresholds for the 'Toxic' category
B.Enable Vertex AI Grounding with a geography knowledge base
C.Configure a safety filter for the 'Geography' category
D.Add a system instruction to not answer geography questions
AnswerD

System instructions constrain model behaviour at inference time without retraining, directly satisfying the requirement to block geography responses. Unlike safety filters or fine-tuning, they apply per-request guidance that the model honours across turns, making them the appropriate control for steering a Vertex AI language model away from a chosen topic.

Why this answer

Vertex AI does not have a predefined safety filter category for 'Geography'. To prevent the model from answering geography questions, the developer should use a system instruction that explicitly tells the model not to respond to geography queries. System instructions are the appropriate mechanism for custom content restrictions, while safety filters are limited to predefined categories such as toxic, harassment, etc.

Exam trap

Candidates often assume safety filters can be configured for arbitrary topics like geography, but Vertex AI safety filters only support predefined categories. The correct approach is to use system instructions to guide model behavior for custom restrictions.

How to eliminate wrong answers

Option A is wrong because adjusting safety filter thresholds for the 'Toxic' category only controls responses related to toxicity (e.g., hate speech, harassment), not geography-specific content; it does not address the requirement to block geography questions. Option B is wrong because enabling Vertex AI Grounding with a geography knowledge base would actually enhance the model's ability to answer geography queries by providing additional context, which is the opposite of what the developer wants. Option D is wrong because while adding a system instruction to not answer geography questions might influence the model, it is not a guaranteed enforcement mechanism—models can still override or ignore instructions, especially if the prompt is rephrased; safety filters provide a more reliable, configurable block.

170
Multi-Selectmedium

Which TWO actions can reduce the cost of using Vertex AI Gemini API? (Choose two.)

Select 2 answers
A.Use batch prediction instead of online
B.Increase the max output tokens
C.Use grounding with Google Search
D.Use a larger model
E.Use context caching
AnswersA, E

Batch prediction processes many prompts asynchronously in a single job, avoiding the per-request overhead and premium pricing of synchronous online calls. This directly satisfies the stem's cost-reduction constraint, since Vertex AI charges less per token for batch workloads than for real-time online inference.

Why this answer

Option A is correct because batch prediction processes many requests asynchronously in a single job and is priced at a discount (typically 50%) compared to online prediction, directly lowering per-request cost for workloads that tolerate latency. Option E is correct because context caching lets you store frequently reused input tokens (e.g., long system prompts or documents) and pay a reduced rate for cached tokens plus a small storage fee, cutting costs when the same context is sent repeatedly. Option B is wrong because increasing max output tokens generates more billable output tokens, raising cost.

Option C is wrong because grounding with Google Search adds grounding charges and does not reduce API cost. Option D is wrong because larger models have higher per-token prices, increasing rather than reducing cost.

Exam trap

Candidates often mistakenly believe that increasing max output tokens or using a larger model improves quality without cost impact, but both directly increase token consumption and per-token pricing.

171
Multi-Selecthard

Which THREE considerations are critical when deploying a generative AI model using Vertex AI Endpoints for a latency-sensitive application? (Choose THREE.)

Select 3 answers
A.Model size and architecture
B.Number of model versions
C.GPU type and number
D.Autoscaling configuration
E.Number of model instances
AnswersA, C, D

Larger models introduce higher latency.

Why this answer

Model size and architecture directly impact inference latency because larger models with more parameters require more computation per request. For latency-sensitive applications, choosing a smaller or distilled model (e.g., Gemma 2B vs. 27B) or using quantization can reduce response times. Vertex AI Endpoints serve the model as-is, so the model's inherent computational cost is the primary driver of per-request latency.

Exam trap

Google Cloud often tests the distinction between configuration choices that affect latency (GPU type, autoscaling, model size) versus operational or lifecycle management choices (version count, manual instance count) that do not directly impact per-request response time.

172
MCQmedium

A logistics company wants employees to ask natural-language questions about shipment trends and have Gemini generate SQL, charts, and narrative summaries directly against data already stored in BigQuery, without moving or duplicating that data. Which Google Cloud capability should the company adopt?

A.Vertex AI Model Garden with a deployed open model
B.Gemini for Google Workspace connected to a BigQuery export sheet
C.Gemini in BigQuery
D.Exporting BigQuery tables to Cloud Storage, then querying them with a custom Gemini application
AnswerC

Gemini in BigQuery brings assistant capabilities directly into the BigQuery experience, generating SQL, explaining results, and helping produce visualizations and summaries over data that stays in place. Because the company wants natural-language analysis against existing BigQuery tables without copying them, this embedded offering matches the requirement precisely.

Why this answer

Gemini in BigQuery is designed to help users write SQL, understand results, and produce summaries and visualizations using data that remains in BigQuery. Since the logistics team wants natural-language insight without duplicating datasets, the embedded BigQuery assistant is the direct fit rather than exports, external deployments, or workspace-side copies.

Exam trap

The trap here is reaching for a generic Gemini model deployment when the requirement is specifically in-place assistance over BigQuery data with schema awareness.

173
MCQmedium

A global bank wants to use Gemini models in Vertex AI to summarize sensitive customer emails. The security team requires that prompts and responses never leave the bank's controlled network perimeter and that access is restricted to approved projects. Which Google Cloud capability should they configure?

A.VPC Service Controls
B.Cloud Armor
C.Cloud CDN
D.Cloud Interconnect
AnswerA

VPC Service Controls create a service perimeter that restricts access to Google Cloud services such as Vertex AI, preventing data exfiltration outside the defined boundary. By placing Vertex AI inside a perimeter with approved projects, the bank ensures prompts and responses cannot leave the controlled network, meeting the security team's requirement.

Why this answer

VPC Service Controls let organizations define a service perimeter around Google Cloud services, including Vertex AI. Resources inside the perimeter can communicate, but access from outside is blocked, which prevents data exfiltration. Combined with IAM policies limiting access to approved projects, this satisfies the bank's requirement that sensitive prompts and responses remain within a controlled boundary.

Exam trap

The trap here is assuming that private connectivity or edge security services such as Cloud Interconnect or Cloud Armor prevent data exfiltration from managed AI APIs, when perimeter controls are what actually enforce service boundaries.

174
MCQmedium

An e-commerce company wants to add a conversational shopping assistant to its mobile app. The assistant must answer product questions using the company's catalog, call a backend API to check live inventory, and escalate to a human agent when the customer requests it. Which Google Cloud offering is designed for this?

A.Gemini for Google Workspace
B.Vertex AI Agent Builder
C.Vertex AI Pipelines
D.Vertex AI Model Garden
AnswerB

Vertex AI Agent Builder is designed to create conversational agents that ground responses in enterprise data, invoke tools or APIs through function calling, and support handoff to human agents. It fits the need to answer catalog questions, check live inventory via a backend API, and escalate on request. This combination of grounding, tool use, and escalation is exactly its purpose.

Why this answer

Vertex AI Agent Builder provides the components to build conversational agents with grounding in enterprise data, function calling for backend APIs, and escalation paths to human agents. It directly supports catalog-grounded answers, live inventory checks, and handoff, matching the scenario. The other services are model catalogs, batch pipelines, or productivity tools that do not deliver agent orchestration for a mobile assistant.

Exam trap

The trap here is assuming a model catalog or pipeline can act as a conversational agent, when agent orchestration, tool calls, and human handoff require Agent Builder.

175
MCQmedium

A logistics company needs a generative AI model that can accept both text and images as input, and produce text output for describing shipping damage. They want to use a Google Cloud model through the Vertex AI API. Which Gemini model capability should they select?

A.Gemini 1.5 Flash with text-only input
B.Vertex AI Vision product for image classification
C.Gemini 1.5 Pro with multimodal input support
D.Vertex AI Embeddings API for text and image vectors
AnswerC

Gemini 1.5 Pro accepts interleaved text, image, and video input and returns text, which matches the requirement to describe shipping damage from photos plus written context. Calling it through the Vertex AI API gives the company managed access to this multimodal capability without hosting the model itself.

Why this answer

A multimodal Gemini model such as Gemini 1.5 Pro can take both text and images and return text, directly satisfying the need to describe shipping damage from photos and written notes. The other services either restrict input to text, produce vectors instead of language, or focus on vision classification rather than generative text output.

Exam trap

The trap here is assuming any Gemini model automatically handles images, when the input modality must be explicitly supported and configured for the chosen model.

176
MCQmedium

You are a generative AI architect for a large e-commerce company. Your team has built a product description generator using Vertex AI's text-bison model. The model is accessed via the Vertex AI API from a web application. You have set the temperature to 0.5 and top_k to 40. The team reports that the generated descriptions are often too generic and lack creativity. They want the descriptions to be more diverse and engaging. You are also concerned about cost, as each API call is billed. Which change should you recommend to increase creativity while managing cost?

A.Keep temperature at 0.5 but reduce top_k to 20.
B.Increase the temperature to 0.8 and keep top_k at 40.
C.Switch to a larger model like text-bison@002 and keep same parameters.
D.Decrease the temperature to 0.2 and increase top_k to 60.
AnswerB

Raising temperature to 0.8 flattens the probability distribution, so lower-probability tokens are sampled more often, producing more diverse and engaging descriptions. Keeping top_k at 40 preserves the existing candidate pool and call volume, so Vertex AI API billing per call stays unchanged.

Why this answer

Increasing the temperature to 0.8 makes the model's output probability distribution flatter, which increases randomness and allows less likely tokens to be selected. This directly addresses the need for more diverse and creative descriptions. Keeping top_k at 40 ensures the model still considers a broad set of candidate tokens, balancing creativity with coherence, and does not increase API call costs since temperature and top_k are inference parameters that do not affect billing.

Exam trap

Google Cloud often tests the misconception that increasing creativity requires a larger model or more expensive resources, when in fact tuning sampling parameters like temperature and top_k is the correct, cost-neutral approach.

How to eliminate wrong answers

Option A is wrong because reducing top_k to 20 narrows the set of candidate tokens, which actually reduces diversity and can make outputs more generic, counteracting the goal of increasing creativity. Option C is wrong because switching to a larger model like text-bison@002 would increase cost per API call (larger models are billed at higher rates) without guaranteeing more creativity; creativity is controlled by sampling parameters, not model size alone. Option D is wrong because decreasing temperature to 0.2 makes the model more deterministic and conservative, reducing creativity, and increasing top_k to 60 does not compensate for the loss of randomness — the net effect is less diverse outputs.

177
MCQmedium

An organization is using Vertex AI Agent Builder to create a customer service agent. They want the agent to be able to hand off to a human agent when it cannot answer a question. What should they configure in the agent's design?

A.Configure 'Slot filling' to collect more info
B.Implement a 'Confirmation' prompt for the user
C.Add an 'Escalation' intent that triggers a human handoff
D.Use a 'Fallback' intent to route to a human
AnswerD

Fallback intent is for unrecognized inputs, not specifically for human handoff.

Why this answer

In Vertex AI Agent Builder (Dialogflow CX), when the agent cannot answer or match a user's request, the conversation triggers a fallback/no-match path. To hand off to a human, you configure that fallback path to route to a live agent via fulfillment, webhook, or live-agent handoff. There is no standard built-in 'Escalation' intent in the product; escalation is an outcome you implement, not a specific intent type.

Exam trap

Candidates may be tempted by the plausible-sounding 'Escalation' option, but the configurable mechanism for a human handoff when the agent cannot answer is the fallback/no-match path routed to a human.

How to eliminate wrong answers

Option A is wrong because 'Slot filling' is used to collect missing information for an intent, not to escalate. Option B is wrong because a 'Confirmation' prompt asks the user to confirm an action, not to hand off. Option D is wrong because a 'Fallback' intent handles unrecognized input but typically provides a default response or reprompt, not necessarily a human handoff unless explicitly configured to do so.

178
MCQmedium

A data scientist uses Vertex AI Model Evaluation to assess a fine-tuned model for sentiment analysis. The evaluation report shows high precision but low recall on the 'negative' class. What is the best course of action to improve recall without sacrificing too much precision?

A.Adjust the prediction threshold for the negative class
B.Switch to a different model architecture (e.g., from BERT to RoBERTa)
C.Collect more labeled examples of negative sentiment and retrain
D.Use a larger pretrained model from Model Garden
AnswerC

Low recall on the negative class means the model misses true negatives, usually from under-representation. Adding more labelled negative examples rebalances the training distribution, lifting recall while preserving precision better than threshold or class-weight tweaks alone.

Why this answer

Collecting more labeled examples of negative sentiment and retraining addresses the root cause of low recall: insufficient or imbalanced training data for the negative class. This improves the model's ability to recognize negative sentiment without sacrificing precision, as the decision boundary is refined with more representative data. Option A (adjusting prediction threshold) can increase recall but typically at the cost of precision, contradicting the goal.

Option B (switching model architecture) is excessive and may not fix data imbalance. Option D (using a larger pretrained model) does not specifically target recall on the negative class.

179
MCQmedium

A global logistics company wants to build a generative AI assistant that can answer operational questions by retrieving information from internal PDF and HTML documents stored in Cloud Storage. They need a managed, serverless retrieval-augmented generation (RAG) capability that requires minimal infrastructure management and integrates with Vertex AI. Which Google Cloud service should they use?

A.Vertex AI Feature Store
B.Cloud SQL with pgvector
C.Vertex AI Search
D.BigQuery ML
AnswerC

Vertex AI Search is a fully managed, serverless search and retrieval service that supports RAG by grounding Gemini responses in enterprise data from Cloud Storage, websites, and other sources. It handles indexing, chunking, and retrieval without requiring the team to manage vector databases or embeddings pipelines, making it the ideal fit for a low-ops RAG solution.

Why this answer

Vertex AI Search delivers a fully managed, serverless retrieval-augmented generation pipeline that ingests documents from Cloud Storage, builds an index, and grounds Gemini responses without requiring the team to manage embeddings or vector databases. The other services are either feature stores, self-managed vector databases, or analytics tools that do not provide end-to-end RAG for unstructured content.

Exam trap

The trap here is assuming that any Google Cloud service that can store embeddings, such as Cloud SQL with pgvector, is a suitable RAG solution, when the requirement is a managed, serverless retrieval service.

180
MCQmedium

A media company wants its editorial team to summarize long internal reports and ask follow-up questions about the content, all inside a Google Workspace interface they already use daily. They prefer not to build any custom application or call APIs. Which Google Cloud generative AI offering should they adopt?

A.Vertex AI Studio
B.Google Cloud Speech-to-Text
C.Cloud Natural Language API
D.Gemini for Google Workspace
AnswerD

Gemini for Google Workspace embeds generative AI directly into Docs, Gmail, Drive, and Meet, so an editorial team can summarize long documents and ask follow-up questions without writing code or building an application. It matches the requirement of working inside an existing Workspace interface, which is exactly the consumption model this offering provides to business users.

Why this answer

Gemini for Google Workspace is the offering designed to bring Gemini capabilities directly into the productivity apps an organization already uses, including document summarization and conversational assistance. Because the team wants those abilities inside Docs and Gmail without developing a custom application, the embedded Workspace assistant fits, whereas developer consoles and narrow pretrained APIs do not match the consumption model or the feature set requested.

Exam trap

The trap here is assuming any Gemini-powered product can summarize documents, when only the Workspace-embedded offering delivers that inside Docs and Gmail without custom development.

181
MCQhard

A game development studio wants to create dynamic non-player character (NPC) dialogues that adapt to player choices. They need a Google Cloud service that allows them to build conversational agents with custom logic and integrate with their game backend. Which service should they use?

A.Vertex AI Prediction
B.Vertex AI Agent Builder
C.Cloud Functions
D.Dialogflow CX
AnswerB

Vertex AI Agent Builder provides a framework to create conversational agents with custom logic, tool integration, and backend connectivity. It supports building complex dialogue flows that can adapt based on player input, and it can call external APIs to fetch game state. This makes it ideal for dynamic NPC dialogues that respond to player choices. The studio can define intents, entities, and fulfillment logic to create immersive interactions.

Why this answer

Vertex AI Agent Builder is designed for creating generative AI-powered conversational agents with custom logic and backend integration. It enables dynamic, context-aware dialogues that can adapt to player choices, making it the best fit for the game studio. Other options either lack generative capabilities or are not tailored for building conversational agents.

Exam trap

The trap here is assuming Dialogflow CX is the only conversational AI tool, overlooking Vertex AI Agent Builder's generative and integration strengths.

182
MCQhard

A software vendor embeds Gemini in its SaaS product and needs each tenant's prompts and outputs isolated so that one tenant's data never appears in another tenant's responses. The vendor also wants usage tracked per tenant for billing. Which design should the team implement on Google Cloud?

A.One project with per-tenant data stores and a shared service account whose credentials are distributed to all tenant workloads.
B.A single shared Google Cloud project and one Gemini endpoint, with tenant identifiers included in the prompt text and per-tenant reporting derived from exported logs.
C.A separate Google Cloud project per tenant with its own Vertex AI resources and service accounts, plus labels or billing export to attribute usage per tenant.
D.One project with VPC Service Controls and per-tenant Cloud Armor rules, without separate identities or data stores for tenants.
AnswerC

Separating tenants into distinct Google Cloud projects creates a clear identity and resource boundary, so one tenant's prompts, data stores, and service accounts cannot be reached by another. Labels on requests and billing export then let the vendor attribute consumption per tenant, satisfying both the isolation and usage-tracking requirements with standard Google Cloud governance primitives.

Why this answer

Strong tenant isolation comes from separating identities and resources, which separate Google Cloud projects per tenant provide, backed by distinct service accounts and Vertex AI resources. Labels and billing export then make consumption attributable per tenant. Prompt-level tagging, shared credentials, or network-only controls do not create the data or identity boundaries multi-tenant isolation requires.

Exam trap

The trap here is believing that including a tenant ID in the prompt or relying on network controls creates real tenant isolation, when only identity and resource separation does.

← PreviousPage 3 of 3 · 182 questions total

Ready to test yourself?

Try a timed practice session using only Google Cloud Gen Ai Offerings questions.