Courseiva

Google Cloud Generative AI Leader Generative AI Leader (Generative AI Leader) — Questions 976–1008

1008 questions total · 14pages · All types, answers revealed

Page 13

Page 14 of 14

976
Multi-Selecthard

A company is moving a GenAI proof-of-concept to production. They need to ensure the system can handle variable traffic and maintain low latency. Which THREE practices should they implement? (Choose 3)

Select 3 answers
A.Implement response caching for common queries
B.Enable auto-scaling for the serving infrastructure
C.Reduce the input context length to the absolute minimum
D.Set up monitoring and alerting on latency metrics
E.Use a single, large instance to handle all traffic
AnswersA, B, D

Caching reduces latency and cost by reusing responses for identical requests.

Why this answer

Response caching stores the outputs of frequently requested queries, allowing the system to serve them instantly without recomputation, which drastically reduces latency for repeated requests and offloads the underlying model. Auto-scaling for the serving infrastructure lets the system dynamically add or remove capacity as traffic varies, preserving low latency during peaks while avoiding over-provisioning during lulls. Monitoring and alerting on latency metrics provides the observability needed to detect degradation early and trigger operational responses before users are affected.

Together these three practices address both variable traffic and low-latency requirements in production.

Exam trap

Google often tests the misconception that minimizing input context length universally improves performance, ignoring the trade-off with output quality, and that a single large instance is simpler and sufficient for production traffic, overlooking scalability and fault tolerance requirements.

977
MCQeasy

What is the primary benefit of using embeddings and vector search in a generative AI application?

A.They improve the model's ability to generate code
B.They reduce the size of the model by compressing weights
C.They enable efficient retrieval of semantically similar content
D.They allow the model to process images directly
AnswerC

Embeddings map text into vectors where semantic similarity corresponds to geometric proximity, letting vector search retrieve passages by meaning rather than exact keyword match. This enables efficient retrieval of semantically similar content to ground generative responses.

Why this answer

Embeddings convert text into dense vector representations that capture semantic meaning, and vector search enables efficient retrieval of semantically similar content by finding nearest neighbors in vector space. This retrieval-augmented generation (RAG) approach grounds the generative AI model in relevant external knowledge, improving accuracy and reducing hallucinations without retraining.

Exam trap

The trap here is that candidates confuse embeddings and vector search with model optimization or multimodal capabilities, when in fact they are a retrieval mechanism for grounding generative outputs in external knowledge. In Google Cloud, this is implemented via Vertex AI Vector Search or the Embeddings API for tasks like RAG.

How to eliminate wrong answers

Option A is wrong because embeddings and vector search are not specifically designed to improve code generation; they enhance retrieval of any text or data, but code generation benefits more from specialized training data and fine-tuning. Option B is wrong because embeddings and vector search do not reduce model size or compress weights; they operate on the input/output side, storing vectors separately, while model compression is achieved through techniques like pruning or quantization. Option D is wrong because embeddings and vector search primarily handle text or other data types via vector representations, not direct image processing; images require separate vision encoders or multimodal models to be processed directly.

978
MCQeasy

What is the key advantage of using vector search for retrieval in a RAG system compared to keyword search?

A.Vector search eliminates the need for a foundation model
B.Vector search can find conceptually similar documents even without exact keyword matches
C.Vector search is faster than keyword search
D.Vector search requires no preprocessing of documents
AnswerB

Vector search embeds queries and documents into a shared space and retrieves by embedding proximity, so semantically related passages surface even when wording differs. Keyword search matches literal terms only, failing the stem's need to find conceptually similar documents without exact keyword overlap.

Why this answer

Vector search in a RAG system encodes documents and queries into dense vector embeddings using a foundation model, then retrieves documents based on semantic similarity in the embedding space. This allows it to find conceptually related documents even when they share no exact keywords with the query, overcoming the lexical gap that limits keyword search.

Exam trap

A common misconception tested in this exam is that vector search is faster than keyword search, but the trap is that while vector search excels at semantic matching, it incurs higher latency and computational overhead compared to the simple inverted index lookup of keyword search.

How to eliminate wrong answers

Option A is wrong because vector search actually requires a foundation model (or an embedding model) to generate the vector representations; it does not eliminate the need for one. Option C is wrong because vector search is generally slower than keyword search due to the computational cost of embedding generation and approximate nearest neighbor (ANN) search, though it offers better recall. Option D is wrong because vector search requires preprocessing of documents to generate and store embeddings, which is a significant upfront step.

979
MCQmedium

A media company is using a foundation model in Vertex AI to summarize long articles. They notice that summaries sometimes omit key details from the middle of very long articles. Which action should they take to improve summary completeness while staying within the model's context window?

A.Add a system instruction telling the model to always include all key details.
B.Split the article into chunks, summarize each chunk, then combine the summaries in a final prompt.
C.Switch to a smaller model variant to reduce latency and cost.
D.Increase the model's temperature setting to encourage more creative output.
AnswerB

Chunking and hierarchical summarization keeps each request within the context window and ensures middle content is processed. By summarizing segments separately and then synthesizing, the model can capture details that would otherwise be dropped due to long-context limitations. This approach is a standard pattern for handling documents that exceed a single prompt's capacity while maintaining completeness.

Why this answer

Long articles can exceed a model's effective context, causing middle content to be overlooked. Chunking the article, summarizing each part, and then combining those summaries ensures all sections are processed and represented. This hierarchical approach stays within context limits and improves completeness, directly addressing the observed omission of key details.

Exam trap

The trap here is thinking that prompt wording or temperature can force a model to retain content that falls outside its effective context window.

980
MCQeasy

A marketing team wants to generate social media posts from product descriptions using Generative AI. They need consistent brand tone and the ability to iterate quickly. Which tool is BEST suited for this task?

A.Duet AI in Google Docs
B.Vertex AI Model Garden
C.Vertex AI Studio
D.Vertex AI Agent Builder
AnswerC

Vertex AI Studio gives prompt design, tuning and rapid iteration against Gemini models in one console, letting the team lock a consistent brand tone through reusable prompts and test variations quickly. It satisfies both the tone-consistency and fast-iteration constraints directly.

Why this answer

Vertex AI Studio is the correct choice because it provides a purpose-built environment for prompt engineering, model tuning, and rapid iteration with foundation models. It allows the marketing team to experiment with prompts, adjust parameters like temperature and top-p, and maintain consistent brand tone through saved prompt templates and versioning, directly supporting the need for quick iteration.

Exam trap

Google often tests the distinction between a general-purpose AI assistant (Duet AI) and a dedicated prompt engineering platform (Vertex AI Studio), leading candidates to choose Duet AI because they confuse document assistance with content generation.

How to eliminate wrong answers

Option A is wrong because Duet AI in Google Docs is an AI-powered assistant for document creation and editing, not a tool for generating social media posts from product descriptions with iterative prompt engineering. Option B is wrong because Vertex AI Model Garden is a repository for discovering and deploying pre-trained models, but it lacks the integrated prompt engineering and iterative testing environment needed for fine-tuning brand tone. Option D is wrong because Vertex AI Agent Builder is designed for building conversational agents and chatbots, not for generating and iterating on social media content from product descriptions.

981
MCQeasy

A developer needs to integrate a GenAI model into an existing customer relationship management (CRM) system. The CRM exposes REST APIs and runs on-premises. Which integration pattern is MOST suitable?

A.Use Apps Script to call the model from Google Sheets
B.Fine-tune a model on CRM data and deploy it on-premises
C.Implement an API-first integration by calling Vertex AI API from the CRM's backend
D.Build a custom Google Workspace add-on
AnswerC

Calling the Vertex AI API from the CRM's backend preserves the on-premises REST architecture, so no inbound connectivity or data migration is required. The API-first pattern lets the existing CRM orchestrate model inference server-side, satisfying the on-premises constraint while keeping credentials and customer data within the trusted boundary.

Why this answer

API-first integration using Vertex AI API allows any system with HTTP capabilities to call the GenAI model. Workspace add-ons and Apps Script are for Google Workspace only. Fine-tuning doesn't help with integration.

982
MCQeasy

A logistics firm wants every generative AI proposal to be judged on whether it reduces cost per shipment. Leadership asks the AI team to define the metric before any project starts. Which practice does this illustrate?

A.Selecting a foundation model based on its published benchmark scores.
B.Adopting a responsible AI review board to approve all generative AI use cases.
C.Choosing a deployment region that minimizes network egress charges.
D.Establishing a measurable business outcome tied to an existing operational key performance indicator.
AnswerD

Tying the generative AI effort to cost per shipment connects the initiative to an operational metric the business already tracks and trusts. That makes value demonstrable, comparable across proposals, and defensible to finance, which is exactly what defining the metric before work begins is meant to achieve.

Why this answer

Value-driven generative AI programs start by naming the business outcome and linking it to a metric the organization already measures, such as cost per shipment. That anchor lets leadership compare proposals, size expected returns, and confirm impact after launch. Model benchmarks, egress savings, and governance approvals are supporting concerns rather than definitions of business value.

Exam trap

The trap here is mistaking a technical or governance activity for a business value definition, when value must be expressed in an operational outcome the business already tracks.

983
MCQhard

A company has deployed a GenAI-powered report generation system using Vertex AI. They notice that the cost is higher than expected. Investigation shows that many requests include very long prompts with repetitive boilerplate text. Which cost optimization strategy is MOST effective?

A.Increase the batch size for batch requests
B.Enable context caching for repeated prompt prefixes
C.Switch to a smaller model size
D.Ignore the cost increase as it will stabilize
AnswerB

Context caching stores the processed repeated prompt prefix so Vertex AI reuses it across requests instead of reprocessing the boilerplate each time. This directly cuts input token charges, targeting the repetitive-prefix pattern identified as the cost driver.

Why this answer

Caching repeated prompt prefixes can significantly reduce token usage and cost. Batch requests help with throughput but not with per-request token savings. Reducing model size may hurt quality.

Ignoring is not a strategy.

984
MCQeasy

A logistics firm's leadership wants a plain-language summary of how generative AI could reduce costs in dispatch operations before approving any budget. The team has no data scientists available. Which first step best aligns with a generative AI business strategy?

A.Ask each regional dispatch manager to independently experiment with public consumer chatbots and report anecdotes.
B.Commission a multi-year research program to train a custom dispatch model on historical route data.
C.Purchase GPU hardware for an on-premises cluster so the firm controls its own model hosting from day one.
D.Run a short, scoped pilot using a managed generative AI model on a sample of dispatch tasks and report measured outcomes.
AnswerD

A scoped pilot with a managed model produces concrete evidence, such as time saved per dispatch decision, without requiring data science hires or large capital. Because managed services remove infrastructure work, a small operations team can execute it. The measured results then give leadership the plain-language, cost-focused summary they requested before any budget commitment.

Why this answer

A short pilot on a managed generative AI model converts an abstract idea into measured dispatch outcomes, such as handling time or routing decision quality, using the team already in place. That evidence lets leadership evaluate cost impact in plain language before committing budget, which is the essence of a staged generative AI business strategy.

Exam trap

The trap here is equating credible generative AI adoption with acquiring hardware or launching research, when a small measured pilot is what actually informs an approval decision.

985
MCQhard

A legal firm wants to automate contract analysis to extract key clauses and risks. They have 10,000 contracts in PDF format. The solution must handle varying layouts and be cost-effective. Which approach is BEST?

A.Use Document AI to convert PDFs to structured text, then use Vertex AI with a prompt that specifies clauses to extract
B.Build a retrieval-augmented generation (RAG) system in Vertex AI Agent Builder
C.Use the Gemini API directly with the raw PDF files as input
D.Fine-tune a Gemini model on 100 annotated contracts and run inference on all contracts
AnswerA

Document AI handles layout parsing, and the structured text is then fed to a GenAI model with a well-designed prompt. This combination is scalable and cost-effective.

Why this answer

Using Document AI to parse PDFs into text, then Vertex AI with a structured prompt for clause extraction combines robust document understanding with flexible GenAI. Fine-tuning on 100 contracts is insufficient for layout variation. Agent Builder is overkill.

Direct Gemini API on raw PDFs loses document structure.

986
Multi-Selecthard

Which THREE are valid methods to reduce bias in generative AI outputs?

Select 3 answers
A.Using only English prompts
B.Increasing model size
C.Using a more diverse training dataset
D.Using safety filters
E.Applying prompt engineering to instruct the model to be fair
AnswersC, D, E

A more diverse training dataset directly addresses representation bias by exposing the model to broader demographic, cultural and linguistic variation during pre-training, reducing the likelihood of skewed associations in generated outputs. This satisfies the stem's requirement for a valid bias-reduction method, since bias frequently originates from unrepresentative or historically skewed training corpora.

Why this answer

Option C is correct because a more diverse training dataset reduces representational bias by exposing the model to a wider range of demographics, cultures, languages, and viewpoints, so the learned distribution is less skewed toward a dominant group. Option D is correct because safety filters (e.g., content moderation classifiers, toxicity detectors, or bias guardrails applied to prompts and outputs) can detect and block biased or harmful generations, mitigating bias at inference time. Option E is correct because prompt engineering that explicitly instructs the model to be fair, neutral, or inclusive (e.g., system prompts specifying balanced representation) steers generation toward less biased outputs without retraining.

Option A is not correct because restricting prompts to English narrows the linguistic and cultural scope, which can amplify bias rather than reduce it. Option B is not correct because increasing model size alone does not remove bias and can even amplify biases present in the training data.

Exam trap

Google Cloud often tests the misconception that increasing model size or using a single language (like English) can solve bias, when in reality these actions can worsen bias by amplifying existing skews or introducing new cultural blind spots.

987
MCQmedium

A data scientist is fine-tuning a large language model for a specialized domain using limited labeled data. To avoid catastrophic forgetting and reduce computational cost, which approach is recommended?

A.Using prompt engineering with in-context learning
B.Full fine-tuning of all model parameters
C.Training a new model from scratch on the domain data
D.Adapter-based fine-tuning using LoRA
AnswerD

LoRA freezes the base weights and trains small low-rank adapter matrices, so the original model's general capabilities are preserved, avoiding catastrophic forgetting. Training only adapters also cuts computational cost, satisfying the limited labelled data and cost constraints.

Why this answer

Adapter-based fine-tuning using LoRA (Low-Rank Adaptation) is recommended because it freezes the pre-trained model weights and injects trainable low-rank matrices into the transformer layers. This approach drastically reduces the number of parameters to update (often by 10,000x), lowering memory and compute requirements, while preserving the original knowledge to prevent catastrophic forgetting on limited domain data.

Exam trap

The Generative AI Leader exam often tests the misconception that prompt engineering (Option A) is a form of fine-tuning, when in fact it is a zero-shot or few-shot inference technique that does not modify model parameters, making it unsuitable for persistent domain adaptation with limited labeled data.

How to eliminate wrong answers

Option A is wrong because prompt engineering with in-context learning does not update model weights, so it cannot adapt the model to a specialized domain with limited labeled data in a persistent manner; it relies on the model's existing knowledge and context window, which is insufficient for deep domain adaptation. Option B is wrong because full fine-tuning of all model parameters updates the entire model, which is computationally expensive and, with limited labeled data, risks catastrophic forgetting of the original pre-trained knowledge. Option C is wrong because training a new model from scratch on domain data requires a massive amount of labeled data and compute resources, defeating the purpose of leveraging a pre-trained LLM and reducing cost.

988
MCQhard

A company is deploying a generative AI application that generates medical reports. They need to ensure the output is factual and minimizes hallucinations. Which approach is most effective?

A.Fine-tune the model with RLHF
B.Set the temperature to 0.0
C.Implement retrieval-augmented generation (RAG) with a curated knowledge base
D.Use prompt engineering to instruct the model to be accurate
AnswerC

RAG grounds generation in retrieved documents from a curated knowledge base, so the model conditions output on verified source content rather than parametric memory alone. This directly reduces hallucination and improves factual accuracy, satisfying the medical report requirement.

Why this answer

Retrieval-Augmented Generation (RAG) is the most effective approach because it grounds the model's output in a curated, authoritative knowledge base of medical data. By retrieving relevant, verified documents at inference time, RAG directly reduces the model's reliance on its parametric memory, which is the primary source of hallucinations in generative AI. This is especially critical in high-stakes domains like medical reporting, where factual accuracy is paramount.

Exam trap

The trap here is that candidates often choose 'Set the temperature to 0.0' because they confuse reducing randomness with eliminating factual errors, but temperature only controls output variability, not the truthfulness of the model's internal knowledge.

How to eliminate wrong answers

Option A is wrong because RLHF (Reinforcement Learning from Human Feedback) optimizes the model for human preference alignment and helpfulness, but it does not provide a mechanism to retrieve or verify facts from an external source, so it cannot reliably prevent hallucinations in factual domains. Option B is wrong because setting temperature to 0.0 makes the model deterministic (always picking the highest-probability token), but it does not correct factual errors stored in the model's weights; the model can still confidently generate false information. Option D is wrong because prompt engineering instructs the model to be accurate, but it is a soft constraint that the model can easily override; without external grounding, the model has no way to verify its own output against a trusted source.

989
MCQmedium

A global news agency is using a generative AI model to summarize breaking news articles in real-time. The model is deployed on Vertex AI across multiple regions (us-central1, europe-west4, asia-southeast1) for low latency worldwide. The agency has a Service Level Objective (SLO) of 99.9% availability and p99 latency under 2 seconds. Recently, during a major event, traffic spiked 10x, and the europe-west4 region experienced latency spikes over 5 seconds and some 503 errors. The team suspects the regional endpoint is under-provisioned. Which combination of actions should they take to meet the SLO consistently?

A.Enable the global endpoint feature in Vertex AI with automatic traffic splitting, and increase the minimum replicas for each regional endpoint
B.Increase the maximum replicas for the europe-west4 endpoint and reduce the min replicas in other regions
C.Implement Cloud CDN caching for common summaries and reduce the number of regions to two
D.Configure a global load balancer with a single Vertex AI endpoint and increase max replicas globally
AnswerA

Global endpoint distributes traffic and increases capacity; higher min replicas prevent cold starts during spikes.

Why this answer

It enables the global endpoint feature with automatic traffic splitting, allowing traffic to be routed to healthy regions and providing failover. Additionally, increasing minimum replicas per region ensures each regional endpoint has baseline capacity to handle spikes, preventing under-provisioning. Option B only increases max replicas in europe-west4, which does not address traffic shifts, and reducing min replicas elsewhere risks capacity issues.

Option C suggests Cloud CDN, which is for static content, not model inference. Option D configures a global load balancer with a single endpoint, which does not optimally use Vertex AI's regional endpoints and may not meet latency SLO.

990
MCQhard

A large enterprise runs a generative AI solution serving millions of daily inference requests. To reduce costs, they propose using serverless endpoints (Vertex AI Prediction) with a custom container, but they notice high latency during cold starts. Which strategy best addresses this problem while minimizing cost?

A.Set a minimum number of replicas to maintain a baseline of always-on instances.
B.Upgrade to GPU-accelerated machines for all replicas.
C.Implement client-side request batching to reduce the number of inference calls.
D.Use prewarmed containers by setting an idle timeout to keep instances alive.
AnswerA

Correct. Setting a minimum number of replicas ensures a baseline of always-on instances, which eliminates cold starts for the majority of requests. This directly addresses the latency spike caused by container initialization and model loading, and the cost is limited to the minimum replicas rather than scaling all instances.

Why this answer

Setting a minimum number of replicas ensures that a baseline of always-on instances is maintained, eliminating cold starts for the majority of requests. This directly addresses the latency spike caused by container initialization and model loading in serverless endpoints, while the cost impact is limited to the minimum replicas rather than scaling all instances.

Exam trap

Google Cloud often tests the misconception that prewarming via idle timeout is a configurable parameter in serverless ML services, but in Vertex AI Prediction, the idle timeout is fixed and not user-adjustable, making minimum replicas the correct approach.

How to eliminate wrong answers

Option B is wrong because upgrading to GPU-accelerated machines increases cost significantly without solving cold start latency; GPUs primarily improve per-request throughput, not initialization time. Option C is wrong because client-side request batching reduces the number of inference calls but does not affect cold start latency; it may even increase perceived latency for individual requests. Option D is wrong because setting an idle timeout to keep instances alive is not a supported mechanism in Vertex AI Prediction; the service uses an internal keep-alive policy, and user-configurable idle timeouts are not available, making this option technically infeasible.

991
MCQeasy

A marketing agency uses generative AI to create social media posts. They need the output to be in a specific JSON format for downstream processing. Which prompt technique should they use?

A.Set the temperature to 0.0
B.Use a model fine-tuned for JSON generation
C.Include a system instruction to output JSON
D.Specify the desired JSON schema in the prompt and use few-shot examples
AnswerD

Declaring the JSON schema in the prompt constrains the model's output structure, while few-shot examples demonstrate the exact key-value pattern expected. Together they steer generation toward valid, parseable JSON, satisfying the downstream processing requirement for a specific format.

Why this answer

Explicitly specifying the desired JSON schema in the prompt combined with few-shot examples provides the most reliable way to enforce structured output from a generative AI model. This technique leverages the model's pattern-matching ability by showing concrete input-output pairs, which is more effective than vague instructions or parameter adjustments alone for achieving exact JSON formatting.

Exam trap

Google often tests the misconception that a simple parameter change (like temperature) or a vague instruction is sufficient to enforce structured output, when in reality explicit schema definition with examples is required for reliable JSON generation.

How to eliminate wrong answers

Option A is wrong because setting temperature to 0.0 only reduces randomness in token selection, making output more deterministic, but does not guarantee the model will output valid JSON or follow a specific schema—it can still produce malformed or non-JSON text. Option B is wrong because while a fine-tuned model may improve JSON generation, fine-tuning requires significant data and compute resources, and the question asks for a prompt technique, not a model customization approach. Option C is wrong because a system instruction to output JSON is a weak constraint—models often ignore or partially follow system instructions, especially for complex schemas, leading to missing fields, extra keys, or incorrect nesting.

992
Multi-Selectmedium

A financial institution is implementing a generative AI chatbot to handle customer inquiries. The institution must comply with regulatory requirements (e.g., GDPR, SOX) and ensure data privacy. Which TWO actions should the institution take?

Select 2 answers
A.Establish a Center of Excellence (CoE) for AI governance to oversee model deployment and monitoring.
B.Use Vertex AI without additional data governance controls to simplify deployment.
C.Use a pre-trained model without customization to reduce development time.
D.Implement model validation and testing to ensure outputs meet regulatory standards.
E.Deploy the model on-premises only to keep data within local infrastructure.
AnswersA, D

A CoE centralises AI governance, giving the financial institution the oversight structure needed to enforce GDPR and SOX controls across model deployment and monitoring. It assigns accountability, standardises review gates and ensures privacy requirements are applied consistently.

Why this answer

Option A is correct because establishing a Center of Excellence (CoE) for AI governance provides the oversight, policy enforcement, and monitoring needed to keep a generative AI chatbot compliant with GDPR and SOX, covering model deployment, risk management, and accountability. Option D is correct because model validation and testing are essential to verify that chatbot outputs meet regulatory standards, detect bias or data leakage, and ensure ongoing compliance before and after deployment. Option B is incorrect because using Vertex AI without additional data governance controls fails to address GDPR and SOX privacy and audit requirements.

Option C is incorrect because a pre-trained model without customization does not by itself satisfy regulatory compliance or data privacy obligations. Option E is incorrect because on-premises deployment alone does not guarantee compliance and may not be feasible or sufficient for the institution's regulatory and operational needs.

Exam trap

The trap here is choosing 'on-premises only' as a silver bullet for data privacy — candidates conflate physical data location with regulatory compliance, but GDPR and SOX require governance, validation, and auditability regardless of where the model runs.

993
MCQhard

A company has been using an on-premises ML infrastructure for generative AI and wants to migrate to Google Cloud. They have a pipeline that fine-tunes a large language model weekly using a proprietary dataset. The migration must minimize downtime and data transfer costs. Which approach best addresses these requirements?

A.Use Vertex AI Pipelines to orchestrate the fine-tuning process, and use Vertex AI Managed Datasets to incrementally sync new data with BigQuery as the source.
B.Use AutoML to train a new model directly from the dataset without fine-tuning.
C.Deploy the existing pipeline on a Google Kubernetes Engine cluster and use Google Cloud Filestore for shared storage.
D.Use Cloud Storage Transfer Service to move all data to Cloud Storage, then set up a Vertex AI custom training job to run the fine-tuning.
AnswerA

Vertex AI Pipelines offers a managed orchestration service that can schedule the weekly fine-tuning workflow with minimal operational overhead. Combined with Vertex AI Managed Datasets, which incrementally sync new data from BigQuery, this approach reduces data transfer costs by avoiding full dataset copies and minimizes downtime by enabling automated, scheduled execution without manual intervention.

Why this answer

Vertex AI Pipelines provides a managed, serverless orchestration service that can run the weekly fine-tuning workflow with minimal operational overhead, while Vertex AI Managed Datasets can incrementally sync new data from BigQuery, reducing data transfer costs by avoiding full dataset copies. This combination minimizes downtime because the pipeline can be triggered on a schedule without manual intervention, and incremental syncs avoid re-transferring the entire proprietary dataset each week.

Exam trap

The trap here is that candidates often assume full data migration (e.g., Cloud Storage Transfer Service) is necessary, overlooking incremental sync capabilities of Vertex AI Managed Datasets with BigQuery, which directly addresses cost and downtime minimization.

How to eliminate wrong answers

Option B is wrong because AutoML is designed for training models from scratch on labeled data, not for fine-tuning an existing large language model, and it would require a full dataset transfer, increasing costs and downtime. Option C is wrong because deploying the existing pipeline on GKE with Filestore for shared storage does not address data transfer costs (Filestore still requires initial data migration) and introduces additional operational complexity for managing Kubernetes clusters, which does not minimize downtime compared to a managed service. Option D is wrong because using Cloud Storage Transfer Service to move all data to Cloud Storage incurs high initial data transfer costs and does not leverage incremental sync capabilities, and the custom training job setup lacks the orchestration and scheduling benefits of Vertex AI Pipelines, leading to more downtime during migration.

994
MCQmedium

A global logistics company wants to build a generative AI assistant that can answer operational questions by retrieving information from internal PDF and HTML documents stored in Cloud Storage. They need a managed, serverless retrieval-augmented generation (RAG) capability that requires minimal infrastructure management and integrates with Vertex AI. Which Google Cloud service should they use?

A.Vertex AI Feature Store
B.Cloud SQL with pgvector
C.Vertex AI Search
D.BigQuery ML
AnswerC

Vertex AI Search is a fully managed, serverless search and retrieval service that supports RAG by grounding Gemini responses in enterprise data from Cloud Storage, websites, and other sources. It handles indexing, chunking, and retrieval without requiring the team to manage vector databases or embeddings pipelines, making it the ideal fit for a low-ops RAG solution.

Why this answer

Vertex AI Search delivers a fully managed, serverless retrieval-augmented generation pipeline that ingests documents from Cloud Storage, builds an index, and grounds Gemini responses without requiring the team to manage embeddings or vector databases. The other services are either feature stores, self-managed vector databases, or analytics tools that do not provide end-to-end RAG for unstructured content.

Exam trap

The trap here is assuming that any Google Cloud service that can store embeddings, such as Cloud SQL with pgvector, is a suitable RAG solution, when the requirement is a managed, serverless retrieval service.

995
MCQmedium

A media company wants its editorial team to summarize long internal reports and ask follow-up questions about the content, all inside a Google Workspace interface they already use daily. They prefer not to build any custom application or call APIs. Which Google Cloud generative AI offering should they adopt?

A.Vertex AI Studio
B.Google Cloud Speech-to-Text
C.Cloud Natural Language API
D.Gemini for Google Workspace
AnswerD

Gemini for Google Workspace embeds generative AI directly into Docs, Gmail, Drive, and Meet, so an editorial team can summarize long documents and ask follow-up questions without writing code or building an application. It matches the requirement of working inside an existing Workspace interface, which is exactly the consumption model this offering provides to business users.

Why this answer

Gemini for Google Workspace is the offering designed to bring Gemini capabilities directly into the productivity apps an organization already uses, including document summarization and conversational assistance. Because the team wants those abilities inside Docs and Gmail without developing a custom application, the embedded Workspace assistant fits, whereas developer consoles and narrow pretrained APIs do not match the consumption model or the feature set requested.

Exam trap

The trap here is assuming any Gemini-powered product can summarize documents, when only the Workspace-embedded offering delivers that inside Docs and Gmail without custom development.

996
MCQhard

An organization wants to ensure that when their employees use the Gemini API via Vertex AI, the grounding searches are restricted to internal company knowledge bases rather than the public web. Which feature should they enable?

A.Enterprise Data Governance on Vertex AI
B.VPC Service Controls
C.Google Search Grounding
D.Vertex AI Search (with private data indexing)
AnswerD

Vertex AI Search indexes private data sources, so grounding queries hit internal knowledge bases instead of the public web. This satisfies the restriction constraint directly, since the API's default grounding would otherwise reach public sources.

Why this answer

Vertex AI Search allows organizations to index private, internal data sources (e.g., documents, databases) and use them as the grounding source for Gemini API queries. By enabling this feature, grounding searches are restricted to the indexed private knowledge base, ensuring no public web results are used. This directly meets the requirement to keep grounding internal.

Exam trap

The trap here is that candidates may confuse 'Google Search Grounding' (which uses the public web) with Vertex AI Search's private data grounding, leading them to select Option C instead of recognizing that Vertex AI Search with private data indexing is the correct solution for internal grounding.

How to eliminate wrong answers

Option A is wrong because Enterprise Data Governance on Vertex AI provides data access controls and audit logging, but it does not restrict grounding sources for Gemini API queries. Option B is wrong because VPC Service Controls create a security perimeter around Google Cloud resources, but they do not control the grounding source used by the Gemini API. Option C is wrong because Google Search Grounding explicitly enables grounding via the public web, which is the opposite of the requirement to restrict to internal knowledge bases.

997
MCQeasy

A data scientist wants to compare the performance of three different foundation models for a text summarization task. They have a labeled dataset of summaries. Which Vertex AI tool should they use to perform this evaluation?

A.Vertex AI RAG Engine - Retrieval evaluation
B.Vertex AI Agent Builder - Agent evaluation
C.Model Garden - Model comparison view
D.Vertex AI Studio - Evaluation
AnswerD

Vertex AI Studio's evaluation capability scores generative model outputs against a labelled reference dataset using metrics such as ROUGE for summarisation. Running all three foundation models through it produces comparable, quantitative results, satisfying the requirement to compare summarisation performance.

Why this answer

Vertex AI Studio includes a built-in Evaluation feature that lets you compare foundation models side-by-side using your own labeled dataset, computing metrics like ROUGE, BLEU, and human-preference scores for tasks such as summarization. It is purpose-built for evaluating generative model outputs against ground-truth references, which matches the data scientist's need exactly.

Exam trap

The trap here is confusing model discovery (Model Garden) or pipeline-specific evaluation (RAG Engine, Agent Builder) with general-purpose generative model evaluation, which lives in Vertex AI Studio.

How to eliminate wrong answers

Option A is wrong because the RAG Engine retrieval evaluation only measures retrieval quality (e.g., recall, precision of retrieved chunks) for RAG pipelines, not the summarization quality of three different foundation models. Option B is wrong because Agent Builder's agent evaluation assesses tool-calling, task completion, and reasoning traces of agents, not text summarization outputs. Option C is wrong because Model Garden's model comparison view is for browsing and discovering models (metadata, pricing, capabilities), not for running quantitative evaluation against a labeled dataset.

998
MCQhard

A game development studio wants to create dynamic non-player character (NPC) dialogues that adapt to player choices. They need a Google Cloud service that allows them to build conversational agents with custom logic and integrate with their game backend. Which service should they use?

A.Vertex AI Prediction
B.Vertex AI Agent Builder
C.Cloud Functions
D.Dialogflow CX
AnswerB

Vertex AI Agent Builder provides a framework to create conversational agents with custom logic, tool integration, and backend connectivity. It supports building complex dialogue flows that can adapt based on player input, and it can call external APIs to fetch game state. This makes it ideal for dynamic NPC dialogues that respond to player choices. The studio can define intents, entities, and fulfillment logic to create immersive interactions.

Why this answer

Vertex AI Agent Builder is designed for creating generative AI-powered conversational agents with custom logic and backend integration. It enables dynamic, context-aware dialogues that can adapt to player choices, making it the best fit for the game studio. Other options either lack generative capabilities or are not tailored for building conversational agents.

Exam trap

The trap here is assuming Dialogflow CX is the only conversational AI tool, overlooking Vertex AI Agent Builder's generative and integration strengths.

999
MCQhard

A regulated healthcare organization needs to use Google Cloud AI services for processing protected health information (PHI). They require a signed Business Associate Agreement (BAA) and HIPAA compliance. Which service provides these assurances?

A.Google AI Studio
B.Colab Enterprise
C.Gemini API directly
D.Vertex AI
AnswerD

Vertex AI is covered by Google Cloud's HIPAA BAA, permitting protected health information processing under a signed agreement. This satisfies the healthcare organisation's requirement for both a BAA and HIPAA-compliant infrastructure for its AI workloads.

Why this answer

Vertex AI is the correct choice because it is the only Google Cloud AI service that offers a signed Business Associate Agreement (BAA) and is explicitly covered under Google Cloud's HIPAA compliance framework. This allows regulated healthcare organizations to process protected health information (PHI) with the necessary contractual and security assurances.

Exam trap

The trap here is that candidates may assume the Gemini API alone is HIPAA-compliant because it is a Google service, but in reality, only when accessed through Vertex AI (which provides the BAA and enterprise controls) does it meet healthcare regulatory requirements.

How to eliminate wrong answers

Option A is wrong because Google AI Studio is a free, web-based prototyping tool that does not offer a BAA or HIPAA compliance, and it is not designed for production workloads involving PHI. Option B is wrong because Colab Enterprise is a managed notebook environment that, while it can use Google Cloud resources, does not itself provide a BAA or HIPAA compliance for PHI processing. Option C is wrong because the Gemini API directly, when accessed via the standard API endpoint, does not include a BAA or HIPAA compliance; only when used through Vertex AI (which wraps the Gemini API) does the BAA and HIPAA coverage apply.

1000
MCQeasy

Which Google DeepMind achievement is recognized for predicting protein structures and advancing drug discovery?

A.AlphaFold
B.AlphaGo
C.Gemini
D.AlphaCode
AnswerA

AlphaFold, developed by Google DeepMind, predicts three-dimensional protein structures from amino acid sequences using deep learning. Accurate structure prediction accelerates drug discovery by revealing binding sites and protein interactions that previously required costly laboratory methods such as crystallography.

Why this answer

AlphaFold is the correct answer because it is a groundbreaking AI system developed by Google DeepMind that predicts the 3D structure of proteins from their amino acid sequences with atomic-level accuracy. This achievement has revolutionized computational biology by solving a 50-year-old grand challenge in molecular biology, enabling significant advances in drug discovery, enzyme design, and understanding disease mechanisms.

Exam trap

The Generative AI Leader exam often tests the distinction between domain-specific AI achievements (like AlphaFold for biology) versus general-purpose AI models (like Gemini) or game-playing AI (like AlphaGo), so candidates may confuse the application area of each DeepMind project.

How to eliminate wrong answers

Option B is wrong because AlphaGo is an AI program that mastered the board game Go using deep reinforcement learning and Monte Carlo tree search, not protein structure prediction. Option C is wrong because Gemini is a multimodal large language model family designed for text, image, audio, and video understanding, not for structural biology or drug discovery. Option D is wrong because AlphaCode is an AI system for competitive programming that generates code solutions to algorithmic problems, not for predicting protein structures.

1001
MCQeasy

A marketing team uses a text generation model to draft campaign copy. They want the output to consistently follow a specific brand voice and include a call to action at the end of every draft. Which technique should they apply?

A.Provide a system instruction that defines the brand voice and requires a call to action, plus one or two exemplar drafts.
B.Set the temperature to zero so the model always produces the same deterministic wording for every campaign.
C.Increase the top-k value so the model samples from a broader set of words and naturally adopts a more distinctive voice.
D.Shorten the prompt to only the product name so the model has maximum freedom to express the brand personality.
AnswerA

A system instruction sets persistent behavioral rules such as tone and mandatory elements, and exemplars demonstrate the exact style. Together they reliably steer the model toward the brand voice and ensure the call to action appears. This is a prompt-level control that needs no training and can be updated quickly as brand guidelines evolve.

Why this answer

System instructions plus exemplars are the standard way to enforce tone and mandatory content without training. The instruction defines persistent rules, and examples show the desired pattern. Sampling changes affect randomness, and shorter prompts remove guidance, so neither reliably produces consistent brand voice and a guaranteed call to action.

Exam trap

The trap here is equating determinism or broader sampling with stylistic control, when only explicit instructions and examples define voice and required elements.

1002
MCQeasy

Which Google service provides free access to Jupyter notebooks with GPU support for prototyping ML models?

A.Colab
B.Kaggle Notebooks
C.Vertex AI Workbench
D.BigQuery Studio
AnswerA

Colab supplies hosted Jupyter notebooks with free GPU runtimes, letting developers prototype ML models without provisioning infrastructure. It directly satisfies the stem's requirement for free notebook access with GPU support, unlike paid or non-notebook alternatives.

Why this answer

Google Colab is a free, hosted Jupyter notebook environment that provides access to GPUs (and TPUs) for prototyping machine learning models without any setup. It integrates with Google Drive and allows users to write and run Python code in the browser. This directly matches the requirement for free Jupyter notebooks with GPU support.

Exam trap

The trap is confusing Colab with Vertex AI Workbench or Kaggle Notebooks — the exam tests that Colab is the free, GPU-enabled Jupyter environment from Google, while Workbench is the paid enterprise option.

How to eliminate wrong answers

Option B is wrong because Kaggle Notebooks, while offering free GPU/TPU notebooks, is a Kaggle platform feature rather than a Google service in the sense the question intends, and Colab is the canonical Google answer for free Jupyter with GPU. Option C is wrong because Vertex AI Workbench is an enterprise managed notebook service on Google Cloud that requires a billing account and is not free. Option D is wrong because BigQuery Studio is a data analytics workspace for SQL and notebooks focused on BigQuery data, not a free GPU-enabled ML prototyping environment.

1003
MCQeasy

A company wants to generate images for slide decks using Gemini in Google Slides. Which Gemini feature in Google Slides should they use?

A.Speaker notes generation
B.Slide layout optimization
C.Image generation from text prompts
D.Summation of slide content
AnswerC

Text-to-image generation converts typed prompts directly into visuals, satisfying the requirement to produce slide imagery without external tools. It operates within Slides, so users create deck graphics in place rather than sourcing stock images separately.

Why this answer

Gemini in Slides can generate images from text prompts, allowing users to create visuals directly within the presentation.

1004
MCQeasy

What is the primary purpose of the transformer architecture in large language models (LLMs)?

A.To generate images from text descriptions
B.To enable parallel processing of tokens and capture long-range dependencies through self-attention
C.To convert text into numerical embeddings for downstream tasks
D.To store and retrieve information from a vector database
AnswerB

Self-attention lets every token attend to every other token simultaneously, so the model captures long-range dependencies without sequential recurrence. This parallel processing across the whole sequence is what makes transformers trainable at the scale LLMs require, unlike RNN architectures that process tokens one at a time.

Why this answer

The transformer architecture's primary purpose is to enable parallel processing of all tokens in a sequence while capturing long-range dependencies through its self-attention mechanism. Unlike recurrent neural networks (RNNs) that process tokens sequentially, transformers compute attention scores between every pair of tokens simultaneously, allowing the model to weigh the relevance of distant tokens without the vanishing gradient problem. This parallelization and global context capture are the foundational innovations that make large language models (LLMs) scalable and effective for tasks like text generation and understanding.

Exam trap

The trap here is that candidates confuse the transformer's core innovation (parallel self-attention for sequence modeling) with auxiliary tasks like embedding generation or retrieval, which are separate components in an LLM pipeline, leading them to pick C or D as plausible but incorrect answers.

How to eliminate wrong answers

Option A is wrong because generating images from text descriptions is the domain of multimodal models like DALL·E or Stable Diffusion, which use diffusion or GAN architectures, not the transformer's primary purpose. Option C is wrong because converting text into numerical embeddings is a preprocessing step (e.g., tokenization and embedding layers) that occurs before the transformer processes tokens, but the transformer's core role is to model relationships between those embeddings via self-attention, not merely to create embeddings. Option D is wrong because storing and retrieving information from a vector database is a retrieval-augmented generation (RAG) technique that supplements LLMs with external knowledge; the transformer architecture itself does not inherently function as a database.

1005
MCQhard

A software vendor embeds Gemini in its SaaS product and needs each tenant's prompts and outputs isolated so that one tenant's data never appears in another tenant's responses. The vendor also wants usage tracked per tenant for billing. Which design should the team implement on Google Cloud?

A.One project with per-tenant data stores and a shared service account whose credentials are distributed to all tenant workloads.
B.A single shared Google Cloud project and one Gemini endpoint, with tenant identifiers included in the prompt text and per-tenant reporting derived from exported logs.
C.A separate Google Cloud project per tenant with its own Vertex AI resources and service accounts, plus labels or billing export to attribute usage per tenant.
D.One project with VPC Service Controls and per-tenant Cloud Armor rules, without separate identities or data stores for tenants.
AnswerC

Separating tenants into distinct Google Cloud projects creates a clear identity and resource boundary, so one tenant's prompts, data stores, and service accounts cannot be reached by another. Labels on requests and billing export then let the vendor attribute consumption per tenant, satisfying both the isolation and usage-tracking requirements with standard Google Cloud governance primitives.

Why this answer

Strong tenant isolation comes from separating identities and resources, which separate Google Cloud projects per tenant provide, backed by distinct service accounts and Vertex AI resources. Labels and billing export then make consumption attributable per tenant. Prompt-level tagging, shared credentials, or network-only controls do not create the data or identity boundaries multi-tenant isolation requires.

Exam trap

The trap here is believing that including a tenant ID in the prompt or relying on network controls creates real tenant isolation, when only identity and resource separation does.

1006
MCQmedium

A company deploying a generative AI assistant wants to allow users to override the AI's suggestions before final actions are taken. Which design pattern does this represent?

A.Human-on-the-loop
B.Human-in-the-loop
C.Human-out-of-the-loop
D.Automated decision-making
AnswerB

Human-in-the-loop places a person at the decision point, letting users review and override the assistant's suggestions before any final action executes. This directly satisfies the stem's requirement for user override prior to commitment, unlike fully automated patterns where the model's output is actioned without intervention.

Why this answer

The Human-in-the-loop (HITL) pattern ensures that a human reviews and can override the AI's suggestions before any final action is executed. This design explicitly maintains human oversight over critical decisions, preventing fully automated actions that could lead to harmful or unintended outcomes.

Exam trap

Google's AI Principles emphasize human oversight. Candidates must distinguish between Human-in-the-loop (active approval) and Human-on-the-loop (monitoring) to correctly identify scenarios where human intervention occurs before vs. after the AI action.

How to eliminate wrong answers

Option A is wrong because Human-on-the-loop implies the human monitors the AI's actions but does not actively intervene in real-time; the system acts autonomously unless the human steps in after the fact. Option C is wrong because Human-out-of-the-loop removes human oversight entirely, allowing the AI to make and execute decisions without any human review or override capability. Option D is wrong because Automated decision-making is a broad category that does not specify any human involvement; it could include fully autonomous systems, which contradicts the requirement for user override before final actions.

1007
MCQhard

An e-commerce company uses a generative AI model to generate product descriptions. They observe that descriptions for high-end products use more sophisticated language compared to budget products, potentially reinforcing class stereotypes. What is the most likely cause, and what should they do to mitigate it?

A.The training data reflects real-world associations; fine-tune with a balanced dataset that includes diverse product descriptions across price ranges
B.The safety filters are too aggressive; relax them
C.The temperature parameter is set too low; increase it to introduce more randomness
D.The model architecture is biased; switch to a different base model
AnswerA

Biased associations in the training corpus skew the model's language toward real-world price stereotypes. Fine-tuning on a balanced dataset containing diverse product descriptions across price ranges directly counteracts this, satisfying the stem's requirement to mitigate stereotype reinforcement at its source rather than masking outputs post-generation.

Why this answer

The bias stems from training data correlations. Fine-tuning on balanced, diverse data can reduce stereotypical associations. The other options either do not address the root cause or are less effective.

1008
MCQhard

A financial services company is using a large language model on Vertex AI to generate summaries of earnings reports. They notice that the model sometimes includes information not present in the source reports, leading to compliance risks. They want to reduce hallucinations by ensuring the model's output is grounded in the provided documents. Which technique should they implement?

A.Using a smaller model
B.Increasing the temperature
C.Retrieval-augmented generation (RAG)
D.Fine-tuning the model on earnings reports
AnswerC

RAG involves retrieving relevant passages from the source documents and providing them as context to the model, which reduces hallucinations by grounding the generation in factual content. This directly addresses the compliance risk by ensuring the summary is based on the earnings reports, not the model's internal knowledge.

Why this answer

Retrieval-augmented generation (RAG) is the most effective technique to ground model outputs in specific documents. By retrieving relevant passages from the earnings reports and including them in the prompt, the model is constrained to generate summaries based on that evidence, significantly reducing hallucinations and compliance risks.

Exam trap

The trap here is assuming fine-tuning eliminates hallucinations; it can adapt style but does not guarantee grounding in a specific input document.

Page 13

Page 14 of 14