Courseiva

Google Cloud Generative AI Leader Generative AI Leader (Generative AI Leader) — Questions 826–900

1008 questions total · 14pages · All types, answers revealed

Page 11

Page 12 of 14

Page 13
826
MCQeasy

A developer needs to use the Vertex AI PaLM API to generate text embeddings for a large corpus of documents. Which model should they use?

A.codey-bison@001
B.textembedding-gecko@001
C.text-bison@001
D.chat-bison@001
AnswerB

textembedding-gecko@001 is the PaLM-family embedding model on Vertex AI, purpose-built to convert text into dense vectors. It satisfies the embedding requirement for a large corpus, unlike generative text models such as text-bison, which produce completions rather than embeddings.

Why this answer

`textembedding-gecko@001` is the specific Vertex AI model designed for generating text embeddings, which convert text into dense vector representations. This model is optimized for semantic similarity, clustering, and retrieval tasks, making it ideal for processing a large corpus of documents. The other models are designed for code generation, text generation, or chat, not embeddings.

Exam trap

The trap here is that candidates may confuse general-purpose text generation models (like `text-bison@001`) with embedding models, assuming any 'text' model can produce embeddings, but only models with 'embedding' in the name are designed for that purpose.

How to eliminate wrong answers

Option A is wrong because `codey-bison@001` is a code generation model, not an embedding model; it generates code snippets or completes code, not vector representations of text. Option C is wrong because `text-bison@001` is a text generation model for tasks like summarization or content creation, not for producing embeddings. Option D is wrong because `chat-bison@001` is a conversational model designed for multi-turn dialogue, not for generating text embeddings.

827
Multi-Selectmedium

A company wants to use GenAI to generate marketing content such as blog posts and social media updates. They need the content to be on-brand and factually accurate. Which TWO features should they use?

Select 2 answers
A.Provide few-shot examples in the prompt to set the brand tone
B.Use a longer context window to include all brand guidelines
C.Fine-tune the model on a large corpus of past marketing content
D.Enable grounding with Google Search for factual accuracy
E.Use Vertex AI Studio to design prompts with no additional grounding
AnswersA, D

Few-shot examples in the prompt demonstrate the desired tone, structure and vocabulary, steering output toward the brand voice. This satisfies the stem's on-brand requirement by conditioning the model on concrete samples rather than relying on generic instructions alone.

Why this answer

Few-shot examples in prompts help maintain brand voice. Grounding with Google Search ensures factual accuracy. Vertex AI Studio is for prompt design but not directly for accuracy.

Fine-tuning may be overkill. Longer context may dilute the message.

828
MCQmedium

A company is building a multilingual customer support chatbot that needs to understand and respond in 20 languages. Which Google model is most suitable for this task?

A.Chirp
B.Imagen
C.Codey
D.Gemini
AnswerD

Gemini is natively multimodal and multilingual, handling 20 languages within one model without per-language pipelines or separate deployments. That built-in multilingual coverage directly satisfies the chatbot's requirement to understand and respond across all 20 languages.

Why this answer

Gemini is a multimodal large language model (LLM) designed for understanding and generating text across multiple languages, making it the most suitable choice for a multilingual customer support chatbot. Unlike specialized models, Gemini's architecture supports over 100 languages natively, enabling it to handle the 20-language requirement without needing separate language-specific models.

Exam trap

Google often tests the distinction between specialized models (like Chirp for audio, Imagen for images, Codey for code) and general-purpose multimodal LLMs (like Gemini) that can handle diverse tasks including multilingual text generation.

How to eliminate wrong answers

Option A is wrong because Chirp is a speech-to-text and text-to-speech model focused on audio processing, not on understanding or generating multilingual text for a chatbot. Option B is wrong because Imagen is a text-to-image generation model, not designed for natural language understanding or multilingual text responses. Option C is wrong because Codey is a code generation model specialized in programming languages and code completion, not in handling natural language conversations across multiple human languages.

829
MCQhard

A company is deploying a GenAI-powered email drafting feature. They want to control costs while maintaining low latency for real-time suggestions. Which strategy is MOST effective?

A.Batch all email drafting requests and run them every hour
B.Implement caching for frequently generated email drafts and use a smaller model variant for real-time requests
C.Use the largest available model and increase the number of tokens per request to generate more complete drafts
D.Use a large model with a longer context window to reduce the number of API calls
AnswerB

Caching repeated draft patterns avoids redundant inference calls, cutting cost and latency, while a smaller model variant serves real-time requests faster than a large model. Together these satisfy the stem's dual constraint of controlled cost and low latency for live suggestions.

Why this answer

Caching common prompt-output pairs reduces API calls for repeated inputs. Choosing a smaller model balances speed and cost. Batching is for offline processing, not real-time.

Long context is more expensive.

830
MCQeasy

A marketing team wants to use generative AI to draft product descriptions. Legal requires that no customer data or proprietary campaign plans appear in any prompt sent to the model. Which practice should the team adopt first?

A.Ask the model provider to sign a data processing addendum covering all prompts submitted by the team.
B.Configure the model with a low temperature to reduce the chance that sensitive text is reproduced.
C.Enable a content filter on the model output to block any sensitive information that appears in drafts.
D.Establish prompt input guidelines and a review step that excludes customer data and proprietary campaign details.
AnswerD

The requirement is about what enters the prompt, so the first control is defining what may and may not be included and verifying it before submission. Clear guidelines plus a review step directly prevent sensitive content from being sent, and they create an auditable practice the legal team can inspect.

Why this answer

Because the restriction concerns what is placed into prompts, the effective first step is to define allowable prompt content and review submissions against it. That prevents sensitive data from leaving the organization and creates evidence of compliance. Temperature settings, output filtering, and contractual addenda do not control what data is included in the prompt, so they cannot satisfy the stated legal constraint.

Exam trap

The trap here is confusing output-side controls with input-side data governance when the restriction is specifically about what is sent to the model.

831
MCQhard

A research team wants to fine-tune a Gemini model on Vertex AI using a dataset of proprietary scientific abstracts. They need to adjust the model's behavior with supervised fine-tuning while keeping the base model's general knowledge. Which Vertex AI capability should they use?

A.Vertex AI Pipelines
B.Vertex AI Vision
C.Supervised fine-tuning of Gemini on Vertex AI
D.Vertex AI Matching Engine
AnswerC

Supervised fine-tuning updates a Gemini model's weights using labeled input-output pairs, teaching it domain-specific patterns while retaining the base model's pretrained knowledge. This is exactly what the research team needs to adapt the model to scientific abstracts without losing general capabilities, and it is supported directly in Vertex AI.

Why this answer

Supervised fine-tuning on Vertex AI lets teams adapt Gemini models with their own labeled examples, adjusting behavior for specialized domains such as scientific literature. The tuned model retains the base model's broad knowledge while learning task-specific patterns. This is the managed capability designed for customizing Gemini, unlike vision, search, or orchestration services that serve different purposes.

Exam trap

The trap here is selecting an orchestration or retrieval service such as Vertex AI Pipelines or Matching Engine when the requirement is specifically to change model weights through supervised fine-tuning.

832
MCQmedium

A developer is using the Vertex AI PaLM API and receives a 429 Resource Exhausted error. What is the most likely cause?

A.The request payload is too large
B.The user has exceeded the allowed number of requests per minute
C.The model is not available in the current region
D.The API key is invalid
AnswerB

The 429 Resource Exhausted status signals quota exhaustion on the Vertex AI PaLM endpoint. Per-minute request quotas cap how many calls a project may issue; exceeding that rate triggers this error, so the cause is too many requests within the minute window.

Why this answer

HTTP 429 Resource Exhausted is the standard error returned when a client exceeds the rate limits or quota assigned to the API — most commonly the requests-per-minute (RPM) limit. In the Vertex AI PaLM API, each project has per-minute quotas for prediction requests, and exceeding them triggers this error. The correct remediation is to implement exponential backoff and/or request a quota increase.

Exam trap

Generative AI Leader often tests the mapping between HTTP status codes and their causes, and candidates frequently confuse 429 (rate/quota) with 400 (bad request) or 403 (permission), leading them to pick payload-size or API-key options.

How to eliminate wrong answers

Option A is wrong because an oversized payload typically returns a 400 Bad Request or 413 Payload Too Large error, not 429 Resource Exhausted — payload size is a request validity issue, not a rate/quota issue. Option C is wrong because a model unavailable in the current region returns a 404 Not Found or a region-specific error indicating the model is not supported, not a 429. Option D is wrong because an invalid API key returns 401 Unauthorized or 403 Forbidden, which are authentication/authorization errors, not resource exhaustion.

833
MCQmedium

A researcher wants to adapt a large language model for a specialized medical terminology domain without retraining the entire model. Which fine-tuning method is MOST parameter-efficient?

A.In-context learning with 50 examples
B.Adapter-based fine-tuning using LoRA
C.RLHF (Reinforcement Learning from Human Feedback)
D.Full supervised fine-tuning of all model weights
AnswerB

LoRA freezes the pretrained weights and injects small trainable low-rank matrices into attention layers, so only a tiny fraction of parameters update. Adapters likewise add compact modules. This adapts the model to medical terminology while avoiding full retraining, satisfying the parameter-efficiency constraint.

Why this answer

LoRA (Low-Rank Adaptation) is the most parameter-efficient fine-tuning method because it injects trainable low-rank matrices into the transformer layers, updating only a tiny fraction (often <1%) of the model's parameters while keeping the original weights frozen. This allows the model to adapt to specialized medical terminology without the memory and compute cost of full fine-tuning, making it ideal for domain adaptation with limited resources.

Exam trap

Google often tests the distinction between 'fine-tuning' and 'prompt engineering'—the trap here is that candidates mistake in-context learning (Option A) for a fine-tuning method because it adapts behavior, but it does not update model parameters, making it ineligible as a parameter-efficient fine-tuning technique.

How to eliminate wrong answers

Option A is wrong because in-context learning with 50 examples does not modify model weights at all; it relies on the prompt context window, which is limited in length and cannot reliably encode specialized medical terminology for consistent generation, making it a zero-shot/prompting technique, not a fine-tuning method. Option C is wrong because RLHF is a training paradigm that aligns model outputs with human preferences using a reward model, but it is not parameter-efficient—it typically requires full model fine-tuning or at least significant weight updates, and it is designed for alignment, not domain-specific knowledge injection. Option D is wrong because full supervised fine-tuning updates all model weights, which is extremely parameter-inefficient (requires storing and computing gradients for billions of parameters), prone to catastrophic forgetting, and demands substantial computational resources, contradicting the requirement for parameter efficiency.

834
MCQeasy

A team wants to build a GenAI application that can interact with external APIs (e.g., to check inventory or place orders). Which Vertex AI component provides this capability?

A.Grounding with Google Search
B.Model Garden
C.Vertex AI Agent Builder
D.Vertex AI Extensions
AnswerD

Vertex AI Extensions let the model call external APIs and services directly, enabling actions such as inventory checks and order placement. This satisfies the stem's requirement for API interaction, since Extensions provide the tool-calling layer rather than merely hosting or tuning models.

Why this answer

Vertex AI Extensions allows a GenAI application to connect to and interact with external APIs, enabling actions like checking inventory or placing orders. It provides a way to extend the model's capabilities by calling external services. This is the component specifically designed for such integrations.

Exam trap

The trap is confusing Agent Builder with Extensions. Agent Builder is for creating agents, but the specific component for API integration is Extensions. Candidates might also think Grounding with Google Search can call any API, but it's limited to web search.

How to eliminate wrong answers

Option A is wrong because Grounding with Google Search is used to ground model responses with web search results, not to call arbitrary external APIs. Option B is wrong because Model Garden is a repository of pre-trained models, not an integration tool. Option C is wrong because Vertex AI Agent Builder is for building conversational agents, but the specific capability to call external APIs is provided by Extensions; Agent Builder may use Extensions, but the question asks for the component that provides the capability.

835
MCQmedium

A retailer is building a product recommendation chatbot using Vertex AI Agent Builder. They want the agent to answer questions about product availability, prices, and promotions, but also to escalate to a human agent when the query is complex. What should they configure in Agent Builder?

A.Create a playbook with a step that transfers to a human via a webhook
B.Define an agent with a 'handoff to human' intent and configure the corresponding flow
C.Integrate a tool that calls a human support API when confidence is low
D.Use Vertex AI Agent Builder's generative fallback to automatically escalate
AnswerB

Agent Builder supports handoff to a human agent through intent and flow configuration.

Why this answer

In Vertex AI Agent Builder, you define intents to handle specific user requests. A 'handoff to human' intent, when triggered, initiates a configured flow that transfers the conversation to a human agent. This is the standard method for escalation.

Option A is incorrect because playbooks define conversation steps but not directly the escalation trigger; they can be part of a flow but not the primary escalation mechanism. Option C is incorrect: while you could integrate a tool to call a human support API, Agent Builder provides a built-in handoff intent and flow, making custom tool integration unnecessary. Option D is incorrect: generative fallback only generates responses when the agent cannot match a query, it does not provide a structured escalation to a human.

836
MCQmedium

A company is using Gemini to generate marketing copy. They want the outputs to be more creative and varied. Which generation parameters should they adjust?

A.Set temperature to 0 and top-k to 1
B.Decrease temperature and increase top-k
C.Increase temperature and adjust top-p to a higher value
D.Increase the context window length
AnswerC

Temperature scales the sampling distribution's randomness, while top-p restricts sampling to the smallest token set whose cumulative probability reaches the threshold. Raising both widens the candidate pool and flattens selection, directly producing the greater creativity and variation the stem requests.

Why this answer

Increasing temperature raises the randomness of token selection, making outputs more creative and varied, while adjusting top-p to a higher value (e.g., 0.9) allows the model to sample from a larger cumulative probability mass of likely tokens, further increasing diversity. Together, these parameters directly control the stochasticity of generation, which is essential for creative marketing copy.

Exam trap

Google often tests the misconception that increasing context length or adjusting a single parameter (like top-k) is sufficient for creativity, when in fact temperature and top-p must be increased together to achieve controlled randomness without generating gibberish.

How to eliminate wrong answers

Option A is wrong because setting temperature to 0 and top-k to 1 forces deterministic, greedy decoding (always picking the most probable token), which eliminates creativity and variation entirely. Option B is wrong because decreasing temperature reduces randomness, making outputs more focused and repetitive, which is the opposite of what is needed for creativity. Option D is wrong because increasing the context window length only allows the model to consider more input tokens (e.g., longer prompt or history), but does not affect the randomness or diversity of token selection during generation.

837
MCQmedium

A regional insurance company wants to launch a generative AI assistant that drafts policyholder responses for its claims team. Before any code is written, the CIO asks the team to produce a document that articulates the intended business outcome, the target user group, the success metrics, and the boundaries of what the assistant may and may not do. Which artifact best matches this request?

A.A Vertex AI model card documenting the training data, evaluation results, and known limitations of the underlying foundation model.
B.A generative AI use case definition that states the business objective, target users, measurable success criteria, and explicit scope boundaries.
C.A service-level objective document specifying latency percentiles and uptime targets for the assistant's API endpoint.
D.A cloud architecture diagram showing the assistant's front end, API layer, retrieval store, and model endpoint.
AnswerB

This artifact captures exactly what the CIO asked for: the business outcome, the intended user group, quantifiable success metrics, and the guardrails defining what the assistant should and should not handle. Framing these elements before implementation keeps the project anchored to measurable value and prevents scope drift once engineering begins, which is the foundational step of a generative AI business strategy on Google Cloud.

Why this answer

A generative AI initiative should begin with a use case definition that ties the technology to a business outcome, names the users it serves, defines how success will be measured, and sets explicit boundaries. This keeps the claims-drafting assistant aligned to measurable value and gives later technical and governance work a stable reference point before any build activity begins.

Exam trap

The trap here is assuming that a governance or architecture artifact such as a model card or diagram satisfies an upfront business framing request, when it describes the solution rather than the business problem.

838
Multi-Selectmedium

A company is deploying a generative AI model for medical diagnosis support. Which THREE considerations are critical for responsible AI?

Select 3 answers
A.Ensure the training data is diverse and representative.
B.Maximize model throughput to handle high volumes.
C.Implement human oversight for all diagnostic suggestions.
D.Provide clear disclaimers about the model's limitations.
E.Use the cheapest model to reduce costs.
AnswersA, C, D

Diverse, representative training data reduces demographic bias, preventing skewed diagnostic suggestions for under-represented patient groups. This satisfies the responsible AI requirement by addressing fairness at the data layer, where bias originates before model training begins.

Why this answer

Option A is correct because diverse and representative training data is essential to reduce bias and ensure the model performs equitably across different patient demographics, which is a core responsible AI principle in medical contexts. Option C is correct because human oversight for all diagnostic suggestions ensures that qualified clinicians review and validate AI outputs, maintaining accountability and preventing harm from erroneous predictions. Option D is correct because clear disclaimers about the model's limitations inform users of its scope and uncertainty, preventing over-reliance and supporting informed decision-making.

Option B is not correct because maximizing throughput addresses performance and scalability, not responsible AI concerns such as fairness, safety, or transparency. Option E is not correct because choosing the cheapest model prioritizes cost over accuracy, safety, and ethical considerations, which is inappropriate for medical diagnosis support.

Exam trap

Google Cloud often tests the distinction between operational metrics (like throughput or cost) and ethical/regulatory requirements (like fairness, transparency, and human oversight) in responsible AI, leading candidates to mistakenly select performance-based options as critical considerations.

839
MCQhard

A university wants to deploy a generative AI study assistant across six faculties. Each faculty has its own curriculum documents and privacy rules, and the central IT team must prevent one faculty's content from appearing in another faculty's answers. Which architecture decision best satisfies this governance requirement?

A.Fine-tune the model separately on each faculty's curriculum and expose a single endpoint that selects the adapter at random.
B.Deploy the assistant in a single project with one service account and apply IAM conditions based on the user's email domain.
C.Use one shared data store for all curricula and rely on prompt instructions telling the model to answer only from the requesting faculty's material.
D.Create a separate grounded data store and retrieval configuration per faculty, and route each request only to its own faculty's store.
AnswerD

Isolating each faculty's curriculum in its own grounded data store means retrieval can only return documents from the requesting faculty, so cross-faculty leakage is prevented by architecture rather than by instruction. Each faculty can also apply its own privacy rules to its store, and central IT retains one platform to operate across all six.

Why this answer

Separating grounded data stores per faculty enforces isolation at the retrieval layer, so answers can only draw on the requesting faculty's curriculum regardless of prompt wording. Each faculty keeps control of its own privacy rules, while central IT operates a single assistant platform. This structural separation is the only option that prevents leakage by design rather than by model compliance.

Exam trap

The trap here is trusting prompt instructions to enforce data boundaries, when retrieval isolation must be enforced in the architecture before the model ever sees the context.

840
MCQeasy

A startup wants to quickly prototype a conversational AI application using Gemini. They need free access during development and do not require VPC controls. Which access tier should they choose?

A.Vertex AI on a pay-as-you-go basis
B.Vertex AI with VPC-SC
C.Gemini API via Cloud Run
D.Google AI Studio free tier
AnswerD

Google AI Studio's free tier provides immediate, no-cost API access to Gemini models, satisfying the startup's prototyping budget constraint. It omits enterprise controls such as VPC Service Controls and data residency guarantees, which the stem explicitly does not require. This makes it the fastest path to building a conversational proof of concept.

Why this answer

Google AI Studio's free tier provides free access to Gemini models for rapid prototyping without requiring VPC controls or any billing setup. This aligns directly with the startup's need for quick, cost-free development iteration before moving to production.

Exam trap

The Generative AI Leader exam often tests the misconception that any Google Cloud service requires a billing account, but Google AI Studio's free tier explicitly bypasses this for prototyping, while options like Vertex AI or Cloud Run always incur costs even at low usage.

How to eliminate wrong answers

Option A is wrong because Vertex AI on a pay-as-you-go basis incurs costs from the start, which contradicts the requirement for free access during development. Option B is wrong because Vertex AI with VPC-SC adds unnecessary VPC security controls and costs, which the startup explicitly does not need. Option C is wrong because the Gemini API via Cloud Run requires a billing account and incurs compute and API usage costs, making it not free for prototyping.

841
Multi-Selecthard

Which THREE factors should be considered when choosing between a fine-tuned model and a prompted foundation model for a generative AI solution? (Select 3)

Select 3 answers
A.Need for domain-specific vocabulary
B.Inference latency requirements
C.Size of training data available
D.Whether the model is open-source
E.Token cost per request
AnswersA, C, E

Fine-tuning can incorporate domain language.

Why this answer

Fine-tuning allows the model to learn domain-specific vocabulary and terminology that may not be well-represented in the foundation model's pre-training data. This is critical for specialized fields like legal, medical, or technical domains where precise language is required for accurate outputs.

Exam trap

Google Cloud often tests the misconception that inference latency is a deciding factor between fine-tuning and prompting, when in reality both can be optimized for speed, and the key differentiators are data availability, domain specificity, and cost per token.

842
MCQeasy

What is the purpose of grounding in Vertex AI?

A.To improve training speed
B.To connect model outputs to verifiable sources
C.To reduce model size for faster inference
D.To enable multi-modal inputs
AnswerB

Grounding links model outputs to verifiable sources such as search results or enterprise data, reducing hallucination. This satisfies the need to trace generated claims back to authoritative evidence, ensuring responses are factually anchored rather than relying solely on parametric knowledge.

Why this answer

Grounding in Vertex AI connects model outputs to verifiable, external sources of information (such as Google Search, enterprise data sources, or third-party databases) to reduce hallucinations and improve factual accuracy. By referencing grounded sources, the model can provide citations and allow users to verify claims, which is critical for enterprise applications requiring trust and compliance.

Exam trap

Google Cloud often tests grounding by conflating it with fine-tuning or prompt engineering, so the trap here is assuming grounding modifies the model's weights or training process, when in fact it is a retrieval-based augmentation layer applied at inference time.

How to eliminate wrong answers

Option A is wrong because grounding does not improve training speed; it is a runtime technique applied during inference to augment responses with real-time data, not a training optimization. Option C is wrong because grounding does not reduce model size or accelerate inference; it may actually add latency due to the retrieval step. Option D is wrong because grounding is not about enabling multi-modal inputs; it specifically addresses output verification and source attribution, whereas multi-modal support is a separate capability for processing images, audio, or video alongside text.

843
MCQmedium

A company is building a customer service chatbot using Vertex AI Agent Builder. The chatbot needs to answer questions based on a large internal knowledge base stored in a Cloud Storage bucket. The team wants to ensure the model can reference the latest documents without fine-tuning. Which configuration should they use?

A.Fine-tune a model on the knowledge base documents
B.Use a pre-built model with no additional configuration
C.Store the documents in BigQuery and use a BigQuery connector
D.Ground the model with a Vertex AI Search data store connected to the Cloud Storage bucket
AnswerD

A Vertex AI Search data store indexes the Cloud Storage documents and grounds responses at query time, so the chatbot cites current content without retraining. This satisfies the no-fine-tuning and latest-documents constraints that embedding-only or static prompt approaches fail.

Why this answer

Grounding with a Vertex AI Search data store is the correct pattern for retrieval-augmented generation (RAG) in Vertex AI Agent Builder. The data store indexes the Cloud Storage documents and the agent retrieves relevant chunks at query time, so the model can cite the latest content without any fine-tuning. This keeps the knowledge base fresh — updating a document in the bucket automatically flows into subsequent answers.

Exam trap

Generative AI Leader often tests the distinction between fine-tuning (bakes knowledge into weights, static) and grounding/RAG (retrieves live external knowledge) — candidates pick fine-tuning when the scenario explicitly says 'without fine-tuning' or 'latest documents'.

How to eliminate wrong answers

Option A is wrong because fine-tuning bakes knowledge into model weights, requires retraining whenever documents change, and is expensive and slow — it does not satisfy the 'latest documents without fine-tuning' requirement. Option B is wrong because a pre-built model with no configuration has no access to the internal knowledge base and will hallucinate or refuse to answer domain-specific questions. Option C is wrong because while BigQuery can be a data source, the question specifies documents already stored in Cloud Storage; adding a BigQuery hop is unnecessary and does not by itself ground the model unless wired through a Vertex AI Search data store.

844
MCQeasy

What is the main advantage of using a model with a larger context window?

A.Better performance on image generation tasks
B.Lower cost per API call
C.Ability to process longer documents or conversations without truncation
D.Faster inference speed
AnswerC

A larger context window raises the number of tokens the model can attend to in one pass, so entire long documents or extended conversations fit without truncation. That directly satisfies the stem's constraint of processing longer input without losing earlier content.

Why this answer

A larger context window allows the model to process and retain more tokens in a single pass, enabling it to handle longer documents, extended conversations, or large codebases without needing to truncate or chunk the input. This is critical for tasks like summarizing entire research papers, maintaining coherent multi-turn dialogues, or analyzing long legal contracts, where preserving full context directly impacts output quality and accuracy.

Exam trap

Google often tests the misconception that 'bigger is always better' by pairing a clear benefit (longer context) with attractive but unrelated options like lower cost or faster speed, tempting candidates to conflate model capability with operational efficiency.

How to eliminate wrong answers

Option A is wrong because context window size is a token-length constraint for text and multimodal inputs, not a direct factor in image generation quality; image generation performance depends on model architecture, training data, and diffusion/transformer design, not context length. Option B is wrong because larger context windows typically increase computational cost per API call due to the quadratic scaling of attention mechanisms (e.g., O(n²) in standard transformers), leading to higher latency and cost, not lower. Option D is wrong because processing more tokens in a larger context window generally increases inference time, as the model must attend to a longer sequence; faster inference is achieved through model optimization, quantization, or pruning, not by expanding context length.

845
Multi-Selecthard

A health-tech startup is fine-tuning a generative AI model on electronic health records (EHR) to assist in clinical decision support. They need to ensure responsible AI practices. Which THREE measures should they implement? (Select three.)

Select 3 answers
A.Automate all decisions to reduce human error
B.Evaluate model outputs for bias across different demographic groups
C.Require a human clinician to review all AI-generated recommendations before action
D.Publish a Model Card that describes the model's intended use, limitations, and performance
E.Train the model exclusively on data from a single hospital to ensure consistency
AnswersB, C, D

Bias evaluation across demographic groups detects disparate performance in the fine-tuned model, a core responsible AI control for clinical decision support. Because EHR data encodes historical inequities, measuring output disparities by subgroup satisfies the requirement to identify and mitigate harm before deployment.

Why this answer

Option B is correct because systematically evaluating model outputs for bias across demographic groups (e.g., using disaggregated performance metrics like sensitivity, specificity, and equal opportunity difference by age, sex, race, and ethnicity) is essential in clinical AI to detect and mitigate health disparities before deployment. Option C is correct because requiring a human clinician to review all AI-generated recommendations preserves meaningful human oversight, keeping the clinician accountable for the final decision and preventing automation bias or unsafe autonomous actions in clinical decision support. Option D is correct because a Model Card documents the model's intended use, limitations, performance metrics, and ethical considerations, providing the transparency and traceability needed for responsible AI governance and regulatory review.

Option A is not appropriate because fully automating all decisions removes human oversight and amplifies automation bias and accountability gaps, which is contrary to responsible AI in healthcare. Option E is not appropriate because training exclusively on data from a single hospital produces narrow, non-generalizable models and can encode site-specific biases, undermining fairness and robustness across diverse patient populations.

846
MCQeasy

A developer needs to generate embeddings for text data to be used in a semantic search application. Which Google Cloud service should they use?

A.Document AI
B.Cloud Translation API
C.Cloud Speech-to-Text
D.Vertex AI Embeddings API
AnswerD

Vertex AI Embeddings API generates dense vector representations from text, which semantic search requires for similarity comparison. It provides managed embedding generation without building custom infrastructure, matching the developer's need for text embeddings on Google Cloud.

Why this answer

Vertex AI Embeddings API is the correct choice because it provides a managed service to generate vector embeddings from text data, which are essential for semantic search applications that rely on understanding meaning rather than exact keyword matches. This API leverages large language models to convert text into high-dimensional vectors, enabling efficient similarity search using vector databases or nearest neighbor algorithms.

Exam trap

The trap here is that candidates may confuse Document AI's ability to extract text from documents with the need to generate embeddings from that text, overlooking that embedding generation is a separate, specialized step required for semantic search.

How to eliminate wrong answers

Option A is wrong because Document AI is designed for document processing tasks like OCR, parsing, and extraction of structured data from documents, not for generating text embeddings. Option B is wrong because Cloud Translation API is used for translating text between languages, not for creating vector representations of text for semantic search. Option C is wrong because Cloud Speech-to-Text converts audio to text, but does not generate embeddings or support semantic search directly.

847
Multi-Selecteasy

Which THREE strategies should be combined to effectively reduce biased outputs in a generative AI model? (Choose three.)

Select 3 answers
A.Implement safety filters targeting hate speech and stereotypes.
B.Conduct human evaluation and feedback loops.
C.Use diverse few-shot examples that represent different demographics.
D.Raise the temperature to increase output variability.
E.Fine-tune the model on a biased dataset to learn patterns.
AnswersA, B, C

Safety filters block explicitly biased content.

Why this answer

Implementing safety filters targeting hate speech and stereotypes directly blocks the generation of biased or harmful content at the output layer. These filters use predefined rule sets or trained classifiers to detect and suppress language that reflects demographic or cultural biases, reducing the risk of the model producing offensive or stereotypical responses.

Exam trap

Google often tests the misconception that increasing randomness (temperature) or training on biased data can somehow reduce bias, when in fact both actions worsen the problem by either amplifying noise or embedding the bias deeper into the model's weights.

848
MCQeasy

A large e-commerce company is experiencing high costs for their generative AI product recommendation system. The system generates personalized product descriptions for millions of users daily. The team wants to reduce cost while maintaining quality. They are using a fine-tuned version of a large foundation model hosted on Vertex AI. The current cost is driven by the number of tokens processed. Which approach should they take?

A.Optimize prompts to generate shorter, more concise descriptions
B.Switch to a larger, more capable foundation model
C.Retrain the model with more product data to improve efficiency
D.Increase the batch size of inference requests
AnswerA

Shorter outputs use fewer tokens, reducing cost.

Why this answer

Prompt engineering to reduce output length decreases token usage per request, directly lowering cost without model changes. Option B (switching to a larger model) increases cost. Option C (increasing batch size) may not reduce per-request cost.

Option D (retraining with more data) does not affect inference cost.

849
Multi-Selecthard

A company is using a large language model for automated translation of legal contracts. They find that the translations sometimes alter the meaning of specific clauses. Which TWO approaches would most effectively preserve the original meaning? (Choose two.)

Select 2 answers
A.Provide the full contract context in a single prompt.
B.Set top-p=0.1 to limit the vocabulary to the most likely tokens.
C.Fine-tune the model on a parallel corpus of legal translations.
D.Use a glossary of key legal terms with their translations.
E.Increase the temperature to allow more creative phrasing.
AnswersC, D

Fine-tuning on domain-specific translations improves accuracy.

Why this answer

Fine-tuning on a parallel corpus of legal translations adapts the model's weights to the specific legal domain, improving fidelity for legal terminology and clause structures by learning from ground-truth translations. In addition, providing a glossary of key legal terms with their approved translations constrains the model to use consistent, correct terminology, reducing meaning-altering substitutions. Together, these two approaches most effectively preserve the original meaning.

Prompt engineering alone (Option A) and parameter tweaks such as top-p or higher temperature (Options B and E) do not address the root cause of semantic drift.

Exam trap

A common misconception is that prompt engineering alone (Option A) or parameter tweaks like top-p and temperature (Options B and E) can substitute for domain-specific fine-tuning or explicit glossary control, when in fact these methods do not address the root cause of semantic drift in specialized translations.

850
MCQeasy

A marketing team uses Gemini in Vertex AI to generate campaign taglines. The first drafts are generic and closely mirror the prompt wording. They want more distinctive, varied taglines from the same model without changing the model itself. Which parameter should they adjust?

A.Increase the maximum output tokens to allow longer taglines.
B.Lower the temperature to make the output more deterministic.
C.Raise the temperature to increase randomness and diversity in token selection.
D.Enable a higher top-p value to consider more cumulative probability mass.
AnswerC

Temperature scales the probability distribution over next tokens; raising it flattens that distribution so less likely words become more probable. That yields more varied and less generic phrasing, which fits the goal of distinctive taglines. The model is unchanged, so this is a low-effort tuning knob. Very high values risk incoherence, so moderate increases are appropriate.

Why this answer

Temperature is the direct control over how sharply the model favors likely tokens. Raising it flattens the distribution and encourages less predictable word choices, which produces more distinctive taglines from the same model. Lowering temperature increases repetition, output token limits only change length, and top-p is a related but secondary truncation control rather than the primary diversity knob.

Exam trap

The trap here is confusing controls that change response length or sampling cutoffs with the primary diversity control, and mistakenly lowering temperature when the goal is more varied output.

851
MCQmedium

A developer is building a real-time speech transcription application for customer support calls. The audio is streamed, and the transcription must be returned with low latency. Which Google Cloud AI service should they use?

A.Natural Language API
B.Cloud Speech-to-Text with streaming recognition
C.Cloud Text-to-Speech
D.Vertex AI with a custom model
AnswerB

Streaming recognition accepts an ongoing audio stream and returns interim and final transcripts as speech arrives, rather than requiring a complete file upload. That incremental processing is what meets the low-latency requirement for live customer support calls.

Why this answer

Cloud Speech-to-Text supports streaming recognition via gRPC or the streaming REST endpoint, returning partial transcripts as audio arrives, which is exactly what low-latency real-time transcription requires. It handles telephony audio with models tuned for phone calls and supports interim results for responsiveness.

Exam trap

Generative AI Leader often tests service-selection confusion — the trap is picking Natural Language API thinking it handles audio, or Text-to-Speech by reversing the direction, when only Speech-to-Text with streaming mode provides real-time transcription.

How to eliminate wrong answers

Option A is wrong because Natural Language API performs text analysis (sentiment, entities, syntax) on already-transcribed text, not audio transcription. Option C is wrong because Text-to-Speech does the reverse — it synthesizes speech from text, not transcribes audio. Option D is wrong because building a custom Vertex AI model is unnecessary and slower to deploy when a managed streaming STT service already meets the requirement.

852
Multi-Selecteasy

A company is choosing a generative AI model for code generation. Which TWO considerations are most important?

Select 2 answers
A.The total number of model parameters
B.Whether the model's training data includes the target programming languages
C.The open-source license of the model
D.The maximum context length supported by the model
E.The latency of the model's inference endpoint
AnswersB, D

Code generation depends on the model having seen the target programming languages during training; without that coverage, syntax, idioms and library usage are unreliable, so training-data language coverage directly determines output quality for the required languages.

Why this answer

Option B is correct because a code-generation model must have been trained on the target programming languages (e.g., Python, Java, C++) to produce syntactically valid and idiomatic code; a model lacking exposure to a language will generate unreliable or non-compiling output. Option D is correct because maximum context length determines how much source code, documentation, and surrounding files the model can ingest at once, which is critical for tasks like multi-file refactoring, repository-level completion, and understanding large codebases. Option A is not decisive because parameter count alone does not guarantee code quality; a smaller model fine-tuned on code can outperform a larger general-purpose model.

Option C is not a primary technical consideration for code-generation capability, though licensing matters for legal/commercial deployment rather than model effectiveness. Option E affects user experience and cost but is a deployment concern, not a core capability consideration for choosing a code-generation model.

Exam trap

The trap here is that candidates often assume more parameters (A) or lower latency (E) are always better, but Google tests the understanding that domain-specific training data relevance (B) and context length (D) are critical for code generation accuracy and handling long code sequences.

853
MCQmedium

A company needs to extract structured data from scanned invoices (invoice number, date, total amount) using a pre-built AI solution. Which Google Cloud service is MOST appropriate?

A.Natural Language AI
B.Translation API
C.Document AI
D.Vision AI
AnswerC

Document AI offers pre-built invoice parsers that extract structured fields such as invoice number, date and total amount from scanned documents, requiring no custom model training. Its specialised document understanding satisfies the stem's demand for a pre-built solution targeting scanned invoices.

Why this answer

Document AI is Google Cloud's purpose-built service for extracting structured data from documents, including scanned invoices, receipts, and forms. It offers pre-built processors like Invoice Parser that use OCR plus specialized ML models to extract fields such as invoice number, date, and total amount with high accuracy, without requiring custom model training.

Exam trap

Generative AI Leader often tests the confusion between general vision/OCR services (Vision AI) and document-specific extraction (Document AI) — candidates must recognize that structured field extraction from invoices requires Document AI's pre-built Invoice Parser, not generic Vision AI.

How to eliminate wrong answers

Option A is wrong because Natural Language AI is designed for text analysis tasks like sentiment analysis, entity extraction, and syntax parsing — it does not parse document layouts or extract structured fields from scanned invoices. Option B is wrong because Translation API translates text between languages and has no document parsing or field extraction capability. Option D is wrong because Vision AI performs general image analysis (object detection, OCR, label detection) but lacks invoice-specific field extraction; it would require custom model development to achieve what Document AI's Invoice Parser does out of the box.

854
MCQmedium

A healthcare startup is developing an AI system to assist radiologists in detecting tumors from X-ray images. Which Google AI Principle is MOST directly applicable to this use case?

A.Incorporate privacy design principles
B.Be built and tested for safety
C.Be socially beneficial
D.Avoid creating or reinforcing unfair bias
AnswerB

Building and testing for safety directly addresses the clinical risk of misdiagnosis in tumour detection, requiring rigorous validation before deployment. This principle governs the model's performance and failure modes, satisfying the constraint that patient harm from inaccurate X-ray interpretation must be prevented through testing.

Why this answer

The principle 'be built and tested for safety' directly applies to medical applications where incorrect detection could harm patients. The other principles are also relevant but safety is the most directly applicable to a diagnostic tool.

855
MCQeasy

A mid-sized insurance company wants to adopt generative AI to help claims adjusters draft customer correspondence. Executives are unsure how to prioritize the first use case and want a framework that balances business value with implementation feasibility. Which first step best aligns with a value-driven generative AI adoption strategy?

A.Wait until competitors publicly report their generative AI results before committing any budget.
B.Begin with the most technically complex use case to prove the organization's engineering capability.
C.Deploy generative AI across all claims functions simultaneously to maximize transformation speed.
D.Select the use case with the highest volume of manual effort and a clear measurable baseline for time saved.
AnswerD

Prioritizing a high-volume, effort-intensive task with an existing measurable baseline lets the organization demonstrate tangible return quickly and build momentum. Drafting correspondence is repetitive, so time saved per adjuster is easy to quantify. This aligns with a value-driven approach that pairs business impact with feasibility rather than chasing the most technically ambitious option first.

Why this answer

A value-driven generative AI strategy starts where effort is high, the task is repetitive, and a baseline already exists to measure improvement. High-volume correspondence drafting offers fast, quantifiable time savings that build credibility for later, more complex initiatives. This sequencing balances impact with feasibility instead of optimizing for technical novelty or waiting on the sidelines.

Exam trap

The trap here is equating ambition with strategy, assuming the most complex or broadest rollout is the best first move when measurable quick wins build the case for scaling.

856
MCQmedium

During model evaluation, a team observes good performance on training data but poor on validation data. Which regularization technique is most appropriate to address this?

A.Add more training data
B.Increase the learning rate
C.Apply dropout
D.Use a larger batch size
AnswerC

Dropout randomly deactivates units during training, forcing the network to learn redundant, distributed representations instead of memorising training samples. This regularisation narrows the gap between training and validation performance, directly addressing the poor generalisation the team observed during model evaluation.

Why this answer

The scenario describes overfitting, where the model memorizes training data but fails to generalize to unseen validation data. Dropout is a regularization technique that randomly deactivates a fraction of neurons during training, forcing the network to learn more robust features and reducing co-adaptation, which directly mitigates overfitting.

Exam trap

Google Cloud often tests the distinction between techniques that improve generalization (regularization) versus those that improve optimization (learning rate, batch size), leading candidates to confuse data augmentation or hyperparameter tuning with regularization methods like dropout.

How to eliminate wrong answers

Option A is wrong because adding more training data can help reduce overfitting but is not a regularization technique; it addresses data scarcity, not the core issue of model complexity. Option B is wrong because increasing the learning rate can cause training instability, divergence, or overshooting of the loss minimum, and does not prevent overfitting. Option D is wrong because using a larger batch size often leads to sharper minima and poorer generalization, potentially worsening overfitting, and is not a regularization method.

857
MCQhard

A gaming company is using Vertex AI Imagen to create concept art. They have a stable pipeline that generates images based on text prompts. Recently, they introduced a new feature: using a reference image to guide the style (image-to-image generation). However, when using a reference image, the generated images often have unnatural color shifts and artifacts. The team suspects that the reference image is being resized to a resolution that the model wasn't trained on. They are using the default Imagen settings. What is the most likely cause and the best solution?

A.Increase the number of inference steps to improve detail.
B.The reference image is being resized to a non-standard aspect ratio; preprocess the image to the recommended resolution and aspect ratio.
C.Reduce the style weight in the image-to-image prompt.
D.Switch to a different image generation model.
AnswerB

Imagen's default preprocessing resizes reference images, and mismatched aspect ratios distort content, producing colour shifts and artifacts. Preprocessing to the recommended resolution and aspect ratio preserves the model's expected input geometry, eliminating the distortion at its source.

Why this answer

The default Imagen settings expect input images at specific resolutions (e.g., 256x256, 512x512, or 1024x1024) and a 1:1 aspect ratio. When a reference image is resized to a non-standard resolution or aspect ratio, the model's internal processing can introduce artifacts and unnatural color shifts due to misalignment with its training distribution. Preprocessing the image to the recommended resolution and aspect ratio ensures the model operates within its optimal input space, eliminating these issues.

Exam trap

The trap here is that candidates may confuse image quality issues with model hyperparameters (like inference steps or style weight) rather than recognizing that the fundamental input preprocessing—specifically resolution and aspect ratio—is the most common cause of artifacts in image-to-image generation with Imagen.

How to eliminate wrong answers

Option A is wrong because increasing inference steps primarily refines noise reduction and detail, but it does not address the root cause of resolution mismatch; it may even amplify artifacts from a poorly resized input. Option C is wrong because reducing style weight controls how strongly the reference image influences the output, but it does not fix the fundamental problem of the reference image being at a non-standard resolution or aspect ratio, which causes artifacts regardless of style influence. Option D is wrong because switching to a different model like Stable Diffusion would not resolve the resolution preprocessing issue; the team would still need to properly resize the reference image for that model, and the problem is with their pipeline, not the model itself.

858
MCQeasy

A marketing company wants to fine-tune a generative AI model to adopt a specific brand voice. Which tuning method is most appropriate?

A.RLHF with general user feedback
B.Grounding with external knowledge base
C.Supervised fine-tuning with labeled examples of the brand voice
D.Prompt engineering with system instructions
AnswerC

Supervised fine-tuning trains the model on labelled input-output pairs, directly shaping its outputs to match the desired brand voice. This satisfies the requirement to adopt a specific style, which prompt engineering alone cannot reliably enforce across all generations.

Why this answer

Supervised fine-tuning (SFT) is the most appropriate method because it directly trains the model on a curated dataset of input-output pairs that exemplify the desired brand voice. By adjusting the model's weights through backpropagation on labeled examples, the model learns to mimic the specific tone, vocabulary, and stylistic patterns of the brand, making it the most precise approach for adopting a fixed voice.

Exam trap

A common misconception is that prompt engineering (Option D) is sufficient for fine-grained style control, when in reality it only provides a weak, non-parametric signal that cannot reliably enforce a consistent brand voice across varied contexts.

How to eliminate wrong answers

Option A is wrong because RLHF with general user feedback optimizes for broad human preferences (e.g., helpfulness, harmlessness) rather than a specific, consistent brand voice; it introduces variance from diverse user ratings that can dilute the target style. Option B is wrong because grounding with an external knowledge base retrieves factual information (e.g., via RAG) to reduce hallucination but does not alter the model's generation style or tone; it cannot teach the model to adopt a brand voice. Option D is wrong because prompt engineering with system instructions provides a static, high-level directive that the model may follow inconsistently, especially for nuanced stylistic constraints; it does not update model weights and fails to embed the brand voice deeply into the model's behavior across diverse prompts.

859
MCQhard

A large insurance company is using generative AI to automate claims processing. They have deployed a custom fine-tuned model on Vertex AI that reads claim documents and extracts key information. Recently, they noticed that the model’s performance degrades over time for certain claim types, leading to incorrect payouts. The team needs to detect and address model drift with minimal manual intervention. They have a data pipeline that captures incoming claims and user feedback on predictions. Which approach should they take?

A.Implement a human review process for all claims the model processes
B.Set up continuous evaluation with automated retraining pipelines based on performance metrics
C.Switch to a simpler rule-based system to avoid drift
D.Manually retrain the model monthly using a snapshot of recent claims
AnswerB

Continuous evaluation monitors live performance metrics against baselines, and automated retraining pipelines refresh the model when drift is detected, correcting degradation with minimal manual intervention. This directly satisfies the requirement to detect and address drift automatically.

Why this answer

It establishes a closed-loop MLOps pipeline where continuous evaluation of performance metrics (e.g., precision, recall, or F1-score on streaming data) triggers automated retraining when drift is detected. This minimizes manual intervention while ensuring the model adapts to distribution shifts in claim types, which is critical for maintaining accurate payouts in production.

Exam trap

Google Cloud often tests the misconception that periodic manual retraining (Option D) is sufficient, but the trap here is that it ignores the need for real-time drift detection and automated response, which is essential for production systems handling high-stakes financial decisions.

How to eliminate wrong answers

Option A is wrong because implementing human review for all claims defeats the purpose of automation and introduces significant operational cost and latency, failing the requirement for minimal manual intervention. Option C is wrong because switching to a simpler rule-based system cannot handle the complexity and variability of claim documents, and it will still suffer from drift as claim patterns evolve over time. Option D is wrong because manually retraining monthly on a snapshot ignores real-time drift detection and may miss sudden shifts between retraining cycles, leading to prolonged periods of degraded performance.

860
Multi-Selecthard

A company is deploying a generative AI chatbot for customer support. They want to ensure that the chatbot does not generate harmful content and that they can customize the safety thresholds. Which TWO features in Vertex AI should they use? (Select 2)

Select 2 answers
A.AutoML Tables
B.Custom safety settings (adjustable thresholds)
C.Model Cards
D.Safety filters
E.People + AI Guidebook
AnswersB, D

Custom safety settings let you tune threshold levels per harm category — hate speech, harassment, sexually explicit, dangerous content — directly satisfying the requirement for adjustable safety thresholds. Configuring these on the Vertex AI model blocks responses exceeding your chosen sensitivity, preventing harmful chatbot output while retaining granular control over filtering strictness.

Why this answer

Option B (Custom safety settings with adjustable thresholds) is correct because Vertex AI's generative AI safety controls let you configure per-category harm thresholds (for example, blocking more or less aggressively for hate speech, harassment, sexually explicit, and dangerous content), which directly satisfies the requirement to customize safety thresholds. Option D (Safety filters) is correct because these filters are the mechanism that actually evaluates model prompts and responses against harm categories and blocks or allows content, preventing the chatbot from generating harmful output. Together, safety filters provide the enforcement while custom safety settings provide the tunable thresholds the company wants.

Option A (AutoML Tables) is incorrect because it is a tabular ML training/prediction service, not a content-safety feature. Option C (Model Cards) is incorrect because it documents model details, intended use, and evaluations but does not filter or block harmful content. Option E (People + AI Guidebook) is incorrect because it is a design guidance resource, not a runtime safety control in Vertex AI.

861
MCQhard

A media company uses a Gemini model to produce short news digests from long articles. Editors complain that digests vary in length and sometimes bury the key fact in the middle. The team wants a repeatable, machine-checkable structure for every digest. Which technique should they use?

A.Provide a response schema with defined fields such as headline, key fact, and supporting details, and request JSON output.
B.Fine-tune the model on a corpus of previously published digests written by editors.
C.Set the temperature to 0 and the topK to 1 to make output fully deterministic.
D.Ask the model to be concise and to put the most important information first in the prompt.
AnswerA

A response schema with controlled generation constrains the model to emit fields in a defined structure, making length and placement of the key fact machine-checkable. The model fills required fields rather than free-forming a paragraph, so editors get consistent digests and downstream systems can validate the JSON programmatically.

Why this answer

A response schema with JSON output turns the digest into a structured object with named fields, so the key fact has a fixed location and length is bounded by the schema. This is directly machine-checkable and repeatable, unlike prose instructions, sampling settings, or style-oriented fine-tuning.

Exam trap

The trap here is treating temperature 0 as a guarantee of structure, when determinism only makes the same input reproducible and says nothing about output shape.

862
MCQeasy

A marketing team wants to generate social media posts with Google Cloud's generative AI, but they have no machine learning experience. They need a no-code interface to experiment with prompts and tune settings like temperature. Which Google Cloud service should they use?

A.Document AI
B.BigQuery ML
C.Dialogflow CX
D.Vertex AI Studio
AnswerD

Vertex AI Studio provides a visual, no-code console where users can design prompts, choose models like Gemini, adjust parameters such as temperature and top-k, and immediately view generated text. It is designed for rapid experimentation without writing code, making it ideal for the marketing team's need to prototype social media content.

Why this answer

Vertex AI Studio is the correct choice because it offers a no-code environment specifically for generative AI experimentation. It allows users to write prompts, select foundation models, and adjust generation parameters, then see results instantly. This matches the marketing team's need to create social media posts without ML expertise.

Exam trap

The trap here is confusing Vertex AI Studio with other Google Cloud AI services that are not designed for no-code generative text experimentation.

863
MCQmedium

A developer is using Gemini 1.5 Pro and needs to process a 2-hour video to answer questions about its content. The video is stored in Cloud Storage. What is the most efficient approach?

A.Extract frames using Video Intelligence API and then send them as images to Gemini
B.Transcribe the video using Chirp, then analyze the text with Gemini
C.Use a custom model fine-tuned on video understanding tasks
D.Send the video file as part of the prompt to Gemini 1.5 Pro
AnswerD

Gemini 1.5 Pro's long context window accepts video directly, so passing the Cloud Storage file in the prompt lets the model reason over the full two-hour content without building a separate frame-extraction or transcription pipeline.

Why this answer

Gemini 1.5 Pro supports video input natively; you can pass the video directly (via GCS URI) and ask questions. Transcribing first adds latency and loses visual context.

864
MCQeasy

A small marketing agency wants to add an AI assistant that summarizes campaign briefs and drafts social posts. The developers have limited cloud experience and prefer a fully managed, serverless way to call Google's Gemini models without provisioning infrastructure. Which approach best meets this requirement?

A.Use the Gemini API through Google AI Studio for rapid prototyping and API-key access
B.Deploy an open model on a Compute Engine VM with an attached GPU
C.Create a Vertex AI custom training job for a bespoke summarization model
D.Build a Cloud Run service that hosts a self-managed transformer container
AnswerA

Google AI Studio and the Gemini API provide a fully managed, serverless way to call Gemini models using an API key, with no infrastructure to provision. This suits a small team with limited cloud experience that wants quick access for summarization and drafting. It is the lowest-friction entry point for experimenting with and shipping lightweight Gemini-powered features.

Why this answer

When the goal is simply to consume Gemini capabilities with minimal setup, the managed Gemini API accessed through Google AI Studio removes infrastructure concerns entirely. It supports summarization and drafting through straightforward API calls, letting a small team focus on prompts and application logic rather than serving stacks. Self-hosted or custom-trained alternatives introduce operational or data-science overhead the scenario does not require.

Exam trap

The trap here is assuming that any Google Cloud compute service, such as Cloud Run, is automatically the serverless answer for calling Gemini, when the managed Gemini API already removes the need to host anything.

865
MCQmedium

A media company uses generative AI to produce personalized news summaries for subscribers. They notice that the summaries sometimes contain factual inaccuracies, leading to customer complaints. The team needs to improve accuracy without slowing down the generation speed. They are using a pre-trained model via Vertex AI. What strategy should they implement?

A.Switch to a larger, more accurate foundation model
B.Fine-tune the model on a dataset of verified news articles
C.Implement retrieval-augmented generation (RAG) with a trusted knowledge base
D.Add a human-in-the-loop review for every summary
AnswerC

RAG grounds generation in a trusted knowledge base by retrieving relevant verified content and injecting it into the prompt, so summaries reflect accurate sources. This improves factual accuracy without retraining or slowing inference, unlike fine-tuning or larger models.

Why this answer

Retrieval-augmented generation (RAG) grounds the model's output in a trusted, external knowledge base, allowing it to retrieve verified facts in real time without retraining. This directly addresses factual inaccuracies while maintaining generation speed, as the pre-trained model remains unchanged and only the retrieval step is added. RAG avoids the latency of human review and the computational cost of fine-tuning or switching models.

Exam trap

Google Cloud often tests the misconception that fine-tuning is the default solution for accuracy issues, but the trap here is that RAG provides a faster, more scalable way to ground outputs in verified data without retraining, which is critical when speed and accuracy must both be maintained.

How to eliminate wrong answers

Option A is wrong because switching to a larger foundation model would increase inference latency and computational cost, contradicting the requirement to not slow down generation speed, and it does not guarantee improved factual accuracy without additional grounding. Option B is wrong because fine-tuning on a dataset of verified news articles requires significant time, data, and compute resources, and it may not prevent hallucinations on unseen topics, while also risking catastrophic forgetting of the model's general capabilities. Option D is wrong because adding a human-in-the-loop review for every summary introduces unacceptable latency and operational overhead, making it impractical for real-time personalized news generation at scale.

866
MCQhard

A company is comparing Google Cloud Vertex AI vs AWS Bedrock vs Azure OpenAI. Their application requires grounding responses with real-time search results from the internet. Which platform's feature uniquely supports this requirement?

A.Gemini API without grounding
B.Vertex AI with Grounding + Google Search
C.Azure OpenAI with Bing Search grounding
D.AWS Bedrock with Knowledge Bases
AnswerB

Vertex AI uniquely provides Grounding with Google Search, allowing real-time web results to be incorporated into model responses.

Why this answer

Vertex AI with Grounding + Google Search is the correct answer because it uniquely provides native, real-time grounding against live internet search results via Google Search, enabling the model to retrieve and cite up-to-date information from the web. This feature is directly integrated into Vertex AI's model serving pipeline, allowing responses to be grounded in current, publicly available data without requiring external API calls or custom retrieval logic.

Exam trap

A common mistake is confusing the distinction between grounding with static knowledge bases (like AWS Bedrock Knowledge Bases) and dynamic, real-time internet search grounding, leading candidates to mistakenly select Azure OpenAI with Bing Search grounding because it also uses a search engine, but the question specifically asks for the platform's unique feature—and Vertex AI's native integration with Google Search is the only one that is built directly into the model serving platform without requiring a separate search service subscription or API key.

How to eliminate wrong answers

Option A is wrong because the Gemini API without grounding does not support any form of internet search grounding; it relies solely on the model's pre-trained knowledge, which is static and cannot incorporate real-time search results. Option C is wrong because Azure OpenAI with Bing Search grounding, while capable of grounding with Bing, is not a unique feature of the Azure platform in this context—the question asks for the platform whose feature uniquely supports the requirement, and Vertex AI's Grounding + Google Search is the only one that is natively built into the model serving infrastructure without requiring a separate search service integration. Option D is wrong because AWS Bedrock with Knowledge Bases is designed for grounding against private, static data sources (e.g., documents, databases) and does not support real-time internet search grounding; it lacks the ability to dynamically query live web content.

867
MCQeasy

What is the key advantage of using adapter-based fine-tuning methods like LoRA compared to full fine-tuning of a large language model?

A.LoRA significantly reduces the number of trainable parameters, making fine-tuning more memory-efficient
B.LoRA is faster at inference time compared to the fully fine-tuned model
C.LoRA eliminates the need for a base model
D.LoRA enables training on a larger dataset than full fine-tuning
AnswerA

LoRA freezes the base weights and injects small trainable low-rank matrices into attention layers, cutting trainable parameters by orders of magnitude. This slashes optimiser and gradient memory, so fine-tuning fits on far cheaper GPUs than full fine-tuning, which updates every weight.

Why this answer

LoRA (Low-Rank Adaptation) injects trainable low-rank matrices into the transformer layers while keeping the original model weights frozen. This drastically reduces the number of trainable parameters (often by 10,000x), which lowers GPU memory requirements for storing optimizer states and gradients during training, making fine-tuning feasible on consumer hardware.

Exam trap

Google often tests the misconception that parameter-efficient methods like LoRA improve inference speed, when in reality they primarily reduce memory during training and do not accelerate inference.

How to eliminate wrong answers

Option B is wrong because LoRA does not change the inference path; the adapter weights are merged into the base model or applied as a separate forward pass, so inference speed is comparable to or slightly slower than a fully fine-tuned model, not faster. Option C is wrong because LoRA is an additive method that requires the original base model to remain frozen; it cannot function without the base model. Option D is wrong because LoRA does not inherently allow training on a larger dataset; the dataset size is independent of the fine-tuning method, and full fine-tuning can also use large datasets if sufficient memory is available.

868
MCQhard

A large enterprise wants to deploy multiple generative AI models across different business units while ensuring cost governance and usage tracking. Which Google Cloud solution is best suited?

A.Use Vertex AI Endpoint with monitoring
B.Deploy each model in a separate project with IAM policies
C.Implement a custom cost allocation using labels
D.Use Cloud Billing budgets and alerts per model
AnswerB

Projects isolate resources and costs per business unit.

Why this answer

Deploying each model in a separate Google Cloud project with IAM policies provides the strongest isolation for cost governance and usage tracking. This approach ensures that each business unit's model usage is billed to its own project, enabling granular cost allocation and independent monitoring without cross-project interference. It also allows per-project budget alerts and usage quotas, directly addressing the enterprise's need for decentralized cost control.

Exam trap

The trap here is that candidates often confuse cost allocation mechanisms (like labels or budgets) with true cost isolation, assuming that tagging or alerting alone can enforce per-business-unit governance without the structural separation that separate projects provide.

How to eliminate wrong answers

Option A is wrong because Vertex AI Endpoint with monitoring tracks model performance and latency but does not inherently isolate costs per business unit; it aggregates usage under a single project, making per-unit cost governance difficult. Option C is wrong because custom cost allocation using labels requires manual tagging and can be inconsistent or incomplete, leading to inaccurate cost tracking; labels are metadata, not a billing boundary. Option D is wrong because Cloud Billing budgets and alerts per model are not natively supported; budgets apply at the project or billing account level, not per model, and cannot enforce cost isolation across multiple business units.

869
MCQmedium

A developer is using the Gemini API to classify customer emails. They want to ensure that the model always returns one of three predefined labels: 'complaint', 'inquiry', or 'feedback'. Which model configuration is MOST appropriate?

A.Set temperature to 1.0 and top-p to 0.9 to allow creativity while constraining via system instructions
B.Fine-tune the model on a dataset of labeled emails to memorize the three classes
C.Use top-k sampling with k=50 and no temperature adjustment
D.Set temperature to 0.0 and use few-shot examples with required labels in the prompt
AnswerD

Setting temperature to 0.0 makes decoding greedy, so the highest-probability token is always chosen, maximising consistency. Few-shot examples in the prompt demonstrate the exact label set, steering the model to emit only 'complaint', 'inquiry' or 'feedback'. Together they satisfy the constraint of returning one of three predefined labels.

Why this answer

Setting temperature to 0.0 makes the model deterministic, minimizing randomness and ensuring consistent output. Combined with few-shot examples that explicitly list the three required labels ('complaint', 'inquiry', 'feedback') in the prompt, this configuration reliably constrains the model to return only those labels, which is the most appropriate approach for a strict classification task.

Exam trap

A common pitfall in the Google Generative AI exam is believing that higher creativity settings (temperature, top-p) are needed for classification tasks, when in fact deterministic settings (temperature 0.0) combined with prompt engineering (few-shot with explicit labels) are the correct approach for strict label constraints.

How to eliminate wrong answers

Option A is wrong because temperature 1.0 and top-p 0.9 maximize randomness and creativity, which is counterproductive for a deterministic classification task where the model must output only three fixed labels. Option B is wrong because fine-tuning on labeled emails would teach the model to generate the labels, but it does not guarantee the model will never output other tokens; fine-tuning is overkill and less reliable than prompt engineering for such a simple constraint. Option C is wrong because top-k sampling with k=50 still introduces randomness and does not force the model to choose only from the three predefined labels; it only limits the pool of candidate tokens to the top 50, which may still include irrelevant tokens.

870
MCQmedium

A product manager wants to generate meeting summaries automatically using Gemini for Google Workspace. They need summaries to be sent to all participants immediately after the meeting ends. Which Gemini feature should they use?

A.Gemini in Google Docs - Help me write
B.Gemini in Gmail - Smart Compose
C.Vertex AI Agent Builder with a meeting transcription model
D.Gemini in Google Meet - Take notes and summaries
AnswerD

Gemini in Google Meet takes notes and generates summaries tied to the meeting lifecycle, distributing them to participants once the call concludes. This directly satisfies the requirement for immediate post-meeting delivery, which Docs or Gmail-based prompting cannot automate.

Why this answer

Gemini in Google Meet can automatically generate meeting summaries and share them after the meeting. The other options are for different Workspace apps or manual use.

871
Multi-Selectmedium

A data scientist wants to train a model using BigQuery ML. Which two statements are true about BigQuery ML? (Choose two.)

Select 2 answers
A.It requires a separate Vertex AI training cluster
B.It supports only linear regression models
C.Data must be exported to Cloud Storage before training
D.Models can be trained using SQL directly on BigQuery data
E.It supports both supervised and unsupervised learning
AnswersD, E

BigQuery ML lets you create and train models using SQL queries directly against data already stored in BigQuery, avoiding data export or separate ML tooling. This satisfies the stem's requirement to train within BigQuery using familiar SQL syntax.

Why this answer

Option D is correct because BigQuery ML's core design lets you create and train models directly on data stored in BigQuery using standard SQL statements such as CREATE MODEL, without moving or exporting the data. Option E is correct because BigQuery ML supports supervised learning (for example, linear regression, logistic regression, and boosted trees) as well as unsupervised learning (for example, k-means clustering and matrix factorization). Option A is wrong because BigQuery ML performs training inside BigQuery itself and does not require a separate Vertex AI training cluster.

Option B is wrong because BigQuery ML supports many model types beyond linear regression, including logistic regression, k-means, and deep neural networks. Option C is wrong because training happens directly on BigQuery tables, so exporting data to Cloud Storage first is unnecessary.

Exam trap

Generative AI Leader often tests the misconception that BigQuery ML requires data export or external training infrastructure — the exam expects recognition that BigQuery ML trains models in-database using SQL and supports both supervised and unsupervised learning.

872
MCQhard

An enterprise wants to use Gemini 1.5 Flash for a real-time chat application with low latency. Which trade-off should they expect compared to Gemini 1.5 Pro?

A.Higher quality and lower latency
B.Lower latency but potentially lower quality
C.Lower cost but longer context window
D.Higher accuracy but slower responses
AnswerB

Flash uses a lighter architecture than Pro, cutting compute per token and so reducing response latency. That speed comes at the cost of reasoning depth and output quality, satisfying the real-time chat constraint while accepting weaker responses on complex prompts.

Why this answer

Gemini 1.5 Flash is specifically optimized for lower latency and cost efficiency, making it ideal for real-time chat applications. However, this optimization comes at the cost of reduced model capacity and reasoning depth compared to Gemini 1.5 Pro, which prioritizes higher quality and accuracy over speed. Therefore, the expected trade-off is lower latency but potentially lower quality.

Exam trap

Candidates often confuse that faster model means better overall, while the core trade-off is between latency and output quality due to model architecture differences in the Gemini family.

How to eliminate wrong answers

Option A is wrong because it claims both higher quality and lower latency, which contradicts the fundamental trade-off between model complexity and speed; Flash sacrifices quality for speed. Option C is wrong because while Flash does offer lower cost, it does not provide a longer context window—both Flash and Pro support up to 1 million tokens, so this is not a distinguishing trade-off. Option D is wrong because it describes higher accuracy but slower responses, which is characteristic of Gemini 1.5 Pro, not the trade-off when choosing Flash over Pro.

873
MCQeasy

A retail company wants to use a generative AI model to create product descriptions from a list of attributes. They want a fully managed service that allows them to quickly experiment with prompts and adjust parameters without managing infrastructure. Which Google Cloud service should they use?

A.AutoML Natural Language
B.Cloud Natural Language API
C.Dialogflow CX
D.Vertex AI Studio
AnswerD

Vertex AI Studio is a Google Cloud console tool that lets you design, test, and manage prompts for generative AI models. It provides an interactive interface to experiment with different models and parameters, such as temperature and top-k, without managing any infrastructure. This directly matches the requirement for quick experimentation with prompts and parameters in a fully managed environment.

Why this answer

Vertex AI Studio is a Google Cloud service that provides a unified interface for experimenting with generative AI models. It allows users to write prompts, adjust model parameters, and see generated outputs in real time. Since the company wants a fully managed service to quickly create product descriptions without infrastructure management, Vertex AI Studio is the correct choice.

Exam trap

The trap here is confusing generative AI tools with traditional NLP services that analyze rather than generate text.

874
MCQeasy

A retail company plans to use Vertex AI's generative AI to create product descriptions. They need to ensure descriptions are factually accurate and do not misrepresent products. Which strategy should they prioritize?

A.Implement human-in-the-loop review
B.Use prompt engineering
C.Use a larger model
D.Increase temperature parameter
AnswerA

Human-in-the-loop review places a person between model output and publication, catching factual errors and misrepresentations before customers see them. For retail product descriptions where accuracy is a hard requirement, this verification step directly satisfies the constraint that generated text must not misstate product attributes.

Why this answer

Human-in-the-loop (HITL) review is the correct strategy because it directly addresses the need for factual accuracy and prevention of misrepresentation. While generative AI can produce fluent text, it lacks a reliable grounding mechanism for product-specific facts, making human oversight essential to catch hallucinations, verify claims, and ensure compliance with advertising standards. This approach aligns with responsible AI practices and is a core recommendation for high-stakes content generation.

Exam trap

Google Cloud often tests the misconception that prompt engineering or model size alone can solve factual accuracy issues, when in reality, generative AI's inherent lack of ground truth makes human validation indispensable for high-stakes content.

How to eliminate wrong answers

Option B is wrong because prompt engineering, while useful for guiding output style and structure, does not guarantee factual accuracy; it cannot prevent the model from generating plausible-sounding but incorrect product details. Option C is wrong because using a larger model may improve fluency and reduce some errors, but it does not eliminate hallucinations or misrepresentations, and can even introduce more subtle inaccuracies. Option D is wrong because increasing the temperature parameter makes the model's output more random and creative, which increases the risk of generating factually incorrect or misleading descriptions, the opposite of what is needed.

875
MCQeasy

A retail company wants to use gen AI for customer service chatbots. They have a large volume of customer interactions. What is the primary business consideration for deploying a gen AI solution?

A.Minimizing latency at any cost
B.Using open-source models only
C.Choosing the most complex model
D.Ensuring data privacy and compliance
AnswerD

Handling large volumes of customer interactions means personal data flows through the chatbot, so data privacy and compliance is the primary business consideration; it governs lawful processing, retention and regulatory risk under GDPR-style rules, directly satisfying the scenario's scale-driven exposure.

Why this answer

Ensuring data privacy and compliance is the primary business consideration when handling customer data, especially with regulatory requirements like GDPR or CCPA. Option A is wrong because while latency matters, minimizing it at any cost is not the primary consideration and can lead to excessive expenses. Option B is wrong because using only open-source models may limit flexibility and can raise data privacy concerns if not properly vetted.

Option C is wrong because choosing the most complex model often increases cost and latency without necessarily improving business outcomes.

876
MCQmedium

A financial services company wants to use generative AI to summarize large volumes of internal documents while ensuring that sensitive data never leaves their virtual private cloud (VPC). They need a solution that provides Gemini models with enterprise-grade security and data residency controls. Which Google Cloud service should they use?

A.Gemini API in Google AI Studio
B.Document AI
C.Vertex AI with Gemini models
D.Cloud Natural Language API
AnswerC

Vertex AI provides Gemini models with enterprise security features such as VPC Service Controls, Customer-Managed Encryption Keys (CMEK), and data residency. It allows the company to keep sensitive data within their VPC and meet compliance requirements, directly addressing the need for secure generative AI on internal documents.

Why this answer

Vertex AI offers Gemini models with enterprise-grade security, including VPC Service Controls and data residency, which are critical for handling sensitive financial data. It enables the company to use generative AI while maintaining strict control over data location and access. The other services either lack generative capabilities or do not provide the necessary security and compliance features for this scenario.

Exam trap

The trap here is assuming that the Gemini API in Google AI Studio offers the same enterprise security controls as Vertex AI, when it is primarily designed for developer ease of use.

877
MCQmedium

A mid-size retail company wants to launch a generative AI assistant that drafts promotional product descriptions for its marketing team. The team expects a rapid pilot, but the CIO insists that any generated text must never expose the company's unreleased product roadmap or pricing data that may exist in internal documents. The company has no dedicated AI engineering staff and prefers a managed approach on Google Cloud. Which strategy best balances rapid pilot delivery with this data-exposure requirement?

A.Use a managed generative AI API with grounding restricted to an approved, curated product catalog and clear system instructions limiting the assistant to that content.
B.Build a custom large language model from scratch using the company's historical marketing copy as the sole training corpus.
C.Deploy an open-weight model on a Compute Engine VM and allow the assistant to search the entire shared drive so the marketing team gets the most complete answers.
D.Fine-tune a foundational model on all internal product documents so the assistant learns the company's writing style and terminology.
AnswerA

A managed API removes infrastructure and model-ops work, enabling a fast pilot without AI engineering staff. Restricting grounding to an approved catalog and using system instructions keeps generation anchored to vetted content, so unreleased roadmap and pricing documents are never retrieved or exposed. This directly satisfies both the speed goal and the CIO's data-boundary requirement.

Why this answer

The scenario couples a fast, low-skill pilot with a hard boundary around sensitive internal content. A fully managed generative AI API eliminates infrastructure and model-tuning overhead, while grounding limited to an approved catalog plus explicit system instructions confines generation to vetted material, preventing unreleased roadmap and pricing data from appearing in drafts. This combination delivers speed without weakening the data-exposure control the CIO demanded.

Exam trap

The trap here is assuming that fine-tuning or self-hosting is required for brand consistency, when grounding and instructions actually control content boundaries without embedding sensitive data in model weights.

878
MCQhard

A healthcare analytics team is evaluating a generative AI model for summarizing clinical notes. They observe that summaries are fluent but occasionally invent details not present in the source notes. They want a practical mitigation that reduces fabricated content without retraining the model. Which approach should they apply first?

A.Raise the temperature to encourage the model to explore alternative phrasings of the notes
B.Fine-tune the model on a large corpus of unrelated medical literature to improve fluency
C.Lower the temperature and instruct the model to answer only from the provided note text
D.Increase the maximum output token limit so the model has more room to explain itself
AnswerC

Lowering temperature reduces random sampling, and an explicit instruction to use only the supplied note constrains the model to the given context. Together they reduce the chance of invented details while keeping the model unchanged. This is a low-cost, immediate mitigation that targets the observed hallucination without retraining.

Why this answer

Reducing temperature and explicitly instructing the model to rely only on the provided note constrains generation toward the source content, cutting down fabricated details without retraining. It is fast, inexpensive, and directly targets the failure mode. Longer output, higher randomness, or unrelated fine-tuning do not improve fidelity and can make hallucination worse.

Exam trap

The trap here is assuming more output space or more training data fixes hallucination, when the first effective lever is constraining sampling and instructing the model to use only the supplied context.

879
MCQhard

An organization is running a large-scale training job for a custom NLP model with a batch size of 2048 and sequence length of 512. They need to minimize training time while keeping costs predictable. Which Google Cloud hardware should they choose?

A.Cloud TPU v5e pods
B.Compute Engine with NVIDIA A100 GPUs
C.Edge TPU devices
D.Compute Engine with NVIDIA T4 GPUs
AnswerA

Cloud TPU v5e pods interconnect thousands of chips over high-bandwidth links, scaling efficiently for large-batch, long-sequence training. This parallelism cuts wall-clock training time, and TPU pod-hour pricing keeps spend predictable, meeting both the speed and cost-certainty constraints.

Why this answer

Cloud TPU v5e pods are purpose-built for large-scale training of transformer-based NLP models, offering high-throughput matrix multiplication and efficient scaling across multiple chips. With a batch size of 2048 and sequence length of 512, TPU v5e pods deliver superior training speed and predictable pricing via reserved capacity, minimizing time-to-train compared to GPU alternatives.

Exam trap

The trap here is that candidates often default to choosing NVIDIA A100 GPUs due to their general popularity, overlooking that TPU pods are specifically optimized for large-scale transformer training with predictable pricing and superior scaling efficiency.

How to eliminate wrong answers

Option B is wrong because NVIDIA A100 GPUs, while powerful, are general-purpose accelerators that lack the dedicated matrix-multiply units (MXU) and high-bandwidth interconnects of TPU pods, leading to higher cost and slower training for large-batch transformer workloads. Option C is wrong because Edge TPU devices are designed for low-power inference at the edge, not for large-scale training, and cannot handle batch sizes of 2048 or sequence lengths of 512. Option D is wrong because NVIDIA T4 GPUs are mid-range inference and training GPUs with lower memory bandwidth and fewer tensor cores, making them unsuitable for large-batch NLP training and resulting in significantly longer training times.

880
MCQeasy

A media company wants to produce short video clips from text prompts for social media campaigns. They need a Google Cloud service that can generate video from text and edit existing video clips, with enterprise-grade controls. Which Google Cloud offering should they choose?

A.Vertex AI Gemini
B.Vertex AI Veo
C.Vertex AI Imagen
D.Vertex AI Chirp
AnswerB

Veo on Vertex AI is Google Cloud's generative video model that creates high-quality video from text or image prompts and supports video editing capabilities. It is integrated with enterprise controls such as IAM, VPC Service Controls, and data governance, making it the right fit for a media company needing controlled video generation and editing.

Why this answer

Veo on Vertex AI is purpose-built for generative video, producing clips from text or image prompts and supporting video editing. It is managed within Vertex AI, so organizations get enterprise security, access control, and compliance features. Other models in Vertex AI handle images, text, or speech, but not video generation and editing in one service.

Exam trap

The trap here is assuming any multimodal model such as Gemini can generate video, when video synthesis specifically requires a dedicated video generation model like Veo.

881
MCQhard

A data scientist fine-tunes a generative image captioning model to describe medical images. The model outputs safe but very generic captions (e.g., 'An image of cells'). The goal is to produce more specific, clinically relevant descriptions. Which approach is most effective?

A.Perform incremental fine-tuning on a curated dataset of detailed medical image captions.
B.Use diverse beam search during decoding to generate multiple caption candidates.
C.Adjust top-k sampling to restrict the vocabulary to medical terms only.
D.Increase the temperature to encourage the model to output longer, more varied captions.
AnswerA

Incremental fine-tuning on curated detailed medical captions shifts the model's output distribution towards specific clinical vocabulary while preserving its existing image understanding. Generic captions persist because the original training data lacked that specificity, which targeted examples now supply.

Why this answer

Incremental fine-tuning on a high-quality dataset of specific medical captions directly teaches the model the desired level of detail. Option B is wrong because diverse beam search generates varied outputs but does not inherently improve specificity; it may produce multiple candidates but does not ensure clinical relevance. Option C is wrong because top-k sampling can reduce the vocabulary but does not guarantee medical accuracy or detail.

Option D is wrong because increasing temperature adds randomness, which may produce longer or more varied captions but often introduces irrelevant words rather than specific clinical terms.

882
MCQmedium

A company wants to deploy a generative AI chatbot for customer service but is concerned about cost unpredictability due to variable usage. Which pricing model should they choose to best manage costs?

A.Committed use discounts
B.Free tier
C.Pay-as-you-go
D.Provisioned throughput
AnswerD

Provides fixed capacity with predictable monthly cost.

Why this answer

Provisioned throughput provides fixed capacity with predictable monthly cost, ideal for managing cost uncertainty. Option A (committed use discounts) requires a commitment but costs can still vary if usage exceeds the committed amount. Option B (free tier) is too limited for a full-scale chatbot.

Option C (pay-as-you-go) is variable and leads to cost unpredictability.

883
Multi-Selectmedium

A data scientist wants to improve the performance of a text classification model for customer feedback. They have a small labeled dataset of 500 examples and a large unlabeled corpus of 100,000 feedback messages. Which TWO strategies would be most effective? (Choose 2)

Select 2 answers
A.Increase the context window of the model
B.Apply semi-supervised learning by pseudo-labeling the unlabeled data
C.Use RAG to retrieve similar examples from the unlabeled corpus during inference
D.Train a model from scratch on the labeled data only
E.Use a pre-trained LLM (e.g., Gemini) and fine-tune on the labeled data
AnswersB, E

Pseudo-labeling assigns predicted labels to the 100,000 unlabeled messages, then retrains on the combined set. This exploits the large unlabeled corpus to compensate for the 500-example labelled set, expanding effective training signal without new annotation.

Why this answer

Option B is correct because semi-supervised learning via pseudo-labeling leverages the large unlabeled corpus: the model is first trained on the 500 labeled examples, then used to predict labels on the 100,000 unlabeled messages, and high-confidence predictions are added back as training data, effectively expanding the labeled set and improving classification performance. Option E is correct because using a pre-trained LLM such as Gemini and fine-tuning it on the small labeled dataset exploits transfer learning: the model already encodes broad language understanding from large-scale pre-training, so only a small amount of task-specific labeled data is needed to adapt it to customer feedback classification, which is far more sample-efficient than training from scratch. Option A is not appropriate because increasing the context window affects how much input text the model can attend to, not the model's ability to learn from limited labels, and does not address the small-labeled-data problem.

Option C is not appropriate because RAG retrieves relevant documents to augment generation at inference time; it does not train or improve the classification model's parameters and is not a training strategy for a text classifier. Option D is not appropriate because training a model from scratch on only 500 labeled examples will almost certainly overfit and underperform compared to transfer learning or semi-supervised approaches.

Exam trap

A common misconception is that RAG can improve model training, but RAG retrieves information during inference and does not augment the training data. In contrast, pseudo-labeling directly expands the training set.

884
Multi-Selecthard

A company is deploying a GenAI system that generates product descriptions. During A/B testing, the new system shows a 20% increase in click-through rate (CTR) but a 15% increase in average cost per query due to the model size. The team wants to optimize cost without sacrificing the CTR gain. Which THREE actions should they take? (Choose three.)

Select 3 answers
A.Batch similar requests to reduce per-request overhead
B.Use a larger model with higher accuracy to further increase CTR
C.Increase the number of few-shot examples in the prompt
D.Switch to a smaller model and re-A/B test to confirm CTR impact
E.Implement response caching for repeated product SKUs
AnswersA, D, E

Batching reduces the number of API calls and can lower cost.

Why this answer

Batching similar requests reduces the per-request overhead by combining multiple inference calls into a single batch, which amortizes the fixed costs (e.g., model loading, token processing) across more outputs. This directly lowers the average cost per query while preserving the model architecture and CTR gains, as the model's output quality remains unchanged.

Exam trap

Google often tests the misconception that adding more few-shot examples always improves output quality, but in reality, it increases token costs and can degrade performance due to context window limits or irrelevant examples.

885
MCQmedium

A research team wants to use Google's AI to generate video content from text prompts for a creative project. Which Google Cloud generative AI model should they use?

A.Imagen
B.Codey
C.Veo
D.Gemini
AnswerC

Veo is Google Cloud's generative video model, accepting text prompts and producing video clips, which directly satisfies the stem's requirement to generate video content from text. Other Gemini and Imagen models output text or images respectively, so they cannot fulfil the video-generation constraint.

Why this answer

Veo is Google's generative AI model specifically designed for text-to-video generation, making it the correct choice for creating video content from text prompts. It is part of Google's Vertex AI model portfolio and supports high-definition video generation with cinematic controls. Imagen, Codey, and Gemini serve different modalities.

Exam trap

Generative AI Leader often tests model-to-modality mapping — candidates confuse Imagen (image) with Veo (video) because both are generative media models with similar-sounding names.

How to eliminate wrong answers

Option A is wrong because Imagen is Google's text-to-image generation model, not video. Option B is wrong because Codey is a family of code generation models (Codey for Code Completion, Codey for Code Chat), not video. Option D is wrong because Gemini is a multimodal large language model for text, image, and code understanding and generation, but it does not natively generate video output.

886
Multi-Selecthard

A healthcare analytics team uses a Gemini model to generate patient-friendly discharge summaries from clinical notes. Clinicians report that summaries occasionally omit critical follow-up instructions. The team must improve recall of these instructions without retraining the model. Which two techniques should they apply? (Choose two.)

Select 2 answers
A.Reduce the maximum output tokens to force the model to be more concise and include only essentials.
B.Increase the temperature so the model produces more varied summaries and captures more details.
C.Enable deterministic decoding and assume the model will then always include every instruction.
D.Add an explicit instruction to extract all follow-up actions into a mandatory checklist section of the summary.
E.Use few-shot prompting with examples that demonstrate surfacing every follow-up instruction in a dedicated section.
AnswersD, E

A direct, unambiguous instruction to place every follow-up action in a required section gives the model a clear structural target and reduces the chance that an item is dropped during free-form generation. Combined with a defined output schema, it makes omissions visible and easier to catch during review.

Why this answer

Few-shot examples and an explicit mandatory checklist section both shape the model toward systematically extracting follow-up actions, improving recall without retraining. Temperature changes, token limits, and deterministic decoding affect randomness or length, not whether required clinical content is captured in the output.

Exam trap

The trap here is equating deterministic or lower-variance decoding with completeness, when consistency and recall are separate properties of a generative system.

887
MCQeasy

Which Google Cloud service allows you to run machine learning models directly using SQL queries on data in BigQuery?

A.Cloud Functions
B.BigQuery ML
C.Vertex AI
D.Dataflow
AnswerB

BigQuery ML lets you create and run machine learning models inside BigQuery using standard SQL syntax, satisfying the requirement to train and predict without exporting data or writing Python. It supports linear regression, logistic regression, k-means, and imported TensorFlow models, so the SQL-only constraint in the stem is met directly.

Why this answer

BigQuery ML enables users to create, train, and deploy ML models using standard SQL, eliminating the need to move data to a separate environment.

888
MCQmedium

A healthcare company is using a generative AI model to summarize patient notes. The model occasionally includes hallucinated medical details. Which strategy best reduces hallucinations?

A.Use a larger model with more parameters.
B.Fine-tune the model on a large corpus of general medical literature.
C.Increase the model's temperature to encourage more diverse outputs.
D.Provide the model with the patient notes as context and instruct it to only use that information.
AnswerD

Grounding the model in the source text and explicitly instructing it to rely only on that context reduces hallucinations because the model can copy or paraphrase from the provided notes. This is a form of retrieval-augmented generation or prompt grounding. It constrains the model's output to the given facts.

Why this answer

Grounding the model by providing the patient notes as context and instructing it to only use that information is the most effective way to reduce hallucinations. This approach, often called retrieval-augmented generation or prompt grounding, forces the model to base its summary on the given text rather than relying on its parametric memory. It also allows for verification against the source.

Exam trap

The trap here is assuming that a larger or fine-tuned model will automatically be more factual, when the key is to constrain generation to the provided source text.

889
Multi-Selectmedium

A data scientist is fine-tuning a generative AI model for customer sentiment analysis. To ensure the fine-tuned model does not inadvertently memorize and reproduce personally identifiable information (PII) from the training data, which THREE practices should they follow? (Select 3)

Select 3 answers
A.Apply differential privacy during fine-tuning
B.Apply model quantization to reduce model size
C.Remove or anonymize all PII from the training data
D.Use only a subset of data that is necessary for the task (data minimization)
E.Use k-fold cross-validation to evaluate the model
AnswersA, C, D

Differential privacy adds calibrated noise during fine-tuning, bounding any single record's influence on learned parameters. This mathematically limits memorisation, so the model cannot reliably reproduce PII from training examples, directly satisfying the stem's requirement to prevent PII reproduction.

Why this answer

Option A is correct because applying differential privacy during fine-tuning adds calibrated noise to the training process, mathematically bounding the influence any single training record can have on the model's parameters and thereby limiting memorization of PII. Option C is correct because removing or anonymizing PII in the training data eliminates the sensitive content before it can ever be learned, which is the most direct defense against the model reproducing it. Option D is correct because data minimization restricts training to only the data necessary for the sentiment-analysis task, reducing the volume of PII exposed to the model and shrinking the attack surface for memorization.

Option B is not correct because quantization only compresses weights to lower precision to reduce size and inference cost; it does not prevent the model from having memorized PII during training. Option E is not correct because k-fold cross-validation is an evaluation/resampling technique for estimating generalization performance and does not remove, mask, or otherwise protect PII in the training data.

Exam trap

Generative AI Leader often tests privacy-preserving techniques; candidates may confuse model optimization techniques (quantization) or evaluation methods (cross-validation) with privacy practices.

890
MCQhard

An enterprise deploys a generative AI chatbot that must comply with GDPR right to deletion. Users can request deletion of their personal data. The chatbot uses a RAG pipeline with a vector database. What is the MOST effective way to handle deletion requests?

A.Delete the user's documents from the vector index and original storage, then rebuild the index
B.Update the user's records in the vector index with anonymized placeholders
C.Add a filter to the chat application to block the user's name from appearing in responses
D.Retrain the LLM from scratch without the user's data
AnswerA

Rebuilding the index after removing the user's documents guarantees no residual embeddings remain, satisfying GDPR erasure. Simply deleting source files leaves vectors containing personal data still queryable, so full index reconstruction is the only method that reliably eliminates every trace.

Why this answer

GDPR Article 17 (right to erasure) requires that personal data be deleted from all systems where it is stored, including derived stores like vector indexes. In a RAG pipeline, the user's documents exist both in original storage and as embeddings in the vector database, so the compliant approach is to delete from both and rebuild or re-index the affected vectors. Simply filtering responses or anonymizing records leaves the personal data recoverable and does not satisfy erasure.

Exam trap

Generative AI Leader often tests the misconception that filtering outputs or anonymizing index entries satisfies deletion — the trap is confusing runtime privacy controls with actual data erasure across all derived stores.

How to eliminate wrong answers

Option B is wrong because replacing records with anonymized placeholders leaves residual embedding vectors that may still encode the individual's personal data, and GDPR requires actual erasure, not pseudonymization of the index entry. Option C is wrong because a response filter is a runtime control, not a deletion — the personal data remains in the vector store and original documents, violating the right to erasure. Option D is wrong because retraining the foundation LLM from scratch is neither necessary nor sufficient: the personal data lives in the RAG corpus and vector index, not in the base model weights, and retraining is prohibitively expensive and does not address the vector store.

891
Multi-Selectmedium

A company wants to build a GenAI-powered customer support chatbot. They require the chatbot to provide accurate answers based on the latest product documentation, and they need to control costs by minimizing token usage. Which TWO strategies should they use?

Select 2 answers
A.Use the largest available model for maximum accuracy
B.Cache frequent queries and their responses
C.Implement Retrieval-Augmented Generation (RAG) to retrieve relevant document chunks
D.Use a longer context window to include entire documents in the prompt
E.Fine-tune a large language model on the product documentation
AnswersB, C

Caching frequent queries and their responses directly reduces token usage: repeated questions are served from the cache rather than re-sent to the model, so no input or output tokens are consumed. This satisfies the stem's cost constraint of minimising token usage, while accurate answers still derive from the latest product documentation.

Why this answer

Option B is correct because caching frequent queries and their responses avoids re-invoking the model for identical requests, directly reducing token consumption and cost for repeated customer questions. Option C is correct because Retrieval-Augmented Generation (RAG) retrieves only the most relevant document chunks from the latest product documentation and injects just those into the prompt, ensuring accurate, up-to-date answers while keeping the token payload small. Option A is not appropriate because using the largest model increases cost per token without addressing retrieval of current documentation.

Option D is not appropriate because stuffing entire documents into a long context window inflates token usage, contradicting the cost-minimization requirement. Option E is not appropriate because fine-tuning on documentation is expensive, must be repeated as docs change, and does not inherently minimize per-query token usage.

Exam trap

Generative AI Leader often tests the misconception that a larger context window or fine-tuning eliminates the need for retrieval and caching, but the exam expects you to recognize that RAG and caching are the primary strategies for balancing accuracy and token cost.

892
MCQmedium

A company is building a document summarization tool using Vertex AI Gemini API. They notice that the model sometimes returns incomplete summaries that miss key points. Which approach is most likely to improve summary quality without increasing token usage significantly?

A.Refine the system instruction to specify the desired summary format and key elements to include
B.Increase the context window to include more of the document
C.Switch to a larger Gemini model (e.g., from 1.0 Pro to 1.5 Pro)
D.Increase the max output token limit to allow longer summaries
AnswerA

Refining the system instruction specifies the required summary format and key elements, steering the model to cover them without enlarging the input. This improves completeness at negligible token cost, unlike approaches that add lengthy examples or extra context.

Why this answer

Refining the system instruction directly addresses the root cause of incomplete summaries by providing explicit guidance on the desired output format and key elements to include. This approach improves the model's adherence to the task without increasing the number of input or output tokens, as it only modifies the instruction text, not the document length or generation limits.

Exam trap

In Google Cloud exams, a common misconception is that increasing model size or token limits directly improves output quality, when in fact prompt engineering—specifically system instructions—is a more efficient and cost-effective lever for controlling model behavior.

How to eliminate wrong answers

Option B is wrong because increasing the context window adds more tokens from the document, which does not guarantee the model will focus on key points and can actually dilute attention, increasing token usage without improving summary completeness. Option C is wrong because switching to a larger Gemini model (e.g., from 1.0 Pro to 1.5 Pro) increases computational cost and token usage (due to larger model overhead) but does not inherently fix the instruction quality; the issue is prompt design, not model capacity. Option D is wrong because increasing the max output token limit allows longer summaries but does not ensure the model includes missing key points; it may simply produce more verbose text without addressing the root cause of omission.

893
Multi-Selectmedium

A company is deploying a generative AI system for medical diagnosis. Which TWO measures are essential for responsible AI in this high-stakes domain?

Select 2 answers
A.Allow the AI to make autonomous decisions in time-sensitive emergencies
B.Ensure a human medical professional reviews all AI-generated diagnoses
C.Publish all patient data used for training to ensure transparency
D.Provide model documentation (Model Cards) to clinicians detailing the system's limitations
E.Use the AI only for administrative tasks, not diagnosis
AnswersB, D

Human-in-the-loop review directly addresses the clinical risk constraint: an erroneous AI diagnosis could cause patient harm. Because generative models can hallucinate confidently, a qualified clinician must validate every output before it informs care, preserving professional accountability and satisfying medical-device oversight expectations for high-stakes decisions.

Why this answer

Option B is correct because in a high-stakes medical diagnosis scenario, keeping a qualified human medical professional in the loop for every AI-generated diagnosis ensures accountability and mitigates the risk of harmful errors, which is a core principle of responsible AI (human oversight). Option D is correct because Model Cards provide structured documentation of a model's intended use, performance metrics, and known limitations, enabling clinicians to understand when and how much to trust the system's outputs. Together, these measures support transparency and human accountability, which are essential in clinical decision-making.

Option A is not appropriate because autonomous AI decisions in emergencies remove the necessary human oversight and could cause serious harm. Option C is wrong because publishing all patient data would violate privacy and confidentiality regulations such as HIPAA, and transparency does not require exposing raw patient data. Option E is incorrect because restricting AI to administrative tasks does not address responsible AI for diagnosis and is not an essential measure for the described use case.

Exam trap

Generative AI Leader often tests the balance between innovation and safety—candidates may select 'autonomous decisions' or 'publish patient data' as transparency measures, confusing openness with responsible AI.

894
MCQeasy

A developer wants to improve the factual accuracy of the model's summaries. Based on the exhibit, what should they do?

A.Enable the support engine.
B.Increase the model's context window.
C.Configure grounding with a knowledge base.
D.Re-train the model with a dataset of facts.
AnswerC

Grounding with a knowledge base retrieves authoritative source content and supplies it to the model at generation time, so summaries cite retrieved facts rather than relying on parametric memory. This directly reduces hallucination and improves factual accuracy of the summaries.

Why this answer

Grounding with a knowledge base is the correct approach because it anchors the model's output to a trusted, external source of facts, directly improving factual accuracy without modifying the model's weights. This technique uses retrieval-augmented generation (RAG) to fetch relevant documents from the knowledge base and inject them into the prompt context, ensuring the summary is based on verified information rather than relying solely on the model's parametric memory.

Exam trap

A common pitfall is mistaking grounding for simply expanding the context window or retraining. Grounding with a knowledge base (RAG) provides direct access to verified facts, which is more effective and efficient than further training or fine-tuning, especially in a production environment.

How to eliminate wrong answers

Option A is wrong because enabling the support engine typically refers to a customer support or troubleshooting tool, not a mechanism for improving factual accuracy in generative AI summaries; it does not provide a knowledge base for grounding. Option B is wrong because increasing the model's context window only allows the model to process more tokens in a single request, but it does not introduce new factual information or correct hallucinations; it may even amplify errors if the additional context is unverified. Option D is wrong because re-training the model with a dataset of facts is a costly, time-consuming process that requires significant computational resources and expertise, and it does not guarantee factual accuracy for unseen or evolving information; grounding with a knowledge base is a more efficient and dynamic solution.

895
MCQeasy

You are a generative AI lead at a healthcare startup developing a system to summarize patient medical records for quick review by doctors. The system uses a fine-tuned LLM. After deployment, doctors report that the summaries often miss critical details like medication dosages and allergy information. The current pipeline preprocesses patient records by extracting text from EHR, feeding it to the LLM, and outputting a summary. The team has limited time and budget. They cannot retrain the model because it is hosted as a managed API. Which action should you take to most effectively improve the summarization quality without changing the model?

A.Increase the maximum output token limit to force the model to include more details.
B.Replace the LLM with a simpler extractive summarization model that selects sentences from the original document.
C.Implement a retrieval-augmented generation (RAG) system that pulls supplementary data from external drug databases.
D.Revise the prompt to explicitly ask for medication dosages and allergies, and format the input text by adding headings (e.g., '### Medications') to emphasize important sections.
AnswerD

Prompt engineering and structured input formatting directly address the omission of dosages and allergies without touching the managed API model, satisfying the no-retraining constraint. Explicit instructions plus section headings steer the LLM's attention to clinically critical fields, a low-cost, rapid fix suited to the limited time and budget.

Why this answer

Prompt engineering is the most effective and cost-efficient way to improve LLM output without retraining or changing the model. By explicitly instructing the model to include medication dosages and allergies, and by structuring the input with clear headings, you guide the model's attention to critical sections, directly addressing the missing details. This approach leverages the LLM's existing capabilities and requires no changes to the hosted API or additional infrastructure.

Exam trap

Google Cloud often tests the misconception that increasing output length or adding external data automatically improves quality, when in fact the most direct and cost-effective fix is to refine the input prompt to guide the model's focus.

How to eliminate wrong answers

Option A is wrong because increasing the maximum output token limit does not force the model to include specific missing details; it only allows longer responses, which may still omit critical information if the prompt does not direct the model's focus. Option B is wrong because replacing the LLM with an extractive summarization model would require retraining or deploying a new model, contradicting the constraint of not changing the model, and extractive methods cannot generate new text to explicitly mention dosages or allergies if they are not present in the original text. Option C is wrong because implementing a RAG system to pull from external drug databases adds complexity, cost, and latency, and does not address the core issue of missing details from the patient's own records; the problem is about extracting existing information, not supplementing with external data.

896
Multi-Selecthard

A company is fine-tuning a Gemma model using Vertex AI. They observe that the model overfits. Which TWO actions should they take to mitigate overfitting?

Select 2 answers
A.Use a larger batch size
B.Increase the number of training epochs
C.Use more diverse data
D.Reduce the learning rate
E.Add dropout during fine-tuning
AnswersC, E

Overfitting arises when training data is too narrow, letting the model memorise specifics rather than general patterns. Expanding the fine-tuning dataset with more varied examples exposes Gemma to broader linguistic patterns, improving generalisation and reducing the gap between training and validation performance.

Why this answer

Option C (Use more diverse data) is correct because overfitting occurs when the fine-tuning dataset is too narrow or unrepresentative, causing the model to memorize specific patterns rather than generalize; expanding the training data with more varied examples exposes the model to a broader distribution and reduces memorization. Option E (Add dropout during fine-tuning) is correct because dropout randomly deactivates a fraction of neurons during training, which acts as a regularization technique that prevents the model from relying too heavily on any single pathway and improves generalization. Option A (Use a larger batch size) is not a reliable overfitting remedy and can even reduce regularization effects from gradient noise, so it is not among the correct answers.

Option B (Increase the number of training epochs) would worsen overfitting by allowing the model to fit the training data even more closely. Option D (Reduce the learning rate) may stabilize training but does not directly address overfitting and can, in some cases, allow the model to converge to a sharper minimum that generalizes poorly.

Exam trap

Google Cloud often tests the misconception that reducing the learning rate or increasing batch size are universal fixes for overfitting, when in fact these hyperparameters primarily affect optimization dynamics rather than regularization.

897
MCQeasy

A retail company wants to integrate generative AI into its customer service chatbot to handle routine inquiries. They have a limited budget and want to launch quickly. Which strategy is most appropriate?

A.Partner with a generative AI vendor for a custom solution
B.Use pre-trained models via Google Cloud's Generative AI Studio API
C.Fine-tune an open-source model on their customer service logs
D.Build a custom LLM from scratch using the company's own data
AnswerB

Pre-trained models accessed through an API remove the cost and time of training or hosting custom models, letting the retailer launch quickly within budget. The API handles inference, so only integration work remains, satisfying both the limited-budget and fast-launch constraints.

Why this answer

Using pre-trained models via Google Cloud's Generative AI Studio API allows the company to leverage existing, powerful models without the high cost and time investment of custom development or fine-tuning. This approach enables rapid deployment on a limited budget by simply integrating the API into their chatbot, handling routine inquiries effectively without requiring extensive machine learning expertise or infrastructure.

Exam trap

Google Cloud often tests the misconception that fine-tuning or custom models are always better for domain-specific tasks, but the trap here is that for routine inquiries with limited budget and time, pre-trained APIs offer the fastest and most cost-effective solution without sacrificing quality.

How to eliminate wrong answers

Option A is wrong because partnering with a generative AI vendor for a custom solution typically involves significant upfront costs, long development cycles, and vendor lock-in, which contradicts the company's limited budget and need for quick launch. Option C is wrong because fine-tuning an open-source model on customer service logs requires substantial computational resources, data preparation, and machine learning expertise, making it slower and more expensive than using a pre-trained API. Option D is wrong because building a custom LLM from scratch is extremely resource-intensive, requiring massive datasets, specialized hardware, and months of training, which is impractical for a company with limited budget and a need for speed.

898
MCQmedium

A research team wants to document the intended uses, limitations, and ethical considerations of their newly trained image classification model. Which Google Cloud tool should they use?

A.Explainable AI SDK
B.Datasheets for Datasets
C.Model Card Toolkit
D.What-If Tool
AnswerC

Model Card Toolkit generates structured Model Cards capturing intended use, limitations, metrics and ethical considerations for trained models. It is purpose-built for documenting image classification models, directly satisfying the team's requirement to record intended uses, limitations and ethical considerations in a standardised format.

Why this answer

The Model Card Toolkit is specifically designed to document the intended uses, limitations, and ethical considerations of machine learning models, including image classification models. It generates a structured model card that provides transparency and accountability, which aligns directly with the team's goal of documenting these aspects.

Exam trap

Candidates often confuse Explainable AI SDK (for local explanations) with Model Card Toolkit (for global documentation) in Google Cloud.

How to eliminate wrong answers

Option A is wrong because the Explainable AI SDK focuses on providing feature attributions and explanations for individual predictions, not on documenting the model's intended uses, limitations, or ethical considerations. Option B is wrong because Datasheets for Datasets is a tool for documenting datasets, not models; it covers dataset characteristics, collection methods, and biases, but does not address model-level documentation. Option D is wrong because the What-If Tool is an interactive visualization tool for exploring model behavior and fairness across different slices of data, but it does not generate a static documentation artifact like a model card.

899
MCQeasy

A small marketing analytics team wants to build a generative AI assistant that can answer questions about their proprietary campaign performance data. They have very limited machine learning engineering resources and want the fastest possible path to a working prototype on Google Cloud. Which approach should they take?

A.Deploy an open-source large language model on a self-managed Google Kubernetes Engine cluster with autoscaling GPU node pools.
B.Train a custom transformer model from scratch on their campaign data using Vertex AI Training with A3 machine types.
C.Create a BigQuery ML remote model that calls a third-party public API and join it to campaign tables in scheduled queries.
D.Use the Gemini models through Vertex AI with a managed RAG Engine backed by their campaign data stored in BigQuery.
AnswerD

The Gemini models on Vertex AI deliver strong general reasoning and generation out of the box, while the managed RAG Engine handles chunking, embedding, retrieval, and grounding against BigQuery data without the team operating any serving or indexing infrastructure. This directly matches the requirement for a fast prototype with minimal ML engineering effort, since the team only supplies data and prompts.

Why this answer

The key requirement is speed to a working prototype with minimal ML engineering, and managed Gemini endpoints plus a managed RAG pipeline satisfy that without infrastructure ownership. The team supplies campaign data and prompts while Google Cloud handles retrieval, grounding, and serving. Options that require training from scratch, self-managing GPU clusters, or calling external APIs all add engineering burden or governance risk that the scenario rules out.

Exam trap

The trap here is assuming that any generative AI prototype must involve training or hosting a model yourself, when managed Gemini endpoints with grounded retrieval remove that burden entirely.

900
MCQhard

A company is deploying a chatbot using Vertex AI and wants to ensure that the model's responses are grounded in Google Search results to reduce hallucinations. Which feature should they enable?

A.Vertex AI Agent Builder
B.Retrieval-Augmented Generation (RAG) with internal documents
C.Gemini API with custom fine-tuning
D.Vertex AI Search for grounding
AnswerD

Vertex AI Search grounding retrieves live Google Search results and passes them as context to the model, constraining generation to verifiable web content. This directly satisfies the stem's requirement to reduce hallucinations by anchoring responses in current search data rather than relying solely on parametric knowledge.

Why this answer

Google Search grounding in Vertex AI allows the model to retrieve real-time information from Google Search to ground answers. RAG uses private data, not web search. Vertex AI Agent Builder includes grounding but is for building agents, not specifically the grounding feature itself.

Page 11

Page 12 of 14

Page 13