Courseiva

Google Cloud Generative AI Leader Generative AI Leader (Generative AI Leader) — Questions 601–675

1008 questions total · 14pages · All types, answers revealed

Page 8

Page 9 of 14

Page 10
601
MCQmedium

A startup develops a generative AI tool for legal document review. To ensure explainability, they want the model to cite specific clauses from source documents when making assertions. Which technique should they use?

A.Fine-tuning on legal documents with citation examples
B.Chain-of-thought prompting
C.Grounding using a retrieval system that provides source documents
D.Prompt engineering to ask for citations
AnswerC

Grounding retrieves passages from the source documents and injects them into the model's context, so each assertion can be tied to a specific clause. This directly satisfies the explainability requirement by making citations traceable to retrieved text rather than relying on parametric memory, which cannot attribute claims to sources.

Why this answer

Grounding with a retrieval system (RAG) provides the model with actual source documents at inference time, so it can cite specific clauses verbatim from those sources rather than generating citations from memory. This directly supports explainability because every assertion can be traced back to a retrieved passage. Fine-tuning and prompting alone cannot guarantee accurate citations because the model may hallucinate references.

Exam trap

Generative AI Leader often tests the misconception that fine-tuning or prompt engineering alone can produce reliable citations, when grounding via retrieval is required to cite actual source content.

How to eliminate wrong answers

Option A is wrong because fine-tuning on citation examples teaches the model a citation style but does not give it access to the actual source documents at inference time — it can still hallucinate clause references. Option B is wrong because chain-of-thought prompting improves reasoning transparency but does not provide source documents, so citations would still be generated from parametric memory. Option D is wrong because prompt engineering asking for citations cannot force the model to cite real clauses it has not been given — it will fabricate plausible-looking references.

602
MCQmedium

A healthcare startup needs to process sensitive patient data using NLP models on Google Cloud. They require HIPAA compliance and the ability to run models within their VPC. Which service should they use to access Gemini models?

A.Gemini API directly via API key
B.BigQuery ML
C.Vertex AI
D.Google AI Studio
AnswerC

Vertex AI hosts Gemini models inside the customer's project, letting workloads run within a VPC through Private Service Connect or VPC Service Controls, and it is covered by Google Cloud's HIPAA BAA. That combination satisfies both the compliance and in-VPC execution constraints.

Why this answer

Vertex AI provides enterprise-grade features including VPC Service Controls, data isolation, and HIPAA BAA. Google AI Studio is free-tier prototyping only and does not offer these compliance or security controls.

603
MCQmedium

A company is piloting a GenAI feature for email drafting in Gmail. They want to measure productivity improvement. Which metric is MOST directly tied to the business goal of reducing time spent on email composition?

A.Adoption rate of the GenAI feature among the pilot group
B.Increase in employee satisfaction survey scores
C.Reduction in average time-to-send per email
D.Reduction in total tokens consumed per email
AnswerC

Time-to-send per email directly quantifies the composition effort the pilot targets, so a reduction evidences the productivity gain. Unlike output volume or subjective satisfaction, it isolates the specific workflow step the GenAI feature accelerates, tying measurement to the stated goal of cutting composition time.

Why this answer

The primary business goal is to reduce the time employees spend composing emails. Measuring the average time-to-send per email directly quantifies this efficiency gain, as it captures the end-to-end duration from initiation to dispatch, which the GenAI feature aims to shorten by generating draft content.

Exam trap

Google often tests the distinction between proxy metrics (like token consumption or adoption) and direct business outcome metrics, so the trap here is that candidates confuse technical efficiency (fewer tokens) with user productivity (time saved).

How to eliminate wrong answers

Option A is wrong because adoption rate measures how many users try the feature, not the actual productivity impact; high adoption could occur even if the feature does not save time. Option B is wrong because employee satisfaction scores are a lagging, subjective indicator that can be influenced by factors unrelated to email composition speed, such as overall morale or feature usability. Option D is wrong because reduction in total tokens consumed per email measures model efficiency or verbosity, not the business-relevant time savings; fewer tokens do not guarantee faster composition due to latency or user review time.

604
MCQeasy

A company wants to estimate the total cost of ownership (TCO) for a gen AI solution on Google Cloud. Which factors are most important?

A.Only model training cost
B.Compute, storage, and API call costs
C.Only inference cost
D.Only compute cost
AnswerB

Compute, storage, and API call costs are the primary recurring drivers of gen AI TCO, since inference consumes GPU or TPU compute, datasets and embeddings consume storage, and model or API invocations are metered per call. These directly determine ongoing operational spend.

Why this answer

The total cost of ownership (TCO) for a generative AI solution on Google Cloud encompasses all operational expenses, including compute (e.g., TPU/GPU instances for training and inference), storage (e.g., Cloud Storage for datasets and model artifacts), and API call costs (e.g., Vertex AI prediction requests). Focusing on a single cost component, such as training or inference alone, ignores the recurring expenses of serving the model and storing data, which often dominate long-term TCO.

Exam trap

Google Cloud often tests the misconception that TCO is dominated by a single cost factor (e.g., training), when in reality, inference and API costs frequently surpass training expenses in production deployments.

How to eliminate wrong answers

Option A is wrong because it ignores inference, storage, and API costs, which are significant for production gen AI solutions where models are queried repeatedly. Option C is wrong because inference cost is only one part of TCO; training, storage, and API overhead also contribute heavily, especially with large models like PaLM 2 or Gemini. Option D is wrong because compute cost alone excludes storage (e.g., model checkpoints, training data) and API call fees (e.g., per-token billing for Vertex AI), leading to an incomplete TCO estimate.

605
MCQmedium

A company is using Vertex AI Model Registry to manage multiple versions of its custom generative model. They want to automatically route a percentage of traffic to a new model version for testing. What should they do?

A.Set up a Cloud Tasks queue to distribute requests
B.Create a new endpoint for each version
C.Deploy both versions to the same endpoint and adjust traffic split settings
D.Use a load balancer in front of the endpoints
AnswerC

Deploying both model versions to one endpoint and adjusting its traffic split routes a defined percentage of requests to the new version. This satisfies the stem's requirement for automatic percentage-based traffic routing for testing, which the Model Registry alone cannot perform.

Why this answer

Vertex AI Endpoints support traffic splitting between model versions.

606
Multi-Selecthard

A company is deploying a generative AI model for medical diagnosis assistance. To comply with both Google's AI Principles and emerging regulations (e.g., EU AI Act), they must ensure appropriate human oversight. Which THREE measures should they implement?

Select 3 answers
A.Use a model with high confidence scores to bypass human review
B.Allow the AI to act autonomously for low-risk cases to reduce workload
C.Require a human clinician to review all AI-generated diagnoses before acting on them
D.Provide an override mechanism that allows the human to reject the AI's recommendation
E.Document the human-in-the-loop process and roles clearly
AnswersC, D, E

Mandatory clinician review before any diagnosis is acted upon places a qualified human directly in the decision loop, satisfying the EU AI Act's requirement for effective human oversight in high-risk medical AI and Google's principle that AI be accountable to people.

Why this answer

Option C is correct because requiring a human clinician to review every AI-generated diagnosis before it is acted upon directly implements human-in-the-loop oversight, which Google's AI Principles and the EU AI Act demand for high-risk medical applications. Option D is correct because an override mechanism ensures the human reviewer retains meaningful authority to reject or correct the AI's recommendation, preventing automation bias and preserving human control. Option E is correct because documenting the human-in-the-loop process, including who reviews what and at which stage, provides the accountability and traceability that regulators and AI governance frameworks require.

Option A is not appropriate because high confidence scores do not eliminate the need for human oversight in a high-risk medical context, and using them to bypass review undermines the required human control. Option B is also not appropriate because allowing autonomous action even for low-risk cases conflicts with the principle of maintaining human oversight for medical diagnosis assistance under the stated compliance requirements.

Exam trap

The trap here is that candidates mistakenly believe high confidence scores or low-risk classifications can justify removing human oversight, but the exam tests that Google's AI Principles and regulatory frameworks like the EU AI Act require human review for all high-risk AI outputs regardless of confidence or perceived risk level.

607
MCQeasy

A marketing team wants to generate consistent brand-aligned social media posts using Vertex AI Studio. Which prompt engineering technique should they use to ensure the output tone matches their brand voice?

A.Set the temperature to 0 and use a long system instruction
B.Provide a few-shot prompt with examples of previous brand-aligned posts
C.Use a zero-shot prompt describing the brand voice
D.Use chain-of-thought prompting to explain the reasoning behind each post
AnswerB

Few-shot prompting supplies concrete examples of prior brand-aligned posts, letting the model infer tone, vocabulary and formatting patterns from them. This conditions output style far more reliably than zero-shot instructions, ensuring generated posts match the established brand voice.

Why this answer

Providing a few-shot prompt with examples of previous brand-aligned posts (B) is the most effective technique because it shows the model concrete patterns of tone, style, and vocabulary, allowing it to generalize and produce consistent output. Few-shot prompting is ideal when the desired style is nuanced and hard to describe abstractly. This directly ensures the generated posts match the brand voice.

Exam trap

The trap is choosing zero-shot with a detailed description or temperature tuning, but the exam expects recognition that few-shot examples are the most reliable way to enforce a specific style like brand voice.

How to eliminate wrong answers

Option A is wrong because setting temperature to 0 makes output deterministic but does not guarantee brand alignment; a long system instruction alone may not capture nuanced tone as effectively as examples. Option C is wrong because a zero-shot prompt describing the brand voice relies on the model's interpretation, which can be inconsistent and lacks concrete examples. Option D is wrong because chain-of-thought prompting improves reasoning for complex tasks, not stylistic alignment; it may even add unnecessary verbosity.

608
MCQmedium

The exhibit shows the output of describing a model on Vertex AI. What does 'modelSource: MODEL_GARDEN' indicate about this model?

A.The model was imported from the Vertex AI Model Garden.
B.The model was trained on Vertex AI from scratch.
C.The model has been exported to Model Garden.
D.The model was fine-tuned using AutoML.
AnswerA

The `modelSource: MODEL_GARDEN` field confirms the model originated from Vertex AI Model Garden, satisfying the stem's requirement to identify the model's provenance. Model Garden supplies curated first-party, open-source and partner models deployable directly into Vertex AI, so this value distinguishes a catalogue-sourced model from one uploaded or trained within the project.

Why this answer

'modelSource: MODEL_GARDEN' explicitly indicates that the model was sourced from Vertex AI Model Garden, which is a curated repository of pre-built and pre-trained foundation models. This field is set when a model is imported from Model Garden, not when it is trained or fine-tuned from scratch within Vertex AI.

Exam trap

The trap here is that candidates confuse 'modelSource' with the model's training or fine-tuning method, assuming 'MODEL_GARDEN' implies the model was trained or fine-tuned on Vertex AI, when in fact it strictly indicates the model was imported from the Model Garden repository.

How to eliminate wrong answers

Option B is wrong because 'modelSource: MODEL_GARDEN' specifically denotes an imported model, not one trained from scratch; models trained on Vertex AI from scratch would have a different source indicator, such as 'CUSTOM' or 'TRAINING_PIPELINE'. Option C is wrong because Model Garden is an import source, not an export destination; exporting a model to Model Garden is not a supported operation—models are imported from Model Garden, not exported to it. Option D is wrong because fine-tuning via AutoML would set a different source field (e.g., 'AUTOML' or 'TRAINING_PIPELINE'), and Model Garden models are typically pre-trained foundation models that may be fine-tuned later, but the source field reflects the origin, not the fine-tuning method.

609
MCQhard

A retail company uses a generative AI model to create personalized product recommendations. The model sometimes generates recommendations that include products the company does not sell. Which technique should be used to prevent the model from generating non-existent products?

A.Increase the model's temperature to allow more creative recommendations.
B.Implement a post-processing filter that checks generated product names against the company's inventory database.
C.Fine-tune the model on a dataset of customer reviews.
D.Use a larger context window to include the entire product catalog in the prompt.
AnswerB

A post-processing filter that validates product names against the actual inventory ensures that only existing products are recommended. This directly prevents the inclusion of non-existent products by catching and removing them before presentation. It is a reliable and straightforward solution for this scenario.

Why this answer

A post-processing filter that cross-references generated product names with the company's inventory database ensures that only real products are recommended. This method is direct, reliable, and does not rely on the model's internal knowledge, effectively preventing hallucinations of non-existent products.

Exam trap

The trap here is assuming that fine-tuning or prompt engineering can completely eliminate hallucinations, when a validation step is often necessary.

610
MCQhard

A streaming platform uses a large generative model for personalized content suggestions. Budget constraints require minimizing inference costs without significantly degrading quality. Which approach is most effective?

A.Deploy the model on higher-end accelerators to save time.
B.Use a distilled version of the model.
C.Implement stronger safety filters to reduce output length.
D.Cache frequent prompts to avoid regeneration.
AnswerB

Distillation trains a smaller model to reproduce the large model's outputs, cutting inference compute and cost while retaining most recommendation quality. This satisfies the budget constraint without the significant quality degradation that cruder reductions would cause.

Why this answer

Distillation trains a smaller 'student' model to mimic a larger 'teacher' model, reducing parameter count and inference latency while retaining most of the recommendation quality. This directly addresses the budget constraint by lowering compute and memory costs per inference, making it the most effective approach among the options.

Exam trap

A common pitfall is assuming that caching frequent prompts or upgrading to higher-end accelerators reduces per-inference costs. Caching only helps with repeated queries, not unique recommendations; hardware upgrades increase fixed costs. Distillation directly reduces model size and inference compute, aligning with cost constraints.

How to eliminate wrong answers

Option A is wrong because deploying on higher-end accelerators increases hardware cost, not reduces it, and while it may save time, the budget constraint demands minimizing inference costs, not just time. Option C is wrong because stronger safety filters do not reduce output length in a meaningful way for cost savings; they add computational overhead for filtering and may degrade user experience by blocking valid suggestions. Option D is wrong because caching frequent prompts only avoids regeneration for identical inputs, but personalized content suggestions are inherently unique per user session, so cache hit rates are low and the approach does not address the core inference cost per unique request.

611
Multi-Selecthard

A bank is deploying a retrieval-augmented generation application on Vertex AI so that a Gemini model answers employee policy questions using the bank's internal document repository. The team wants the model's responses to cite source documents and to reduce fabricated content. Which two capabilities should they implement to ground the model in the bank's own content? (Choose two.)

Select 2 answers
A.Gemini safety filters configured to block sensitive topics
B.Vertex AI embeddings with vector search over the document corpus
C.Grounding with Google Search
D.Vertex AI Search grounding with the internal document index
E.Vertex AI Pipelines scheduled retraining of the Gemini model
AnswersB, D

Generating embeddings for the bank's documents and storing them in a vector index such as Vertex AI Vector Search allows retrieval of semantically similar passages for each query. Those retrieved passages become the grounding context for the Gemini model, enabling responses that reflect internal policy content and can reference the retrieved sources, which fulfills the grounding and citation objectives.

Why this answer

Grounding responses in the bank's own content requires retrieving relevant internal passages and supplying them to the model. Indexing documents with Vertex AI Search or generating embeddings and retrieving them through vector search both provide that context, which improves factual accuracy and supports citations to source documents. Training or safety-filter adjustments do not retrieve proprietary content at inference time and therefore cannot satisfy the citation requirement.

Exam trap

The trap here is treating model retraining or safety filtering as a substitute for retrieval, when grounding in proprietary documents is fundamentally a retrieval-at-inference-time problem.

612
MCQeasy

A project manager wants to understand which Google Cloud generative AI services are subject to the 'Prohibited Use' policy. Where can they find the most up-to-date information?

A.Google Cloud documentation
B.Google's AI Principles
C.The Google Cloud Acceptable Use Policy
D.The Gemini Terms of Service
AnswerC

The Google Cloud Acceptable Use Policy documents which services fall under the Prohibited Use policy, making it the authoritative, current source. This satisfies the stem's need for the most up-to-date information on generative AI service coverage.

Why this answer

The Google Cloud Acceptable Use Policy (AUP) is the authoritative document that defines prohibited uses of Google Cloud services, including generative AI offerings. It is regularly updated to reflect current legal, ethical, and security requirements, making it the most reliable source for the most up-to-date information on prohibited use cases. The AUP explicitly covers restrictions on generating harmful content, engaging in illegal activities, and violating intellectual property rights, which directly apply to generative AI services.

Exam trap

This exam often tests the distinction between high-level ethical principles (AI Principles) and enforceable policy documents (Acceptable Use Policy), leading candidates to mistakenly choose the broader, aspirational document over the specific, binding one.

How to eliminate wrong answers

Option A is wrong because Google Cloud documentation provides general guidance on service features and best practices but does not serve as the definitive policy document for prohibited use; the AUP is the binding policy. Option B is wrong because Google's AI Principles are high-level ethical commitments that guide AI development and use, but they are not a specific, enforceable policy document detailing prohibited uses of services. Option D is wrong because the Gemini Terms of Service govern the use of the Gemini product specifically, not the broader set of Google Cloud generative AI services, and they do not replace the overarching Acceptable Use Policy.

613
MCQmedium

A marketing team uses a generative AI model to produce ad copy. They notice that when they set the temperature parameter to 0.1, the outputs are very similar across runs, but when they set it to 0.9, the outputs vary widely. They want a balance between creativity and consistency for a campaign that requires some variation but also brand alignment. Which temperature value should they choose?

A.0.5
B.0.0
C.2.0
D.1.0
AnswerA

A moderate temperature like 0.5 balances randomness and coherence. It allows the model to explore alternatives without drifting too far from likely outputs, which suits a campaign needing both variation and brand consistency. This setting is a common middle ground between deterministic and highly random generation.

Why this answer

A temperature of 0.5 provides a middle ground between deterministic and highly random outputs. It allows for some creativity and variation while keeping the text coherent and aligned with brand guidelines. Lower values like 0.0 reduce variation, while higher values like 1.0 or 2.0 increase randomness and risk off-brand content.

Exam trap

The trap here is assuming that higher temperature always yields better creativity without considering the trade-off in coherence and brand alignment.

614
MCQhard

A logistics company plans to use generative AI to draft responses to customer shipment inquiries. Legal requires a documented process for reviewing model outputs and handling harmful or inaccurate content before responses reach customers. Which practice should be built into the solution?

A.Fine-tune the model on past customer emails so it learns the company's tone and never produces harmful content.
B.Restrict the assistant to internal use only and skip any output review since customers never see it.
C.Deploy the model directly to customers and rely on users to report bad responses afterward.
D.Implement human-in-the-loop review with logging and content safety filters before responses are sent.
AnswerD

Human-in-the-loop review ensures a person validates outputs before they reach customers, while logging creates an auditable record and safety filters catch harmful content automatically. Together these satisfy the legal requirement for a documented review process. This layered approach also produces data for improving prompts and evaluating model behavior over time.

Why this answer

A documented review process for generative AI outputs typically combines automated safety filters, human review before customer delivery, and logging that preserves an audit trail. Reactive reporting, fine-tuning for tone, or avoiding review altogether do not provide the proactive controls and documentation legal demands. The layered human-in-the-loop design satisfies both safety and accountability requirements.

Exam trap

The trap here is assuming fine-tuning or post-hoc reporting substitutes for an explicit review and audit process, when governance requires proactive controls.

615
MCQeasy

A developer is using the Gemini API to generate code snippets. They notice the outputs often contain deprecated API calls. Which parameter adjustment or prompt strategy would most effectively encourage the model to use current APIs?

A.Add a system instruction specifying 'Use the most recent API version and avoid deprecated functions.'
B.Set top-p to 0.5 to reduce output diversity
C.Provide one few-shot example of a correct API call
D.Set temperature to 1.5 to increase creativity
AnswerA

A system instruction sets persistent behavioural guidance that conditions every response, so explicitly directing the model toward current API versions and away from deprecated functions steers generation more reliably than per-request wording. This directly targets the deprecated-call pattern observed in the outputs.

Why this answer

Adding a system instruction that explicitly directs the model to 'Use the most recent API version and avoid deprecated functions' directly influences the model's behavior at the prompt level. The Gemini API supports system instructions that act as persistent, high-level guidance, steering the model toward preferred output patterns—in this case, avoiding deprecated API calls. This is the most effective and direct method to enforce current API usage without altering sampling parameters or relying on limited examples.

Exam trap

This question tests the misconception that adjusting sampling parameters (like temperature or top-p) or providing a single example can reliably enforce content constraints, when in fact system instructions are the designed mechanism for persistent behavioral guidance in production-grade APIs.

How to eliminate wrong answers

Option B is wrong because setting top-p to 0.5 reduces the cumulative probability mass of token choices, which narrows output diversity but does not inherently bias the model toward current APIs; it may even suppress rare but correct modern API tokens. Option C is wrong because a single few-shot example provides only one instance of a correct API call, which is insufficient to override the model's training data bias toward deprecated APIs; the model may still default to older patterns. Option D is wrong because increasing temperature to 1.5 amplifies randomness and creativity, which can increase the likelihood of hallucinated or incorrect API calls, including deprecated ones, rather than encouraging adherence to current standards.

616
MCQhard

A company is required by the EU AI Act to ensure high-risk AI systems are transparent and auditable. They are using a proprietary model from a vendor. Which step is CRITICAL?

A.Implement custom safety filters on the model outputs
B.Ask the vendor to provide a Model Card and datasheets for the training data
C.Use a larger context window to capture all interactions
D.Train an internal model from scratch to replace the vendor model
AnswerB

Model Cards and training datasheets document intended use, evaluation results, limitations and data provenance, giving the deployer the evidence needed for transparency and audit obligations under the EU AI Act. Vendor contracts or accuracy benchmarks alone cannot demonstrate the system's design, risks and training data characteristics.

Why this answer

Under the EU AI Act, high-risk AI systems must be transparent and auditable, which requires documentation of the model's intended purpose, training data characteristics, performance metrics, and limitations. A vendor-provided Model Card and datasheets for the training data are the standard artifacts that satisfy this transparency and auditability obligation for a proprietary model. Without them, the deploying company cannot demonstrate compliance or perform meaningful risk assessment.

Exam trap

The trap is confusing output safety controls (filters) with regulatory transparency artifacts (Model Cards and datasheets), which are the actual audit evidence required.

How to eliminate wrong answers

Option A is wrong because custom safety filters on outputs address content moderation and harm prevention, not the transparency and auditability documentation required for high-risk systems. Option C is wrong because a larger context window is a technical capability that affects how much input the model can process; it has no bearing on regulatory transparency or audit trails. Option D is wrong because training an internal model from scratch is a costly, time-consuming alternative that does not by itself guarantee compliance and is not required when a vendor model can be documented via Model Cards and datasheets.

617
Multi-Selectmedium

A company is deploying a generative AI system for medical diagnosis support. To comply with Google's AI Principles and regulatory requirements, which TWO actions are essential? (Select 2)

Select 2 answers
A.Implement a human-in-the-loop review for all diagnostic suggestions
B.Publish a Model Card for the model
C.Ensure the system complies with GDPR and other privacy regulations for patient data
D.Use SynthID to watermark all output
E.Use a larger model to improve accuracy
AnswersA, C

Human-in-the-loop review keeps a qualified clinician accountable for every diagnostic suggestion, satisfying Google's AI Principles requirement that high-stakes medical AI remains subject to human oversight and cannot autonomously determine care. This directly addresses the regulatory constraint that generative outputs supporting diagnosis must be verified by a responsible professional before affecting patient treatment.

Why this answer

Option A is correct because Google's AI Principles require that high-stakes applications such as medical diagnosis keep humans in the loop, ensuring a qualified clinician reviews and approves every AI-generated diagnostic suggestion before it affects patient care. Option C is correct because processing patient data for diagnosis support triggers legal obligations under GDPR and comparable privacy regulations, including lawful basis, data minimization, and protection of special-category health data. Option B is not essential here: a Model Card is useful documentation for transparency, but it is not a mandatory action for regulatory compliance in this scenario.

Option D is not essential: SynthID watermarks AI-generated content to aid provenance, which is irrelevant to diagnostic decision support. Option E is not essential: using a larger model may improve accuracy but does not by itself satisfy AI Principles or regulatory requirements.

618
MCQhard

A healthcare organization is using Vertex AI to build a generative AI application that summarizes patient notes. They need to ensure that the model does not inadvertently generate or expose protected health information (PHI) in its outputs. Which Google Cloud feature should they implement to detect and redact sensitive data in the model's responses?

A.Vertex AI Model Monitoring
B.VPC Service Controls
C.Cloud Identity-Aware Proxy (IAP)
D.Cloud Data Loss Prevention (DLP) API
AnswerD

Cloud DLP API can inspect text for sensitive data like PHI and redact or mask it before it is stored or displayed. By integrating DLP into the application's output pipeline, the organization can automatically detect and redact PHI from model responses, ensuring compliance with regulations like HIPAA. This is the appropriate service for data loss prevention and sensitive data redaction.

Why this answer

Cloud DLP API is designed to discover, classify, and protect sensitive data such as PHI. By integrating it into the output pipeline of the generative AI application, the healthcare organization can automatically detect and redact PHI from model responses, ensuring compliance and preventing data leaks.

Exam trap

The trap here is confusing access control or monitoring services with data inspection and redaction; only DLP is built to detect and redact sensitive data in text.

619
MCQhard

A financial analytics firm wants to prototype prompts against several Gemini model versions quickly, compare outputs side by side, and then export the winning prompt configuration into a production application with enterprise controls. Which combination of Google Cloud offerings best supports this workflow?

A.Vertex AI for prototyping, then Google AI Studio for production
B.Cloud Natural Language API for prototyping, then Vertex AI for production
C.Google AI Studio for prototyping, then Vertex AI for production deployment
D.Gemini for Google Workspace for both prototyping and production
AnswerC

Google AI Studio is built for fast prompt experimentation with Gemini models, letting the firm compare outputs and iterate without infrastructure overhead. Once a prompt configuration is proven, Vertex AI provides the enterprise controls, IAM, logging, and quotas needed for production. Together they cover the full path from prototype to governed deployment described in the scenario.

Why this answer

Prototyping prompts against multiple Gemini models is fastest in Google AI Studio, and moving the validated configuration into a governed environment is what Vertex AI provides. The productivity suite and the language analysis API cannot perform side-by-side generative prompt comparison or serve as the production deployment target.

Exam trap

The trap here is assuming the prototyping tool and the production platform are interchangeable, when only one is designed for governed application deployment.

620
MCQeasy

Which Google Cloud generative AI model is specifically designed for code generation tasks?

A.Gemini
B.PaLM 2
C.Imagen
D.Codey
AnswerD

Codey is Google Cloud's model family fine-tuned specifically for code generation, completion and chat about code, so it directly matches the stem's requirement for a code-focused generative AI model rather than a general-purpose text model.

Why this answer

Codey is Google's model fine-tuned for code generation, completion, and chat. PaLM 2 is a general purpose LLM, Gemini is multimodal, and Imagen is for image generation.

621
MCQhard

A financial services company wants to build a generative AI application that can answer questions based on their internal documents, which are constantly updated. They need the model to cite sources and avoid hallucination. Which Google Cloud service should they use?

A.Vertex AI Model Garden
B.Vertex AI Pipelines
C.Vertex AI Feature Store
D.Vertex AI Search
AnswerD

Vertex AI Search is designed for enterprise search and retrieval-augmented generation (RAG) over proprietary documents. It indexes content, supports frequent updates, and provides grounded responses with citations. This directly addresses the need for accurate, source-cited answers from internal documents.

Why this answer

Vertex AI Search provides enterprise-grade search and RAG capabilities, enabling applications to retrieve relevant passages from internal documents and generate answers with citations. It handles dynamic document updates, reducing hallucination by grounding responses in the source material.

Exam trap

The trap here is confusing Vertex AI Search with Vertex AI Model Garden, but only Vertex AI Search offers built-in document indexing and citation support.

622
MCQmedium

A logistics company built a Gemini-powered assistant that answers driver questions about routes and hours-of-service rules. The assistant performs well on common questions but produces fabricated regulatory citations when asked about rare edge cases. The team has a curated set of correct answers for these edge cases and wants the model to adopt that behavior reliably. Which approach best fits?

A.Apply supervised fine-tuning on the curated edge-case examples to teach the desired response pattern.
B.Lower the topP value so the model only considers the most likely tokens and avoids fabricating citations.
C.Add the curated answers to the system instruction and rely on the model to generalize.
D.Increase the model's context window by switching to a long-context variant and pasting the full regulations.
AnswerA

Supervised fine-tuning is designed to teach a model a specific input-to-output behavior using labeled examples. With a curated set of correct edge-case answers, tuning adjusts the model so it reproduces the desired citation style and content, which is more reliable than describing the behavior in a prompt for rare cases.

Why this answer

Supervised fine-tuning uses the curated input-output pairs to adjust the model's weights so it reliably reproduces correct edge-case answers, which is the intended use of labeled examples. Prompt stuffing, sampling controls, and larger context windows do not durably change behavior for rare inputs.

Exam trap

The trap here is believing that narrowing sampling with topP or enlarging the context window will fix factual errors, when those knobs do not teach the model new domain answers.

623
MCQeasy

According to Google's AI Principles, which of the following is a key commitment regarding privacy?

A.AI systems should incorporate privacy design principles, including notice and consent
B.AI systems should share all data with third parties for transparency
C.AI systems should collect as much data as possible to improve accuracy
D.AI systems should store data indefinitely for future analysis
AnswerA

Google's AI Principles commit to privacy by design, embedding protections such as notice and consent into systems from the outset. This satisfies the question's focus on privacy commitments rather than fairness, safety or accountability principles.

Why this answer

Google's AI Principles explicitly commit to incorporating privacy design principles, such as notice and consent, into AI systems. This means that AI systems should be designed with privacy safeguards from the outset, ensuring users are informed about data collection and have control over their data, aligning with frameworks like GDPR and privacy-by-design.

Exam trap

Google often tests the misconception that transparency requires full data sharing or that maximizing data collection is always beneficial for AI accuracy, when in fact privacy principles mandate data minimization and user consent.

How to eliminate wrong answers

Option B is wrong because sharing all data with third parties for transparency violates core privacy commitments; transparency does not require indiscriminate data sharing, and Google's principles emphasize data minimization and user consent. Option C is wrong because collecting as much data as possible to improve accuracy contradicts the principle of data minimization and could lead to privacy violations; accuracy should be balanced with privacy, not achieved at any cost. Option D is wrong because storing data indefinitely for future analysis violates the principle of data retention limits and user control; AI systems should only retain data as long as necessary for specified purposes, with mechanisms for deletion.

624
MCQmedium

A developer wants to add real-time speech transcription to a customer call center application. They need low latency and high accuracy for multiple languages. Which Google AI API is most appropriate?

A.Speech-to-Text API
B.Natural Language API
C.Text-to-Speech API
D.Translation API
AnswerA

Google Cloud Speech-to-Text API delivers streaming recognition with low latency, supporting real-time transcription of live audio. It covers over 125 languages and variants with high accuracy, satisfying the call centre's multilingual requirement. Its synchronous streaming mode returns interim results as speech occurs, which batch alternatives cannot match.

Why this answer

The Speech-to-Text API is the correct choice because it is specifically designed to convert audio into text in real time, supporting over 125 languages and variants with low-latency streaming. It offers features like automatic punctuation, speaker diarization, and domain-specific models (e.g., phone call) that directly meet the requirements of a customer call center application needing high accuracy across multiple languages.

Exam trap

The Generative AI Leader exam often tests the distinction between APIs that process text (Natural Language, Translation) versus those that process audio (Speech-to-Text, Text-to-Speech), and the trap here is confusing the direction of conversion (speech-to-text vs. text-to-speech) or assuming a translation API can handle raw audio input.

How to eliminate wrong answers

Option B is wrong because the Natural Language API analyzes text for entities, sentiment, and syntax, but it does not process audio or perform speech transcription. Option C is wrong because the Text-to-Speech API converts text into spoken audio, which is the opposite direction of the required speech-to-text functionality. Option D is wrong because the Translation API translates text between languages but cannot transcribe speech from audio input.

625
Multi-Selectmedium

Which THREE steps are required to secure a generative AI pipeline that uses Vertex AI and involves sensitive customer data?

Select 3 answers
A.Use VPC Service Controls to create a perimeter around Vertex AI resources
B.Apply IAM roles with least privilege and use service accounts for the pipeline
C.Expose the prediction endpoint publicly with an API key
D.Enable data encryption at rest using Cloud KMS
E.Disable audit logging to reduce data exposure
AnswersA, B, D

VPC Service Controls establish a service perimeter that blocks data exfiltration from Vertex AI endpoints, directly satisfying the requirement to protect sensitive customer data during pipeline operations. This network-level containment prevents unauthorised projects or identities from reaching the model and its training data, even if IAM permissions are misconfigured.

Why this answer

Option A is correct because VPC Service Controls lets you define a service perimeter around Vertex AI resources, preventing data exfiltration of sensitive customer data even if credentials are compromised. Option B is correct because applying IAM roles with least privilege and using dedicated service accounts for the pipeline enforces fine-grained access control and limits the blast radius of any compromised identity. Option D is correct because enabling data encryption at rest with Cloud KMS (customer-managed encryption keys) protects sensitive customer data stored in Vertex AI datasets, models, and related storage from unauthorized access at the storage layer.

Option C is incorrect because exposing the prediction endpoint publicly with only an API key removes network-level protections and is not a recommended security control for sensitive data. Option E is incorrect because disabling audit logging reduces visibility and accountability, which weakens security and compliance rather than strengthening them.

Exam trap

The trap here is that candidates may confuse API key authentication (Option C) as a valid security measure, but for sensitive data, API keys lack identity binding and are considered a weak secret, whereas VPC Service Controls and IAM provide defense-in-depth.

626
Multi-Selectmedium

A media company is using Vertex AI to build a generative AI application that creates personalized news summaries. They want to ensure the model's outputs are factually grounded in their curated article database and that the application can scale to thousands of concurrent users. Which two Google Cloud services should they use to achieve these goals? (Choose two.)

Select 2 answers
A.Vertex AI Search
B.Cloud Dataflow
C.Vertex AI Endpoints
D.Cloud CDN
E.Vertex AI Pipelines
AnswersA, C

Vertex AI Search provides managed retrieval over the curated article database, enabling the generative model to ground its summaries in factual content. It handles indexing, semantic search, and integration with Gemini, which directly supports the requirement for factual grounding and reduces hallucinations.

Why this answer

Vertex AI Search grounds the generative summaries in the curated article database, ensuring factual accuracy. Vertex AI Endpoints provides the scalable online serving infrastructure needed to handle thousands of concurrent users. Together, they address both grounding and scalability, while the other services focus on data processing, content delivery, or pipeline orchestration.

Exam trap

The trap here is confusing data processing or orchestration services like Cloud Dataflow or Vertex AI Pipelines with the retrieval and serving components required for a grounded, scalable generative AI application.

627
Multi-Selectmedium

A machine learning team is using Vertex AI to train a custom model. They want to optimize hyperparameters automatically. Which TWO steps are necessary to set up hyperparameter tuning in Vertex AI? (Choose TWO)

Select 2 answers
A.Enable Vertex AI Experiments
B.Use a custom container with a GPU
C.Enable distributed training across multiple nodes
D.Define a hyperparameter metric in the training code
E.Create a HyperparameterTuningJob with parameter specifications
AnswersD, E

Hyperparameter tuning needs a measurable objective, so the training code must report the metric to optimise, such as validation accuracy, via the defined hyperparameter metric. Without it, Vertex AI cannot evaluate trials or compare configurations, so this step is required alongside the tuning job.

Why this answer

Option D is correct because Vertex AI hyperparameter tuning requires the training application to report a target metric (for example, by writing it to the Vertex AI summary or using the hypertune package) that the tuning service can optimize; without a defined hyperparameter metric, the service has no objective to maximize or minimize. Option E is correct because you must create a HyperparameterTuningJob resource that specifies the search algorithm, the parameters with their ranges or discrete values, the metric to optimize, and the training job to run. Options A, B, and C are not required: Vertex AI Experiments is for tracking and comparing runs, not for enabling tuning; a custom container with a GPU is only one possible training setup and GPUs are not mandatory for tuning; and distributed multi-node training is unrelated to the hyperparameter tuning setup itself.

Exam trap

The trap here is confusing hyperparameter tuning with other Vertex AI features like Experiments or distributed training, leading candidates to select unnecessary options. The exam often tests whether you know the minimal required components: a tuning job and a reported metric.

628
MCQhard

Refer to the exhibit. A developer sees this error when trying to deploy a model from Vertex AI Model Registry. What is the most likely cause?

A.The region is not supported
B.The developer used the model display name instead of the full resource name
C.The model is not published
D.The model is in a different project
AnswerB

Deployment APIs require the fully qualified resource name, formatted as projects/{project}/locations/{location}/models/{model}. Supplying only the display name fails resource resolution, producing the error, since display names are not unique identifiers within the Vertex AI Model Registry.

Why this answer

The error occurs because Vertex AI Model Registry requires the full resource name (e.g., 'projects/{project}/locations/{region}/models/{model_id}') to deploy a model, not just the display name. The display name is a human-readable label that is not unique within a project, while the full resource name uniquely identifies the model version. Using the display name causes the API to fail with a 'not found' or 'invalid argument' error.

Exam trap

Google Cloud often tests the distinction between display names (non-unique, human-readable) and resource names (unique, API-required) in cloud services like Vertex AI, where candidates mistakenly assume display names can be used interchangeably with resource identifiers.

How to eliminate wrong answers

Option A is wrong because Vertex AI supports model deployment in all regions where the service is available, and the error message does not indicate a regional restriction. Option C is wrong because a model can be deployed from the registry even if it is not published to the public; publishing is only required for sharing with external users or making it available in the Model Garden. Option D is wrong because the error would reference a cross-project permission issue (e.g., 'permission denied' or 'resource not found in project'), not a display name mismatch.

629
MCQmedium

A financial analyst is using a large language model to generate executive summaries from lengthy earnings call transcripts. The summaries often miss key financial figures and include irrelevant details. Which technique should be used to improve the relevance and accuracy of the summaries?

A.Use few-shot prompting with examples of well-structured summaries.
B.Increase the maximum output token limit.
C.Fine-tune the model on a large corpus of general text.
D.Increase the model's temperature setting.
AnswerA

Few-shot prompting provides the model with examples of desired output, guiding it to focus on relevant information and format. By showing examples that highlight key financial figures and omit irrelevant details, the model can learn to replicate that behavior. This technique is effective for improving relevance and accuracy in summarization tasks.

Why this answer

Few-shot prompting is a powerful technique to guide the model's output by providing examples. In this scenario, showing the model examples of summaries that include key financial figures and exclude irrelevant details helps it learn the desired pattern, thereby improving relevance and accuracy.

Exam trap

The trap here is assuming that increasing token limits or temperature will improve summarization, when actually they can degrade quality.

630
MCQhard

A company deploys a GenAI-powered code review assistant. During evaluation, they find that the assistant often suggests security vulnerabilities as improvements. What is the MOST likely cause?

A.The model was trained on a dataset with many insecure code examples
B.The model's temperature is set too low
C.The model is too small for code generation tasks
D.The prompt does not include a security constraint
AnswerA

Training data containing insecure code patterns teaches the model that such patterns are acceptable improvements, so it reproduces them during review. This directly satisfies the stem's constraint: the assistant recommends vulnerabilities because its learned distribution reflects the insecure examples it was trained on, rather than secure coding practise.

Why this answer

The most likely cause is that the model was trained on a dataset containing many insecure code examples. A GenAI code review assistant learns patterns from its training data; if that data includes prevalent security vulnerabilities (e.g., SQL injection, buffer overflows), the model will internalize those patterns as 'normal' or even 'desirable' improvements. This leads to the assistant suggesting insecure code changes because it is statistically replicating the flawed logic it was exposed to during training.

Exam trap

A common misconception tested in this context is that prompt engineering alone (e.g., adding a security constraint) can override fundamental training data biases, when in fact the model's learned weights from the training corpus are the dominant factor in output quality.

How to eliminate wrong answers

Option B is wrong because setting the temperature too low (e.g., near 0) makes the model more deterministic and conservative, reducing randomness and the likelihood of suggesting unusual or insecure patterns; it would not cause the model to actively suggest vulnerabilities. Option C is wrong because model size (number of parameters) affects capability and fluency, not the tendency to generate insecure code; a small model can still produce secure suggestions if trained on secure data, while a large model trained on insecure data will replicate those flaws. Option D is wrong because while a missing security constraint in the prompt might fail to guide the model away from vulnerabilities, the root cause is the training data; even with a security constraint, a model trained on insecure examples may still suggest vulnerabilities due to its ingrained patterns, and the question asks for the 'most likely' cause, which is the data quality issue.

631
Multi-Selecthard

A product team is designing a generative AI application to create personalized email campaigns. They want to ensure the model produces high-quality, relevant content while minimizing risks such as bias and harmful outputs. Which two practices should the team implement? (Choose two.)

Select 2 answers
A.Fine-tune the model on a small dataset of previously successful email campaigns without any validation.
B.Implement content filters and safety thresholds to block harmful or inappropriate outputs.
C.Conduct regular fairness evaluations of the model's outputs across different demographic groups.
D.Disable all logging and monitoring to protect user privacy and reduce overhead.
E.Set the model's temperature to the maximum value to encourage diverse and creative emails.
AnswersB, C

Content filters and safety thresholds automatically detect and suppress toxic, biased, or otherwise harmful text before it reaches recipients. This is a critical layer of defense in generative AI applications, especially for customer-facing campaigns. It helps maintain brand safety and compliance with regulations, directly addressing the goal of minimizing harmful outputs.

Why this answer

Regular fairness evaluations help detect and reduce bias across demographic groups, ensuring equitable and appropriate content. Content filters and safety thresholds act as a safeguard to block harmful or inappropriate outputs before they reach users. Together, these practices address both bias and harmful content, aligning with responsible AI development.

Other options either increase randomness, risk overfitting, or remove necessary oversight.

Exam trap

The trap here is focusing only on creative output and overlooking the need for systematic bias detection and automated safety filters.

632
MCQhard

A financial services firm wants to deploy a generative AI assistant that summarizes earnings call transcripts for its analysts. The firm's risk committee requires that the assistant never produce investment recommendations and that all outputs be traceable to source transcript passages. Which design choice most directly enforces both constraints?

A.Store all prompts and responses in BigQuery and run a nightly batch job to flag outputs that look like recommendations.
B.Deploy the assistant with a content filter that blocks any output containing financial terminology.
C.Use retrieval-augmented generation over the transcripts and apply a system instruction that prohibits recommendations, returning cited source chunks.
D.Fine-tune a Gemini model on historical earnings transcripts and deploy it with a low temperature setting.
AnswerC

RAG grounds every response in retrieved transcript passages and can return citations to those passages, satisfying traceability. A system instruction that explicitly forbids investment recommendations constrains the model's behavior at generation time. Together they directly address both the no-recommendation policy and the requirement that outputs map back to source text.

Why this answer

Grounding responses in retrieved transcript passages with RAG gives analysts verifiable citations, while a system instruction establishes a hard behavioral boundary against investment recommendations. This combination enforces both the provenance and the policy constraint at generation time, rather than relying on post-processing or broad content filters that cannot distinguish summary from advice.

Exam trap

The trap here is treating traceability and policy enforcement as model-tuning problems, when they are better solved through retrieval grounding and explicit generative instructions.

633
MCQeasy

A healthcare startup is developing a generative AI system to assist doctors in diagnosing rare diseases. According to Google's AI Principles, what is the MOST important requirement before deployment?

A.The model must achieve at least 99% accuracy on a held-out test set
B.The model must be trained on the most recent medical literature
C.The startup must publish the model's architecture in a peer-reviewed journal
D.The system must include a mechanism for human review of all diagnostic suggestions
AnswerD

Google's AI Principles require human oversight for high-stakes domains such as healthcare diagnosis. Retaining a clinician to review every diagnostic suggestion satisfies this, since the system only assists rather than determines care, keeping accountability with a qualified professional.

Why this answer

Google's AI Principles emphasize that AI systems should be socially beneficial and avoid creating or reinforcing unfair bias, and for high-stakes domains like healthcare, human oversight is critical. The principle of 'be accountable to people' requires that AI systems include mechanisms for human review and feedback, especially for diagnostic suggestions that could affect patient safety.

Exam trap

The trap is assuming that a quantitative metric like 99% accuracy is the most important requirement, when Google's AI Principles prioritize human oversight and accountability over raw performance metrics in high-stakes applications.

How to eliminate wrong answers

Option A is wrong because a specific accuracy threshold (99%) is not a requirement in Google's AI Principles; the focus is on overall safety, fairness, and human oversight, not a single metric. Option B is wrong because training on recent literature is a good practice but not a stated principle requirement; the principles focus on broader ethical deployment. Option C is wrong because publishing architecture in a peer-reviewed journal is not mandated by Google's AI Principles; transparency is encouraged but not in that specific form.

634
MCQeasy

A retail company wants to use a generative AI model to create unique product descriptions for thousands of items. They need the model to produce varied, human-like text without being explicitly programmed for each product. Which core capability of generative AI does this scenario primarily rely on?

A.Forecasting future sales trends
B.Clustering similar products based on features
C.Generating novel content based on patterns learned from training data
D.Classifying products into predefined categories
AnswerC

Generative AI models learn statistical patterns from large datasets and can produce new, original content such as text, images, or code. In this scenario, the model generates unique product descriptions by leveraging its training on diverse text, without explicit programming for each item. This is the fundamental capability that enables creative and varied output.

Why this answer

Generative AI is designed to create new content by learning patterns from existing data. In this scenario, the model generates unique product descriptions without explicit programming, which is the essence of generative AI. The other options describe discriminative or predictive tasks that do not produce novel text, so they fail to meet the company's need for varied, human-like descriptions.

Exam trap

The trap here is confusing generative AI with traditional discriminative models that only classify or predict, rather than create new content.

635
MCQeasy

A company is using Vertex AI to deploy a text generation model for a chatbot. They want to reduce the response latency. Which configuration change is most effective?

A.Enable model quantization
B.Use a smaller model variant
C.Increase the number of GPUs
D.Use a larger batch size
AnswerB

A smaller model variant reduces the number of parameters processed per token, cutting inference compute and therefore response latency. This directly satisfies the stem's constraint of lowering latency for the Vertex AI chatbot, since generation time scales with model size rather than with prompt formatting or endpoint region.

Why this answer

Using a smaller model variant directly reduces the number of parameters and computational operations required per inference, which lowers latency. In Vertex AI, smaller models like `text-bison@002` have fewer layers and attention heads than larger counterparts, resulting in faster token generation without requiring hardware changes.

Exam trap

Google Cloud often tests the misconception that increasing compute resources (GPUs) or batch size always reduces latency, when in fact these optimizations target throughput, not per-request response time.

How to eliminate wrong answers

Option A is wrong because model quantization (e.g., reducing weights from FP32 to INT8) can reduce memory footprint and improve throughput, but it does not guarantee lower latency per request and may introduce accuracy trade-offs; it is not the most effective single change for latency reduction. Option C is wrong because increasing the number of GPUs can improve throughput for batch processing but does not reduce per-request latency; in fact, it may increase communication overhead and cost without speeding up individual inference. Option D is wrong because using a larger batch size increases throughput for concurrent requests but actually increases the latency for each individual request, as the model processes more sequences together before returning results.

636
MCQmedium

A company is using Vertex AI Agent Builder to create a travel booking agent. They want the agent to book flights and hotels dynamically. What action type should they use?

A.Dynamic call
B.Static call
C.Webhook
D.Notification
AnswerC

Webhook actions let the agent call external APIs to perform dynamic transactions such as booking flights and hotels, returning live results. Static responses or simple text replies cannot execute the real-world booking the scenario requires.

Why this answer

Vertex AI Agent Builder uses webhooks to integrate with external systems for dynamic, real-time operations like booking flights and hotels. A webhook allows the agent to make HTTP calls to external APIs (e.g., a travel booking service) to fetch or update data during a conversation, enabling dynamic booking actions. Static or notification actions cannot handle the two-way, real-time data exchange required for live reservations.

Exam trap

The trap here is that candidates confuse 'dynamic call' (a generic term) with the actual Vertex AI Agent Builder mechanism, or assume 'notification' can handle bidirectional data exchange, when only webhooks provide the required synchronous HTTP callback for real-time operations.

How to eliminate wrong answers

Option A is wrong because 'Dynamic call' is not a recognized action type in Vertex AI Agent Builder; the platform uses webhooks for dynamic interactions, not a separate 'dynamic call' concept. Option B is wrong because 'Static call' refers to predefined, non-interactive responses or data lookups that cannot handle real-time booking logic or external API calls. Option D is wrong because 'Notification' is a one-way push mechanism (e.g., sending alerts) and does not support the request-response pattern needed to execute a booking transaction.

637
MCQhard

A researcher wants to use Google's AlphaFold for a project. What is the primary capability of AlphaFold?

A.Generating realistic human speech
B.Playing the game of Go at superhuman level
C.Predicting 3D protein structures from amino acid sequences
D.Generating code from natural language descriptions
AnswerC

AlphaFold predicts a protein's three-dimensional structure directly from its amino acid sequence, solving the longstanding folding problem. This satisfies the researcher's requirement by outputting atomic-level structural models, a capability distinct from sequence alignment or molecular docking tools.

Why this answer

AlphaFold, developed by Google DeepMind, is specifically designed to predict the 3D structure of proteins from their amino acid sequences. This capability solves a fundamental challenge in biology, as the function of a protein is largely determined by its 3D shape, and experimental methods like X-ray crystallography are time-consuming and expensive. AlphaFold achieves this using a deep learning architecture that integrates multiple sequence alignment (MSA) and pairwise distance predictions to model the spatial coordinates of atoms.

Exam trap

Google often tests the distinction between its own DeepMind projects (AlphaGo vs. AlphaFold vs. AlphaZero), so the trap here is confusing the domain of game-playing AI with the domain of scientific prediction, leading candidates to pick Option B if they recall AlphaGo's fame but not AlphaFold's specific purpose.

How to eliminate wrong answers

Option A is wrong because generating realistic human speech is the primary capability of text-to-speech models like WaveNet or Tacotron, not AlphaFold, which focuses on structural biology. Option B is wrong because playing the game of Go at superhuman level is the achievement of AlphaGo and AlphaZero, which use reinforcement learning and Monte Carlo tree search, not the protein folding task AlphaFold was built for. Option D is wrong because generating code from natural language descriptions is the domain of large language models like Codex or GPT-4, not AlphaFold, which has no code generation functionality.

638
MCQeasy

A data analyst needs to run a simple regression model directly on data stored in BigQuery without moving data to another platform. Which service should they use?

A.TensorFlow on Compute Engine
B.BigQuery ML
C.Vertex AI Training
D.Google Colab
AnswerB

BigQuery ML executes regression directly inside BigQuery using SQL, so the data never leaves the platform — satisfying the no-movement constraint. Training and prediction run server-side against the stored table, unlike exporting to Vertex AI or a notebook, which would require extracting data first.

Why this answer

BigQuery ML (B) is correct because it allows users to create and execute machine learning models using standard SQL syntax directly on data stored in BigQuery, without needing to export data to a separate platform. This service is specifically designed for running regression, classification, and other models natively within BigQuery, leveraging its serverless architecture and built-in ML capabilities.

Exam trap

The trap here is that candidates often confuse Vertex AI Training (a full-featured ML platform) with BigQuery ML, not realizing that Vertex AI requires data export and more setup, while BigQuery ML is purpose-built for in-database modeling with minimal overhead.

How to eliminate wrong answers

Option A is wrong because TensorFlow on Compute Engine requires moving data out of BigQuery to a virtual machine, where you must manually manage infrastructure, install dependencies, and write custom training code, which contradicts the requirement of not moving data. Option C is wrong because Vertex AI Training is a managed ML platform that typically requires exporting data from BigQuery to Cloud Storage or a dataset in Vertex AI, and it involves more complex pipeline setup than a simple regression model. Option D is wrong because Google Colab is a Jupyter notebook environment that runs in the cloud but requires data to be loaded from BigQuery into a DataFrame, moving it out of BigQuery's native storage, and it does not provide a direct SQL-based modeling interface.

639
MCQeasy

In the transformer architecture, what is the role of the attention mechanism?

A.It normalizes the output of each layer
B.It decides which parts of the input to focus on when generating each token
C.It predicts the next token directly
D.It converts tokens into numerical vectors
AnswerB

Attention computes weighted relevance between every token pair, letting the model weigh distant context rather than relying on sequential recurrence. For each generated token, query–key dot products produce weights over value vectors, so the decoder dynamically prioritises the input positions most relevant to that specific output step.

Why this answer

The attention mechanism in the Transformer architecture computes a weighted sum of all input token representations, allowing the model to dynamically focus on the most relevant parts of the input sequence when generating each output token. This is achieved through learned query, key, and value projections that produce attention scores, enabling the model to capture long-range dependencies and contextual relationships. Option B correctly identifies this core function of selectively attending to input elements during token generation.

Exam trap

Candidates often mistake the attention mechanism's role in focusing on input parts with the final prediction layer's role in outputting the next token, leading them to select Option C.

How to eliminate wrong answers

Option A is wrong because normalization of layer outputs is performed by layer normalization, not the attention mechanism; attention computes relevance weights, not normalization statistics. Option C is wrong because predicting the next token directly is the role of the final linear layer and softmax over the vocabulary, while the attention mechanism provides contextualized representations that feed into that prediction. Option D is wrong because converting tokens into numerical vectors is the function of the embedding layer (token embeddings), not the attention mechanism, which operates on those vectors to compute attention scores.

640
MCQmedium

A data scientist observes that a text generation model consistently produces outputs that stereotype certain genders. According to Google's AI Principles, what is the BEST first step?

A.Evaluate the model's bias using a diverse test set across genders
B.Immediately stop using the model and delete it
C.Fine-tune the model on a gender-balanced dataset
D.Add a disclaimer that the model may exhibit bias
AnswerA

Measuring bias with a diverse test set across genders establishes evidence of the stereotyping before any remediation. This satisfies the principle of avoiding unfair bias by first quantifying disparate outputs, so subsequent mitigation targets the actual observed skew rather than assumptions.

Why this answer

Google's AI Principles emphasize that the first step in addressing bias is to evaluate and measure it using appropriate tools and diverse datasets. This aligns with Principle #2: 'Avoid creating or reinforcing unfair bias,' which requires testing models across relevant demographic groups before taking corrective action. Without evaluation, any subsequent mitigation steps would lack a baseline and could be ineffective or counterproductive.

Exam trap

The Generative AI Leader exam often tests the misconception that mitigation (like fine-tuning or disclaimers) should be the immediate response, rather than the correct first step of systematic evaluation and measurement of bias.

How to eliminate wrong answers

Option B is wrong because immediately stopping use and deleting the model is an overreaction that violates the principle of 'Be socially beneficial' — the model may still provide value if bias is addressed, and deletion prevents any learning from the bias. Option C is wrong because fine-tuning on a gender-balanced dataset is a mitigation step that should only be taken after evaluation to understand the specific nature and extent of the bias; premature fine-tuning could introduce new biases or fail to address root causes. Option D is wrong because adding a disclaimer is a transparency measure, not a first step — it acknowledges bias without measuring or understanding it, which violates the principle of 'Be accountable to people' by avoiding proactive bias detection.

641
MCQeasy

Which Google initiative provides a set of interactive, open-source tools to help UX designers and product managers build human-centered AI products?

A.TensorFlow Privacy
B.Model Cards
C.Datasheets for Datasets
D.People + AI Guidebook (PAIR)
AnswerD

PAIR's People + AI Guidebook supplies interactive, open-source design tools and patterns specifically for UX designers and product managers building human-centred AI, matching the stem's audience and resource type. It is Google's dedicated human-centred AI design resource.

Why this answer

The People + AI Guidebook (PAIR) is Google's initiative that offers a collection of interactive, open-source tools, methods, and case studies specifically designed to help UX designers and product managers create human-centered AI products. It focuses on the human side of AI, providing frameworks for user research, prototyping, and evaluation to ensure AI systems are intuitive, fair, and useful. Unlike technical libraries, PAIR addresses the design and product management challenges unique to AI, making it the correct answer.

Exam trap

The trap here is confusing technical AI ethics/documentation tools (like Model Cards or Datasheets) with human-centered design resources, so candidates may pick an option that sounds related to responsible AI but is not aimed at UX designers and product managers.

How to eliminate wrong answers

Option A is wrong because TensorFlow Privacy is a technical library for training machine learning models with differential privacy, aimed at developers and data scientists, not UX designers or product managers. Option B is wrong because Model Cards are standardized documentation artifacts that describe a model's performance, limitations, and ethical considerations, but they are not an interactive set of tools for building products. Option C is wrong because Datasheets for Datasets is a documentation framework for dataset provenance and composition, not a toolkit for human-centered AI design.

642
Multi-Selecteasy

Which TWO Google Cloud services can be used together to implement a RAG (retrieval-augmented generation) pipeline? (Select 2)

Select 2 answers
A.Cloud SQL
B.Vertex AI Vector Search
C.Bigtable
D.Vertex AI PaLM API
E.Cloud Functions
AnswersB, D

Provides vector similarity search for retrieval.

Why this answer

Vertex AI Vector Search (option B) is correct because it provides a managed vector database for storing and querying embeddings, which is essential for the retrieval step in a RAG pipeline. It enables semantic similarity search over large datasets, allowing the system to fetch relevant context documents based on a user query.

Exam trap

Google Cloud often tests the misconception that any database (like Cloud SQL or Bigtable) can serve as a vector store for RAG, but they lack native vector indexing and similarity search, making them unsuitable for efficient retrieval at scale.

643
MCQmedium

A company wants to embed a generative AI writing assistant into their Google Docs workflow. The assistant should help users draft emails and reports based on prompts. Which Google Workspace feature should they leverage?

A.Google Workspace Add-ons with Vertex AI
B.Vertex AI API integrated via Apps Script
C.Gmail Smart Compose
D.Duet AI in Google Docs (Gemini for Google Workspace)
AnswerD

Gemini for Google Workspace, previously Duet AI, embeds generative assistance directly inside Google Docs, letting users draft emails and reports from prompts without leaving the document. This satisfies the requirement to integrate a writing assistant into the existing Docs workflow.

Why this answer

Duet AI in Docs (now Gemini for Google Workspace) provides 'Help me write' functionality for drafting content. Apps Script can be used for custom add-ons but requires development. Vertex AI API is for external integration.

Smart Compose is for Gmail only.

644
Multi-Selecteasy

Which TWO are components of the Vertex AI Generative AI Studio?

Select 2 answers
A.Dataflow
B.Model Garden
C.Pipeline templates
D.Cloud Functions
E.Prompt Editor
AnswersB, E

Model Garden is a catalogue within Vertex AI Generative AI Studio for discovering, testing and deploying foundation models, including Google's own and third-party options. It is one of the studio's core components, satisfying the question's requirement.

Why this answer

Model Garden (B) is a core component of Vertex AI Generative AI Studio, providing a curated catalog of foundation models from Google and third parties (e.g., PaLM, Gemini, Llama) that users can browse, test, and deploy. Prompt Editor (E) is also a built-in Studio component that lets users interactively design, test, and refine prompts against foundation models without writing code. Dataflow (A) is a managed Apache Beam service for batch and streaming data pipelines, unrelated to the Studio's model/prompt tooling.

Pipeline templates (C) belong to Vertex AI Pipelines for orchestrating ML workflows, not the Generative AI Studio interface. Cloud Functions (D) is a serverless compute service for event-driven code, not a Generative AI Studio component.

Exam trap

Google Cloud often tests the distinction between core generative AI studio components (like Model Garden and Prompt Editor) and broader GCP services (like Dataflow or Cloud Functions) that are not part of the studio, leading candidates to select familiar but incorrect options.

645
MCQeasy

What is the primary purpose of Google's Content Safety filters in Vertex AI?

A.To filter out low-quality training data
B.To ensure the model only generates content from a curated set of sources
C.To block generated content that contains hate speech, violence, or sexually explicit material
D.To improve the model's accuracy on safe content
AnswerC

Vertex AI's Content Safety filters apply configurable thresholds across harm categories — hate speech, violence, sexually explicit material and harassment — to block or flag both prompts and generated responses. This directly satisfies the stem's focus on safety enforcement, since filtering occurs at the API layer before content reaches users, independent of model choice or tuning.

Why this answer

Google's Content Safety filters in Vertex AI are designed to block generated content that violates safety policies, specifically targeting hate speech, violence, and sexually explicit material. This is a core component of responsible AI deployment, ensuring that model outputs adhere to ethical guidelines and legal requirements. The filters operate by analyzing the generated text or images against predefined safety categories, not by assessing data quality or source curation.

Exam trap

The trap here is that candidates may confuse Content Safety filters with data quality filters or source restrictions, assuming they improve model accuracy or curate training data, when in fact they are purely safety mechanisms applied at inference time.

How to eliminate wrong answers

Option A is wrong because Content Safety filters are not used to filter out low-quality training data; that function is handled by data preprocessing and curation pipelines, not by inference-time safety filters. Option B is wrong because Content Safety filters do not restrict the model to a curated set of sources; they block specific types of harmful content regardless of source, and the model can still generate from its full training distribution. Option D is wrong because the primary purpose is not to improve accuracy on safe content but to prevent the generation of unsafe content; accuracy improvements are a separate concern addressed by model tuning and evaluation.

646
Multi-Selectmedium

A tech company wants to ensure that their generative AI model does not produce harmful content. They plan to use Google Cloud's content safety features. Which two methods can they use to customize content safety? (Choose two.)

Select 2 answers
A.Use the default safety filters without any modifications
B.Disable all safety filters for maximum model creativity
C.Adjust safety thresholds for different categories like hate speech and violence
D.Define a custom blocklist of prohibited words or phrases
E.Train a separate model to detect harmful content
AnswersC, D

Per-category safety thresholds let the company tune sensitivity independently for hate speech, violence and other harm types, raising or lowering blocking strictness to match its risk appetite. This satisfies the stem's customisation requirement without retraining the underlying model.

Why this answer

Option C is correct because Google Cloud's generative AI safety features (such as those in Vertex AI) let you configure per-category safety thresholds, so you can raise or lower sensitivity for categories like hate speech, harassment, sexually explicit content, and violence to match your tolerance policy. Option D is correct because you can supply a custom blocklist of prohibited terms or phrases, which the service checks in addition to the built-in safety filters, giving you organization-specific control over banned content. Option A is not a customization method—using default filters unchanged is the absence of customization.

Option B is wrong because disabling all safety filters removes protection rather than customizing it, and it is not a supported way to ensure harmful content is blocked. Option E is not part of Google Cloud's content safety customization features; training a separate detection model is a different, external approach rather than configuring the built-in safety controls.

647
MCQhard

An AI team is building a customer support chatbot for a telecom company using a fine-tuned LLM on Vertex AI. The model performs well on common issues but fails to answer correctly for rare or novel problems, often providing plausible-sounding but incorrect solutions. The team has a large corpus of internal troubleshooting documents. They want to minimize incorrect answers while keeping latency low. Which approach should they take?

A.Switch to a larger base model (e.g., Gemini Ultra) without any retrieval.
B.Implement a retrieval-augmented generation (RAG) pipeline using Vertex AI Search to fetch relevant documents before generating answers.
C.Collect more data on rare issues and continue fine-tuning the model weekly.
D.Use a few-shot prompt with 10 examples of rare problems and solutions.
AnswerB

RAG grounds generation in retrieved internal troubleshooting documents, so rare or novel queries draw on authoritative content rather than parametric guesses. Vertex AI Search supplies relevant passages at inference time, reducing hallucinated answers while keeping latency low.

Why this answer

Implementing a RAG pipeline with Vertex AI Search allows the chatbot to retrieve relevant troubleshooting documents from the internal corpus in real-time, grounding the LLM's responses in authoritative sources. This approach directly addresses the problem of plausible-sounding but incorrect answers for rare/novel issues without requiring retraining, and it keeps latency low by fetching only the most relevant documents before generation.

Exam trap

Google often tests the misconception that fine-tuning or larger models alone can solve knowledge gaps, when in fact retrieval-augmented generation is the standard approach for grounding LLM outputs in up-to-date, domain-specific documents without retraining.

How to eliminate wrong answers

Option A is wrong because switching to a larger base model without retrieval does not solve the core issue of hallucination on rare/novel problems; larger models can still generate plausible-sounding but incorrect answers when they lack specific knowledge, and they often increase latency and cost. Option C is wrong because collecting more data on rare issues and fine-tuning weekly is resource-intensive, may lead to catastrophic forgetting of common issues, and cannot keep pace with the long tail of novel problems that emerge dynamically. Option D is wrong because a few-shot prompt with 10 examples is insufficient to cover the vast space of rare problems, and the model may still hallucinate when the input does not closely match any example, especially without retrieval grounding.

648
Multi-Selecthard

A financial services firm is deploying a generative AI model to assist in loan approval decisions. To comply with regulatory requirements for fairness and explainability, which THREE actions should they take? (Choose 3)

Select 3 answers
A.Add SynthID watermarks to all model outputs
B.Increase the model size to improve accuracy
C.Evaluate the model for bias using diverse test sets
D.Implement chain-of-thought reasoning to explain loan decisions
E.Design a human-in-the-loop process with override capability
AnswersC, D, E

Evaluating with diverse test sets directly satisfies the fairness and explainability constraints by exposing disparate impact across protected groups before deployment. Bias testing quantifies performance gaps between demographic cohorts, producing evidence regulators expect for lending decisions. Without it, discriminatory patterns in training data remain undetected, breaching fair-lending obligations.

Why this answer

Option C is correct because evaluating the model for bias using diverse test sets is essential to detect disparate impact across protected groups (e.g., race, gender, age) and satisfy fairness regulations such as ECOA and fair-lending requirements. Option D is correct because implementing chain-of-thought reasoning produces intermediate reasoning steps that make each loan decision traceable and explainable to regulators, auditors, and applicants, directly supporting explainability obligations. Option E is correct because a human-in-the-loop process with override capability ensures meaningful human review of consequential credit decisions, allowing qualified staff to correct erroneous or unfair AI recommendations and meet accountability requirements.

Option A does not belong because SynthID watermarks only mark AI-generated content for provenance and do not address fairness or explainability of loan decisions. Option B does not belong because increasing model size may improve accuracy but does nothing to guarantee fairness or provide the required explanations, and larger models can even amplify opaque behavior.

Exam trap

The Generative AI Leader exam often tests the distinction between technical safeguards (like watermarks) and governance actions (like bias evaluation and explainability), leading candidates to mistakenly select watermarks as a fairness measure when they are only for content attribution.

649
MCQmedium

Which Google AI model was the first to demonstrate that transformers could be pre-trained bidirectionally on a large corpus, leading to major improvements in language understanding?

A.GPT-3
B.AlphaGo
C.Transformer (the paper)
D.BERT
AnswerD

BERT pre-trains transformers bidirectionally using masked language modelling, so each token attends to both left and right context simultaneously. This bidirectional pre-training on a large corpus produced the language-understanding gains the stem describes, unlike unidirectional models such as GPT, which only attend to preceding tokens.

Why this answer

BERT (Bidirectional Encoder Representations from Transformers) was the first model to demonstrate that transformers could be pre-trained bidirectionally on a large corpus (BooksCorpus and English Wikipedia). By using a masked language model (MLM) objective, BERT conditions on both left and right context simultaneously, unlike previous unidirectional models, leading to significant improvements on 11 NLP benchmarks at its release.

Exam trap

In Google exams, candidates often confuse the original Transformer paper (introducing the architecture) with BERT's specific contribution of bidirectional pre-training, leading to selection of Option C instead of D.

How to eliminate wrong answers

Option A is wrong because GPT-3 is a unidirectional (autoregressive) transformer model that predicts the next token left-to-right, not bidirectionally, and it was released after BERT. Option B is wrong because AlphaGo is a reinforcement learning model for playing the board game Go, not a language model, and it uses convolutional neural networks and Monte Carlo tree search, not bidirectional transformer pre-training. Option C is wrong because the Transformer paper ("Attention Is All You Need") introduced the transformer architecture itself but did not demonstrate bidirectional pre-training on a large corpus; it was a supervised translation model, not a pre-trained language model.

650
Multi-Selectmedium

Which TWO are benefits of using retrieval-augmented generation (RAG) over fine-tuning?

Select 2 answers
A.No need for training
B.Higher accuracy on all tasks
C.More up-to-date information
D.Reduced model size
E.Lower latency
AnswersA, C

RAG eliminates the training step entirely: it supplies relevant context at inference time via a retrieval index, so the model weights stay frozen. This satisfies the stem's benefit of avoiding the compute, cost and labelled-data overhead that fine-tuning demands, letting you update knowledge by refreshing documents instead of retraining.

Why this answer

Option A is correct because RAG does not modify the model's weights; it retrieves relevant documents at inference time and injects them into the prompt, so no training or fine-tuning run is required to add new knowledge. Option C is correct because RAG can pull from an external, continuously updated knowledge base or index, letting the system answer with information newer than the model's static training cutoff, whereas fine-tuning bakes in knowledge only as of the training data used. Options B, D, and E are not benefits guaranteed by RAG: it does not yield higher accuracy on all tasks (retrieval quality can introduce errors), it does not reduce the underlying model's size since the same LLM is still used, and it typically adds retrieval overhead that can increase, not lower, latency.

Exam trap

Google Cloud often tests the misconception that RAG reduces latency or model size, when in fact it increases system complexity and inference time due to the retrieval step, while fine-tuning keeps the model unchanged in size and latency.

651
MCQeasy

A company wants to build a chatbot that answers questions using their internal knowledge base. Which approach is most suitable?

A.Use Retrieval-Augmented Generation (RAG)
B.Fine-tune a model on the knowledge base
C.Train a new model from scratch
D.Use zero-shot prompting with no context
AnswerA

RAG retrieves relevant context and generates answers, perfect for knowledge base Q&A.

Why this answer

RAG retrieves relevant passages from the internal knowledge base at query time and injects them into the model's context, so the chatbot answers with grounded, up-to-date, source-specific information without retraining. It is the standard pattern when the knowledge base changes frequently and citations or freshness matter. Fine-tuning, training from scratch, or zero-shot prompting do not reliably surface private, changing content.

Exam trap

The trap is assuming fine-tuning is the default way to add private knowledge; exams test whether you know RAG is preferred for dynamic, factual, source-grounded retrieval.

How to eliminate wrong answers

Option B is wrong because fine-tuning bakes knowledge into weights, which is expensive, slow to refresh, and prone to hallucination when facts change; it is better for style/format adaptation than factual retrieval. Option C is wrong because training a new model from scratch requires massive compute, data, and expertise, and still would not guarantee retrieval of current internal documents. Option D is wrong because zero-shot prompting with no context gives the model no access to the internal knowledge base, so answers will be generic or fabricated.

652
MCQeasy

Which of the following is a key principle in Google's AI Principles that directly addresses the need to avoid creating or reinforcing unfair bias?

A.Avoid creating or reinforcing unfair bias
B.Uphold high standards of scientific excellence
C.Be socially beneficial
D.Be accountable to people
AnswerA

Google's AI Principles explicitly list avoiding creation or reinforcement of unfair bias as a core objective, addressing fairness directly. The principle commits Google to preventing unjust impacts on people, particularly those relating to sensitive characteristics such as race, gender, and similar protected attributes.

Why this answer

Google's AI Principles explicitly state 'Avoid creating or reinforcing unfair bias' as a standalone principle. This principle directly mandates that AI systems must be designed and tested to mitigate biases in training data, model outputs, and deployment contexts, ensuring fairness across demographic groups. It is the most direct response to the question's focus on avoiding unfair bias.

Exam trap

The Generative AI Leader exam often tests the distinction between principles that directly address bias versus those that are related but broader, so candidates may confuse 'Be socially beneficial' or 'Be accountable to people' as the correct answer because they seem to cover fairness, but they lack the explicit focus on avoiding unfair bias.

How to eliminate wrong answers

Option B is wrong because 'Uphold high standards of scientific excellence' addresses rigor, reproducibility, and methodological soundness, not the specific mitigation of unfair bias. Option C is wrong because 'Be socially beneficial' is a broader principle about overall positive impact, which includes but does not specifically target the avoidance of unfair bias. Option D is wrong because 'Be accountable to people' focuses on transparency, oversight, and redress mechanisms, not the direct prevention of bias in model design or data.

653
MCQmedium

Refer to the exhibit. A data scientist runs the gcloud command and sees the model listed. However, when they try to deploy the model to an endpoint, they get an error: 'Model is not deployable'. What is the most likely reason?

A.The model is still in training and not yet ready.
B.The model was imported from a custom container but without a serving specification or artifact.
C.The model does not have the correct IAM permissions assigned to the deployment service account.
D.The region for the endpoint is different from the model's region.
AnswerB

Importing a custom container without a serving specification or artifact leaves Vertex AI no defined prediction routine or weights to load, so the model registers but cannot be deployed. The serving specification is what makes a model deployable to an endpoint.

Why this answer

A model imported from a custom container must include a serving specification (e.g., a `predict` route) and an artifact (e.g., a saved model file) to be deployable. Without these, Vertex AI cannot determine how to serve predictions, resulting in the 'Model is not deployable' error. The `gcloud` command listing the model only confirms its registration, not its readiness for deployment.

Exam trap

Google Cloud often tests the misconception that a model listed in the registry is automatically deployable, but the trap here is that Vertex AI separates model registration from deployment readiness, requiring explicit serving configuration for custom containers.

How to eliminate wrong answers

Option A is wrong because if the model were still in training, it would not appear in the model list via `gcloud`; Vertex AI only registers a model after training completes. Option C is wrong because IAM permissions affect the deployment action itself (e.g., who can deploy), not the deployability status of the model; the error 'Model is not deployable' is a model-level validation, not an authorization failure. Option D is wrong because region mismatch between the endpoint and model would cause a resource-location error, not a 'Model is not deployable' error; Vertex AI enforces regional consistency but does not block deployment based on region alone.

654
Multi-Selecteasy

A company is prompt engineering a model for customer support. They want to reduce hallucination (false information) in responses. Which TWO techniques are most effective? (Choose two.)

Select 2 answers
A.Implement RAG to retrieve relevant documents for context
B.Provide 3 few-shot examples of conversations
C.Reduce max output tokens to 150
D.Add a system instruction: 'Only answer based on the provided context.'
E.Increase temperature to 1.2
AnswersA, D

RAG provides factual grounding, reducing hallucination.

Why this answer

Retrieval-Augmented Generation (RAG) grounds the model's output in external, verifiable documents retrieved from a knowledge base. By providing relevant context at inference time, RAG significantly reduces the likelihood of the model fabricating information, as it can reference and paraphrase from the retrieved sources rather than relying solely on its parametric memory.

Exam trap

A common mistake in this exam is to think that adjusting parameters like temperature or output tokens directly reduces hallucination, when in fact only techniques that constrain the model's knowledge source (like RAG and strict system instructions) are effective.

655
MCQhard

An enterprise is deploying a customer-facing chatbot using a foundation model on Vertex AI. They need to ensure the model does not produce toxic outputs. Which combination of settings and features should they implement?

A.Use Reinforcement Learning from Human Feedback (RLHF) during fine-tuning
B.Enable Vertex AI Model Monitoring and set up alerts for toxic outputs
C.Reduce the temperature to 0.0 and increase top-k to 50
D.Configure safety filters and safety settings in the model deployment to block harmful categories
AnswerD

Safety filters and configurable safety settings block harmful content categories such as harassment, hate speech and dangerous material before responses reach users. Applied at the model deployment layer, they satisfy the requirement to prevent toxic outputs in a customer-facing chatbot without altering the underlying foundation model.

Why this answer

Safety filters and safety settings in Vertex AI model deployment are the direct mechanism to block harmful categories of output at inference time. These settings allow administrators to define thresholds for categories like toxicity, harassment, and hate speech, ensuring the model refuses to generate prohibited content without requiring retraining or post-hoc monitoring.

Exam trap

Google often tests the distinction between training-time alignment techniques (like RLHF) and inference-time safety controls (like safety filters), tempting candidates to choose a fine-tuning approach when the question explicitly asks for deployment settings to prevent toxic outputs.

How to eliminate wrong answers

Option A is wrong because RLHF is a fine-tuning technique that aligns model behavior based on human preferences, but it does not provide real-time blocking of toxic outputs during inference; it only influences the model's general behavior after training. Option B is wrong because Vertex AI Model Monitoring is designed for detecting data drift and performance anomalies, not for filtering or blocking toxic content in real-time responses. Option C is wrong because reducing temperature to 0.0 makes the model deterministic and increasing top-k to 50 broadens token selection, which does not prevent toxicity; these parameters control randomness, not content safety.

656
Multi-Selectmedium

A company is considering whether to use Vertex AI's Generative AI Studio. Which TWO are benefits?

Select 2 answers
A.It is always cheaper than using third-party APIs
B.It integrates seamlessly with Vertex AI Pipelines for MLOps
C.It generates outputs that are always more accurate than custom models
D.It provides built-in tools for prompt engineering and iterative testing
E.It requires no coding or machine learning expertise to use
AnswersB, D

Integration allows automating deployment, monitoring, and retraining.

Why this answer

Vertex AI Generative AI Studio is designed to work natively with Vertex AI Pipelines, enabling users to incorporate generative models into end-to-end MLOps workflows for automation, monitoring, and retraining. This integration allows seamless orchestration of prompt tuning, model evaluation, and deployment within the same managed environment, reducing operational overhead.

Exam trap

Google Cloud often tests the misconception that 'no-code' tools eliminate the need for any ML expertise, but the trap here is that Generative AI Studio still requires understanding of prompt engineering, model evaluation, and cost trade-offs to avoid poor outputs or unexpected expenses.

657
MCQmedium

A software company wants to provide users with a clear understanding of when and why their AI system may produce incorrect answers. Which tool from the Responsible AI toolkit should they use to communicate model limitations?

A.People + AI Guidebook
B.PAIR Explorables
C.Model Cards
D.Datasheets for Datasets
AnswerC

Model Cards document a model's intended use, performance metrics, and known limitations, directly satisfying the requirement to communicate when and why incorrect answers occur. Unlike dashboards or error-analysis tools, they are static disclosure artefacts designed for stakeholders, making them the appropriate Responsible AI toolkit component for transparently explaining limitations to users.

Why this answer

Model Cards are designed to communicate model performance, intended use, and limitations to stakeholders in a standardized format.

658
MCQmedium

A company wants to generate high-quality product images from text descriptions for an e-commerce catalog. They need photorealistic results. Which model and approach should they choose?

A.Use Veo for video generation and extract frames
B.Fine-tune Gemini 1.5 Pro on product images
C.Use Imagen on Vertex AI with appropriate prompts
D.Use Codey to generate code that renders images
AnswerC

Imagen on Vertex AI is purpose-built for text-to-image synthesis, producing photorealistic outputs from descriptive prompts, which directly satisfies the catalogue's photorealism requirement. Its prompt-based generation maps text descriptions to high-fidelity product imagery without training custom models, matching the scenario's need for quality results from text alone.

Why this answer

Imagen on Vertex AI is specifically designed for high-quality, photorealistic text-to-image generation, making it the ideal choice for creating product images from text descriptions. It leverages advanced diffusion models to produce detailed and visually accurate outputs that meet the requirements of an e-commerce catalog.

Exam trap

Google often tests the distinction between models specialized for different modalities (text, image, video, code) to see if candidates recognize that a dedicated image generation model like Imagen is required for photorealistic text-to-image tasks, rather than repurposing video or code models.

How to eliminate wrong answers

Option A is wrong because Veo is a video generation model, and extracting frames from video would introduce motion artifacts, temporal inconsistencies, and lower resolution compared to a dedicated image generation model, failing to achieve photorealistic results. Option B is wrong because Gemini 1.5 Pro is a multimodal large language model optimized for understanding and generating text, code, and reasoning, not for high-fidelity image generation; fine-tuning it on product images would not produce photorealistic outputs as it lacks a diffusion-based image generation architecture. Option D is wrong because Codey is a code generation model designed to produce code snippets, not to render images; using it to generate code that renders images would require additional rendering engines and would not directly produce photorealistic images from text descriptions.

659
Multi-Selectmedium

A development team is integrating a large language model into a healthcare application. They need to reduce the risk of generating harmful medical advice. Which THREE measures should they implement? (Choose three.)

Select 3 answers
A.Use a safety filter to block outputs containing harmful medical terminology.
B.Implement RAG to retrieve verified medical information from trusted sources.
C.Fine-tune the model on a curated dataset of medical textbooks.
D.Include a disclaimer in the system instruction that the model is not a doctor.
E.Set the temperature to a very high value to ensure diverse outputs.
AnswersA, B, C

Safety filters directly block harmful content at inference time.

Why this answer

Implementing a safety filter that blocks outputs containing harmful medical terminology directly mitigates the risk of generating dangerous advice. This acts as a post-processing guardrail, intercepting model outputs that include terms associated with diagnoses, dosages, or procedures that could lead to patient harm. It is a standard practice in high-stakes domains to layer such filters on top of the generative model.

Exam trap

The Generative AI Leader exam often tests the misconception that disclaimers or system instructions alone are sufficient safety measures, when in fact they do not technically prevent the model from generating harmful content—only post-hoc filtering or architectural controls like RAG and fine-tuning can reduce the risk at the output level.

660
MCQhard

A hospital network wants patients to describe symptoms in a mobile app and receive immediate guidance, but the clinical knowledge base changes weekly and the network must be able to update answers without retraining a model. They also need the assistant to escalate to a nurse when confidence is low. Which Google Cloud approach best meets these requirements?

A.Fine-tune a Gemini model on the clinical knowledge base each week and deploy it to a Vertex AI endpoint
B.Deploy an open model from Vertex AI Model Garden and manage the serving infrastructure on Google Kubernetes Engine
C.Build a Vertex AI Agent Builder agent grounded on a frequently refreshed Vertex AI Search data store, with a tool that triggers nurse escalation
D.Use the Gemini API directly with a long system prompt containing the entire clinical knowledge base
AnswerC

Vertex AI Agent Builder supports agents that combine grounding on enterprise data stores with tools and function calling. Pointing the agent at a Vertex AI Search data store that is refreshed weekly keeps answers current without retraining, and defining a tool that triggers nurse escalation satisfies the handoff requirement. This architecture addresses both the changing knowledge base and the escalation workflow in a managed way.

Why this answer

The requirements pair a knowledge base that changes frequently with a need to escalate to a human, which maps to an agent grounded on a refreshable data store plus a tool for escalation. Updating the search data store keeps content current without retraining, while function calling handles the nurse handoff. Fine-tuning, giant prompts, or self-managed open models fail one or both constraints.

Exam trap

The trap here is reaching for fine-tuning to inject knowledge, when frequently changing content is better served by grounding on a refreshable data store rather than baked-in weights.

661
MCQeasy

A project manager wants to automatically generate weekly status reports from meeting notes and project data. The team uses Google Workspace. Which built-in capability is the QUICKEST to implement?

A.Use Gemini for Workspace (Duet AI) in Google Docs and Google Meet to generate summaries and reports
B.Select a model from Model Garden and deploy it as a private endpoint for report generation
C.Use Vertex AI Studio to design a custom prompt and call the Gemini API from a custom app
D.Write a Google Apps Script to call the Gemini API and format the report
AnswerA

Gemini for Workspace is already embedded in Google Docs and Meet, so it can summarise meeting notes and draft reports without custom development or third-party integration. This satisfies the stem's Google Workspace constraint and its quickest-to-implement requirement.

Why this answer

Gemini for Workspace (Duet AI) can generate summaries directly in Google Docs and Meet, leveraging existing data with no custom development. Custom API integration or Model Garden would require more effort. Apps Script is for custom automation, not built-in.

662
MCQhard

A healthcare organization needs a generative AI model to answer medical questions using proprietary clinical guidelines. They have a large dataset of doctor-patient interactions. Should they fine-tune a pre-trained model or use Retrieval-Augmented Generation (RAG)?

A.Use RAG to reduce inference costs by skipping model updates.
B.Use RAG to retrieve relevant guidelines during inference, avoiding frequent retraining.
C.Use prompt engineering to encode all guidelines into the system prompt.
D.Fine-tune the model on the clinical guidelines and interactions.
AnswerB

RAG retrieves the relevant clinical guideline passages at inference time and supplies them as context, so answers stay grounded in proprietary content without retraining. This satisfies the need to reflect frequently updated guidelines, since the index can be refreshed without touching model weights.

Why this answer

RAG is preferred because it retrieves the most current clinical guidelines from an external knowledge base during inference, avoiding the need for constant retraining when guidelines change. This is especially important in healthcare where regulations update frequently. Fine-tuning a pre-trained model on proprietary interactions might lead to overfitting or outdated knowledge.

Option A is incorrect because RAG does not necessarily reduce inference costs; it adds overhead for retrieval. Option C is incorrect because prompt engineering cannot encode all proprietary guidelines; it is limited by context window size. Option D is incorrect because fine-tuning requires retraining when guidelines change, which is less flexible than RAG.

663
MCQhard

A financial services firm wants to deploy a generative AI assistant that summarizes analyst research for internal advisors. The security team requires that the assistant never expose confidential client data in its responses, and the business wants to move quickly without building a custom model from scratch. Which approach best balances speed, control, and data protection on Google Cloud?

A.Fine-tune a foundation model on all historical analyst reports, including confidential client sections.
B.Deploy an open-source model on a single Compute Engine VM and rely on network firewalls for data protection.
C.Store all analyst reports in a public Cloud Storage bucket so the model can retrieve them without authentication.
D.Use a managed foundation model with grounding on an access-controlled document store and Sensitive Data Protection filters on inputs and outputs.
AnswerD

Grounding the assistant on an access-controlled corpus ensures it summarizes only documents the advisor is permitted to see, while Sensitive Data Protection inspection on prompts and responses adds a layer that can detect and redact confidential client identifiers. This uses managed models, so the firm avoids building from scratch, and it keeps sensitive data out of model weights, which supports both speed and the security team's non-exposure requirement.

Why this answer

Combining a managed foundation model with grounded retrieval over an access-controlled store and Sensitive Data Protection filtering keeps confidential client data out of model weights and out of responses. Advisors receive summaries only from documents they are authorized to view, and inspection adds detection and redaction for sensitive identifiers. This satisfies the security requirement while avoiding the cost and delay of training a custom model.

Exam trap

The trap here is assuming that network perimeter controls or fine-tuning alone protect confidential data, when the real risk is sensitive content appearing in model context or generated responses.

664
MCQeasy

A data scientist needs to generate high-quality images from text prompts using Google Cloud. Which service should they use?

A.Imagen
B.PaLM 2
C.Gemini Pro Vision
D.Codey
AnswerA

Imagen is Google Cloud's text-to-image generation model, converting natural-language prompts directly into high-quality images. It satisfies the stem's requirement for image generation on Google Cloud, unlike Vertex AI's text-focused Gemini models or unrelated analytics services. The data scientist should therefore select Imagen.

Why this answer

Imagen is Google Cloud's text-to-image diffusion model. PaLM 2 and Gemini are primarily text models; Codey is for code generation.

665
Multi-Selectmedium

Which TWO of the following are best practices for prompt engineering?

Select 2 answers
A.Provide context and examples in the prompt
B.Append random noise to prompts to improve creativity
C.Use clear and specific instructions
D.Always use the maximum possible number of tokens
E.Use negative prompts to discourage undesired outputs
AnswersA, C

Supplying context and few-shot examples steers the model toward the desired output format and domain, reducing ambiguity. It satisfies the best-practise criterion by grounding responses in relevant information rather than relying on the model's parametric memory alone.

Why this answer

Option A is correct because supplying relevant context and few-shot examples in the prompt grounds the model, clarifies the expected format and intent, and measurably improves output relevance and accuracy. Option C is correct because clear, specific instructions reduce ambiguity, constrain the model's response space, and yield more predictable, on-target results. Together, these two practices are widely recommended in prompt engineering guidance for both large language models and generative AI services.

Option B does not belong because adding random noise degrades coherence and reliability rather than improving genuine creativity. Option D does not belong because maximizing token count wastes context window and cost without improving quality; concise, purposeful prompts are preferred. Option E does not belong because negative prompts are a feature specific to certain image-generation tools and are not a general best practice for prompt engineering across LLMs.

666
MCQeasy

Which Google resource provides interactive visualizations and exercises to help AI practitioners understand concepts like fairness, interpretability, and privacy?

A.PAIR Explorables
B.Model Cards
C.People + AI Guidebook
D.TensorFlow Fairness Indicators
AnswerA

PAIR Explorables are interactive, browser-based visualisations and exercises from Google's People + AI Research team, letting practitioners experiment with fairness, interpretability and privacy concepts hands-on. This interactive format matches the stem's requirement for visualisations and exercises rather than static documentation or policy guidance.

Why this answer

PAIR (People + AI Research) Explorables is the correct answer because it is a Google resource specifically designed to provide interactive visualizations and hands-on exercises that help AI practitioners grasp complex concepts like fairness, interpretability, and privacy. Unlike static documentation or tools, Explorables allow users to manipulate parameters and see real-time effects, making abstract responsible AI principles tangible and actionable.

Exam trap

Google often tests the distinction between educational/interactive resources and practical implementation tools, so the trap here is that candidates confuse Model Cards or Fairness Indicators (which are about applying fairness) with PAIR Explorables (which are about learning fairness concepts through interaction).

How to eliminate wrong answers

Option B (Model Cards) is wrong because Model Cards are standardized documentation templates that disclose model performance, intended use, and fairness evaluations, but they do not offer interactive visualizations or exercises for learning concepts. Option C (People + AI Guidebook) is wrong because it is a static design guide with best practices and patterns for human-AI interaction, not an interactive learning tool with visualizations and exercises. Option D (TensorFlow Fairness Indicators) is wrong because it is a suite of tools for computing and visualizing fairness metrics on model evaluations, but it is a practical debugging tool, not an educational resource with interactive exercises to teach underlying concepts.

667
MCQhard

A financial institution is using Vertex AI to generate personalized investment advice. They need to ensure that the model's responses are grounded in the latest regulatory documents and do not include outdated or fabricated information. Which feature should they implement to achieve this?

A.Fine-tuning the model on historical advisory emails
B.Grounding with Vertex AI Search
C.Using a larger context window
D.Increasing the model's temperature setting
AnswerB

Grounding with Vertex AI Search allows the model to retrieve relevant information from a specified data store, such as a repository of regulatory documents, and use that information to generate responses. This ensures that the advice is based on the latest documents and reduces hallucinations. It is the correct approach for keeping responses current and factual without retraining the model.

Why this answer

Grounding with Vertex AI Search enables the model to dynamically retrieve and cite information from a connected data store, such as a set of regulatory documents. This ensures that the generated advice is based on the most current and authoritative sources, reducing the risk of outdated or fabricated content. The other options do not provide real-time grounding and are either static or counterproductive for factual accuracy.

Exam trap

The trap here is thinking that fine-tuning or a larger context window can solve the need for up-to-date information, but only grounding with a search service provides dynamic retrieval.

668
Multi-Selectmedium

A team is preparing to use Vertex AI Studio to prototype a generative AI application. They want to understand which factors directly influence the quality and safety of the model's responses. (Choose two.)

Select 2 answers
A.The configured safety filters and their threshold settings
B.The color scheme of the application's user interface
C.The physical region where the developer's laptop is located
D.The programming language used to call the Vertex AI API
E.The clarity and specificity of the prompt provided by the user
AnswersA, E

Safety filters with adjustable thresholds determine which categories of content the model will block or allow, directly affecting the safety of responses. Tightening thresholds reduces harmful output, while loosening them permits more content. These settings are a core control for responsible deployment and must be tuned to the application's risk profile.

Why this answer

Prompt construction and safety filter configuration are the two levers that most directly shape a generative model's responses in Vertex AI Studio. A precise prompt improves relevance and accuracy, while safety thresholds govern what content is permitted. Presentation choices, client location, and programming language affect the surrounding application but not the model's generation or filtering behavior, so they do not influence response quality or safety.

Exam trap

The trap here is assuming that implementation details like the calling language or UI design affect model output, when only prompt and safety configuration do.

669
MCQeasy

A data scientist needs to fine-tune a foundation model for a sentiment analysis task without managing infrastructure. Which Google Cloud service should they use?

A.Compute Engine
B.BigQuery ML
C.Cloud Run
D.Vertex AI Model Garden
AnswerD

Vertex AI Model Garden supplies foundation models with tuning workflows that run on managed infrastructure, so no cluster provisioning is needed. It directly satisfies the stem's constraint of fine-tuning for sentiment analysis without infrastructure management, unlike raw compute or self-hosted alternatives.

Why this answer

Vertex AI Model Garden is the correct service because it provides a curated hub of foundation models that can be fine-tuned with managed infrastructure, eliminating the need for the data scientist to provision or manage servers. It supports one-click deployment and fine-tuning workflows for sentiment analysis, directly addressing the requirement to avoid infrastructure management.

Exam trap

The trap here is that candidates often confuse BigQuery ML's ability to train models on tabular data with the capability to fine-tune large language models, but BigQuery ML does not support fine-tuning of foundation models for NLP tasks.

How to eliminate wrong answers

Option A is wrong because Compute Engine is an IaaS offering that requires the user to manually provision, configure, and manage virtual machines, which contradicts the requirement of not managing infrastructure. Option B is wrong because BigQuery ML is designed for creating and executing machine learning models using SQL queries on structured data in BigQuery, not for fine-tuning large foundation models for natural language tasks like sentiment analysis. Option C is wrong because Cloud Run is a serverless container platform for running stateless HTTP-driven applications, but it does not provide native support for fine-tuning foundation models; it would require the user to build and manage the fine-tuning pipeline themselves.

670
MCQmedium

A company fine-tunes a model using Vertex AI and notices the model's performance drops on the original training task (e.g., language understanding) after fine-tuning for a new task (e.g., summarization). What could be the cause?

A.Data leakage
B.Model quantization
C.Catastrophic forgetting
D.Underfitting
AnswerC

Catastrophic forgetting occurs when gradient updates for the new summarisation task overwrite the weights encoding the original language-understanding capability. This directly explains the performance drop on the earlier task, satisfying the stem's observation of degraded prior-task accuracy after fine-tuning.

Why this answer

Catastrophic forgetting occurs when a neural network loses previously learned knowledge upon being fine-tuned on a new task. In this scenario, fine-tuning the model for summarization overwrites the weights responsible for language understanding, causing performance degradation on the original task. This is a well-known limitation of sequential fine-tuning in deep learning.

Exam trap

Google Cloud often tests the distinction between catastrophic forgetting and underfitting, as candidates may mistakenly think the model simply didn't learn the new task well, rather than recognizing that it forgot the original task due to weight overwriting.

How to eliminate wrong answers

Option A is wrong because data leakage refers to the inadvertent exposure of target information during training, which would typically inflate performance metrics rather than cause a drop on the original task. Option B is wrong because model quantization reduces numerical precision (e.g., from FP32 to INT8) to improve inference speed and memory efficiency, but it does not inherently cause performance loss on a previously learned task; any accuracy loss from quantization is generally uniform across tasks. Option D is wrong because underfitting means the model fails to capture patterns in the training data, resulting in poor performance on both the original and new tasks, not a selective drop on the original task after fine-tuning.

671
MCQhard

A healthcare organization is using a generative AI model to summarize patient discharge instructions. They need to ensure the summaries are accurate and do not omit critical information. Which technique should they implement to reduce the risk of omissions?

A.Increase the model's token limit to allow longer summaries that can include more details.
B.Fine-tune the model on a large dataset of medical summaries to improve its ability to capture key points.
C.Implement a retrieval-augmented generation (RAG) system that retrieves relevant sections of the discharge instructions and includes them in the prompt.
D.Use chain-of-thought prompting to encourage the model to reason step by step.
AnswerC

RAG ensures the model has access to the actual discharge instructions. By retrieving and including relevant sections, the model is less likely to omit critical information because it is grounded in the source text. This directly addresses the risk of omissions by providing the necessary context.

Why this answer

Retrieval-augmented generation (RAG) is the best choice because it grounds the model in the actual discharge instructions, reducing the chance of omissions. By retrieving and including relevant text in the prompt, the model can generate summaries that are comprehensive and accurate. Other techniques like fine-tuning or chain-of-thought do not directly ensure that all critical information from the source is captured.

Exam trap

The trap here is thinking that a more powerful model or longer output will automatically capture all critical details, when the real issue is ensuring the model references the source document.

672
MCQhard

A large enterprise is using Vertex AI to deploy a generative AI model for internal document summarization. They need to ensure that the model's responses are based on the most current internal documents and that the model does not hallucinate. They also want to minimize latency and cost. Which feature of Vertex AI should they implement?

A.Grounding with Vertex AI Search
B.Model evaluation
C.Batch prediction
D.Model fine-tuning
AnswerA

Grounding with Vertex AI Search allows the model to retrieve relevant information from a specified data store, such as internal documents, and use it to generate responses. This reduces hallucinations and ensures responses are based on current, proprietary data. It also minimizes the need for fine-tuning, which can be costly and time-consuming.

Why this answer

Grounding with Vertex AI Search is the correct feature because it enables the model to retrieve and cite relevant passages from a designated data store, ensuring responses are based on the latest internal documents. This approach reduces hallucinations and avoids the need for frequent fine-tuning, which can be costly and slow. It also supports low-latency responses by integrating retrieval with generation.

Exam trap

The trap here is assuming that fine-tuning is the default solution for domain adaptation, when grounding is more effective for dynamic, up-to-date information.

673
Multi-Selectmedium

A company is using Vertex AI to generate report summaries. They want to measure ROI. Which three metrics should they track? (Choose THREE)

Select 3 answers
A.Error rate (e.g., factual errors in summaries)
B.Number of employees trained on the system
C.Number of active users per month
D.Accuracy score of generated summaries compared to human-written ones
E.Time saved per report (minutes)
AnswersA, D, E

Reducing errors improves trust and saves correction time.

Why this answer

Error rate directly measures the quality and reliability of generated summaries, which is a key factor in determining ROI. If summaries contain factual errors, they require human correction, reducing the time savings and potentially introducing business risks. Tracking error rate helps quantify the cost of inaccuracies against the benefits of automation.

Exam trap

The trap here is that candidates confuse adoption metrics (active users, training count) with ROI metrics, but ROI for Vertex AI requires direct financial or efficiency quantification, not just usage statistics.

674
Multi-Selecthard

A company is deploying a document summarization solution using Vertex AI. They want to minimize cost while maintaining quality. Which three strategies should they implement? (Choose THREE)

Select 3 answers
A.Use prompt compression to cut token usage
B.Choose a smaller, specialized model (e.g., Gemini 1.5 Flash)
C.Implement response caching for repeated queries
D.Use the largest available model for best quality
E.Batch multiple document summarization requests together
AnswersB, C, E

Gemini 1.5 Flash delivers lower per-token inference pricing than larger Gemini Pro models, directly satisfying the stem's cost-minimisation constraint. Its summarisation quality remains sufficient for document summarisation, so the quality requirement holds. Selecting a smaller specialised model therefore reduces spend without materially degrading output.

Why this answer

Using caching reduces repeated processing, choosing a smaller model lowers token cost, and batching requests minimizes overhead. Prompt compression is not a standard Vertex AI feature, and using the largest model increases cost.

675
MCQhard

A developer uses the Vertex AI Python SDK to call a Gemini model for structured JSON output. However, the model often returns malformed JSON. Which parameter should the developer set in the generation configuration to enforce valid JSON output?

A.Set the temperature to a lower value (0.1) to reduce variation.
B.Set the 'response_mime_type' parameter to 'application/json'.
C.Include few-shot examples of the desired JSON format in the system prompt.
D.Switch to a smaller model to reduce complexity.
AnswerB

Setting `response_mime_type` to `application/json` constrains Gemini's decoding to emit syntactically valid JSON, directly satisfying the stem's requirement for structured output. This parameter enforces the output format at generation time, eliminating the malformed responses the developer currently receives from the Vertex AI Python SDK.

Why this answer

Setting `response_mime_type` to `'application/json'` in the generation configuration instructs the Gemini API to constrain the model's output to valid JSON format. This parameter leverages the model's native structured output capability, ensuring the response adheres to JSON syntax without relying on post-processing or prompt engineering.

Exam trap

Google Cloud often tests the misconception that prompt engineering (e.g., few-shot examples or temperature tuning) can reliably enforce structured output, when in fact the correct approach is to use the API's native structured output parameter like `response_mime_type`.

How to eliminate wrong answers

Option A is wrong because lowering temperature reduces randomness but does not enforce structural constraints; the model can still produce malformed JSON due to token-level deviations. Option C is wrong because few-shot examples in the system prompt improve formatting consistency but do not guarantee valid JSON output, as the model may still generate syntax errors or deviate from the schema. Option D is wrong because switching to a smaller model reduces capacity and may increase the likelihood of malformed output, and model size does not address the need for structured output enforcement.

Page 8

Page 9 of 14

Page 10