Courseiva

Google Cloud Generative AI Leader Generative AI Leader (Generative AI Leader) — Questions 1–75

1008 questions total · 14pages · All types, answers revealed

Page 1 of 14

Page 2
1
Multi-Selectmedium

A retail company wants to build an internal knowledge base chatbot using Vertex AI. They need to ensure the chatbot only answers from approved company documents and can handle updates without retraining. Which TWO components should they include? (Choose 2)

Select 2 answers
A.RAG Engine to connect the chatbot to a document index
B.A vector store (e.g., Vertex AI Vector Search) to index the documents
C.Fine-tuned model on company documents
D.Model Garden to select a pre-trained model
E.Apps Script to update the document index
AnswersA, B

RAG Engine retrieves relevant passages from an indexed corpus of approved company documents and supplies them as grounding context, so answers stay within that content. Because the index updates independently, documents can change without retraining the underlying model.

Why this answer

Option A is correct because RAG Engine (Retrieval-Augmented Generation) is the Vertex AI component that grounds a chatbot's responses in an external document index, ensuring answers come only from approved company documents rather than the model's parametric knowledge. Option B is correct because a vector store such as Vertex AI Vector Search is required to embed and index the documents so that relevant chunks can be retrieved at query time, and the index can be updated with new documents without retraining the underlying model. Together, RAG Engine plus a vector store satisfy both requirements: grounded answers from approved sources and updates via index refresh instead of retraining.

Option C is not appropriate because fine-tuning bakes document knowledge into model weights, which requires retraining to reflect updates and does not restrict answers to approved documents. Option D is not sufficient because Model Garden only lets you discover and select pre-trained or foundation models; it provides no grounding or document retrieval capability. Option E is incorrect because Apps Script is a Google Workspace automation tool and is not used to build or update a Vertex AI vector index.

Exam trap

The trap is assuming fine-tuning is needed to make a model 'know' company documents; the exam expects you to recognize RAG plus a vector store as the no-retrain, grounded-answer pattern.

2
MCQhard

A healthcare startup uses a generative model fine-tuned on general medical literature to provide preliminary diagnostic suggestions from patient text. The model frequently misses rare diseases and sometimes suggests common conditions that are unlikely given the symptoms. The startup has a curated dataset of rare disease case reports and wants to improve the model’s sensitivity to rare conditions without sacrificing overall accuracy. They cannot afford to retrain the entire model from scratch. The model is deployed on Vertex AI Prediction with low latency requirement. Which approach should they take?

A.Perform continued fine-tuning on the rare disease dataset using a low learning rate.
B.Add a system prompt instructing the model to consider rare diseases more carefully.
C.Reduce top-p sampling to focus on high-probability tokens, assuming rare diseases have lower probability.
D.Implement a human-in-the-loop system: for outputs with low confidence or suspected rare disease, route to a human expert.
AnswerD

Human-in-the-loop catches edge cases without retraining, preserving accuracy for common conditions.

Why this answer

Implementing a human-in-the-loop process for rare disease flags combines AI with expert review, catching misses while maintaining speed for common cases. Option A is wrong because prompt engineering alone may not teach the model about rare diseases. Option B is wrong because increasing top-p restricts vocabulary but doesn't inject knowledge.

Option C is wrong because fine-tuning again might cause catastrophic forgetting of common conditions.

3
MCQhard

A media company is building an internal tool that generates first-draft marketing copy from campaign briefs. Legal insists that the tool never reproduce copyrighted third-party text verbatim, and the content team wants a measurable way to compare draft quality across prompt revisions. Which two-part approach best addresses both needs?

A.Lower the model temperature to zero and require editors to sign off on every draft before it is used.
B.Enable Vertex AI safety filters and recitation checking, and use Vertex AI evaluation to score draft quality across prompt versions.
C.Add a keyword blocklist that strips any phrase appearing in a public style guide from generated drafts.
D.Deploy the largest available model with maximum output tokens and rely on human editors to catch any copied passages.
AnswerB

Recitation checking detects when generated content closely matches training data and can flag or block it, directly addressing the verbatim-copying concern. Vertex AI evaluation provides repeatable metrics on draft quality, giving the content team an objective basis to compare prompt revisions instead of relying on subjective impressions.

Why this answer

Recitation checking inspects generated output for passages that closely match training data and can block or flag them, which is the systematic safeguard legal requires. Vertex AI evaluation supplies repeatable quality scores, so prompt revisions can be compared with evidence rather than opinion. Together they cover both the compliance and the measurement objectives in one design.

Exam trap

The trap here is believing that lower temperature or human review eliminates verbatim reproduction, when recitation is a distinct detection capability that must be explicitly enabled and measured separately.

4
MCQhard

A large enterprise is deploying a generative AI system for automated contract review. The system must provide confidence indicators for its legal analysis. How should confidence indicators be implemented to maximize transparency?

A.Show the top-k most likely outcomes without probabilities
B.Provide a numerical confidence score between 0 and 1 for each conclusion
C.Display a binary pass/fail indicator for each analysis
D.Hide confidence indicators to avoid confusing users
AnswerB

A numerical score between 0 and 1 quantifies certainty for each conclusion, letting reviewers judge how much weight to place on the legal analysis. This satisfies the transparency requirement by exposing the model's confidence rather than presenting every conclusion as equally reliable.

Why this answer

Providing a numerical confidence score between 0 and 1 for each conclusion directly quantifies the model's certainty, enabling users to assess the reliability of each legal analysis. This approach maximizes transparency by allowing legal professionals to calibrate their trust in the AI's output, which is critical for high-stakes contract review where false positives or negatives carry significant risk.

Exam trap

Google often tests the misconception that binary outputs (pass/fail) are sufficient for transparency, but the trap here is that binary indicators hide the model's uncertainty, which is exactly what confidence scores are designed to reveal in responsible AI deployments.

How to eliminate wrong answers

Option A is wrong because showing the top-k most likely outcomes without probabilities omits the crucial uncertainty information; users cannot gauge how much more likely one outcome is over another, which undermines transparency in legal decision-making. Option C is wrong because a binary pass/fail indicator oversimplifies the model's output, hiding the nuanced confidence levels that are essential for evaluating ambiguous contract clauses or borderline legal interpretations. Option D is wrong because hiding confidence indicators entirely defeats the purpose of transparency, leaving users with no insight into the model's reliability and potentially leading to blind trust or unwarranted skepticism.

5
Multi-Selectmedium

A company uses a generative AI model to produce financial reports. They want to ensure content safety and prevent the generation of harmful or misleading information. Which TWO Google Cloud features should they configure? (Choose 2)

Select 2 answers
A.Fine-tune the model on a dataset of approved financial reports
B.Log all prompts and responses in Cloud Logging
C.Enable Google's safety filters for hate speech, violence, and sexual content
D.Use SynthID to watermark all outputs
E.Configure custom content controls in Vertex AI to block financial misinformation
AnswersC, E

Google's safety filters block harmful categories such as hate speech, violence and sexual content before responses reach users, directly satisfying the content-safety constraint in the stem. Configuring these thresholds on the generative AI model prevents toxic output in financial reports, though they do not address factual accuracy or misleading claims.

Why this answer

Option C is correct because Google's built-in safety filters in Vertex AI (via the SafetySettings on GenerativeModel requests) let the company block or threshold harmful categories such as hate speech, harassment, violence, and sexually explicit content, directly enforcing content safety on generated financial reports. Option E is correct because Vertex AI supports configurable content controls, including custom safety attributes and blocklists, which can be tuned to detect and block domain-specific harmful output such as financial misinformation before it reaches users. Option A is not correct here because fine-tuning on approved reports improves style and factual alignment but does not by itself provide runtime safety filtering against harmful or misleading content.

Option B is not correct because logging prompts and responses in Cloud Logging is for auditing and observability, not for preventing unsafe generation. Option D is not correct because SynthID watermarks AI-generated content for provenance and detection, but it does not block or prevent harmful or misleading output.

Exam trap

The trap here is that candidates often confuse post-hoc logging (Option B) or watermarking (Option D) with proactive content safety controls, failing to recognize that only runtime filtering mechanisms like safety filters and custom content controls can prevent generation of harmful outputs.

6
MCQhard

A financial services firm uses a fine-tuned model for contract analysis. They observe that the model's performance degrades after a few months because contract language evolves. The team wants to maintain accuracy without full retraining. What is the MOST cost-effective approach?

A.Switch to a larger base model and use zero-shot prompting
B.Retrain the model from scratch every quarter with all historical data
C.Perform incremental fine-tuning with a small representative sample of new contracts
D.Use Vertex AI Model Monitoring to detect drift and alert, then manually adjust prompts
AnswerC

Incremental fine-tuning updates only the model's weights using a small sample of recent contracts, avoiding the compute cost of full retraining while adapting to evolving language. This directly satisfies the cost-effectiveness constraint and restores accuracy as contract wording drifts.

Why this answer

Incremental fine-tuning updates the existing fine-tuned model's weights using a small, representative sample of recent contracts, adapting to evolving language at a fraction of the cost of full retraining while preserving prior knowledge. This directly addresses drift without the expense of rebuilding from scratch or the inaccuracy of prompt-only fixes.

Exam trap

Generative AI Leader often tests the misconception that monitoring plus prompt tweaking is equivalent to model adaptation, when drift in a fine-tuned model requires weight updates, not just prompt changes.

How to eliminate wrong answers

Option A is wrong because switching to a larger base model with zero-shot prompting discards the domain-specific fine-tuning already paid for and typically underperforms a tuned model on specialized contract language. Option B is wrong because full retraining from scratch every quarter with all historical data is the most expensive option and unnecessary when only incremental drift needs correction. Option D is wrong because monitoring detects drift but does not fix it — manually adjusting prompts is a stopgap that does not restore the fine-tuned model's accuracy on evolving contract terminology.

7
Multi-Selectmedium

A company is establishing governance practices for generative AI models. Which three actions are essential for responsible AI deployment?

Select 3 answers
A.Use model versioning to track changes.
B.Regularly audit model outputs for bias.
C.Monitor for data leakage from training data.
D.Implement a human review process for critical decisions.
E.Open-source the model to ensure transparency.
AnswersA, B, D

Versioning ensures reproducibility and accountability for model updates.

Why this answer

Model versioning (Option A) is essential because it enables tracking of changes to generative AI models over time, ensuring reproducibility, rollback capability, and compliance with governance policies. Without versioning, it becomes impossible to audit which model produced a specific output, undermining accountability and regulatory adherence.

Exam trap

Candidates often confuse operational security practices (like data leakage monitoring) with core governance actions (like versioning, auditing, and human review), leading them to select Option C as essential when it is actually a secondary security measure.

8
MCQhard

A team is fine-tuning a large language model on custom data using Vertex AI. They find that the training loss decreases but validation loss increases. What is the best course of action?

A.Increase the number of training epochs.
B.Reduce the model size or add dropout regularization.
C.Increase the learning rate.
D.Switch to a smaller batch size.
AnswerB

Adding dropout regularisation directly counteracts the overfitting causing validation loss to diverge from training loss. Reducing model capacity likewise limits memorisation of the custom fine-tuning data. Both address the generalisation gap the stem describes, where training loss falls while validation loss rises, restoring alignment between the two curves.

Why this answer

The increasing validation loss while training loss decreases is a classic sign of overfitting, where the model memorizes the training data but fails to generalize. Reducing model size or adding dropout regularization directly combats overfitting by limiting the model's capacity or introducing noise during training, which forces the model to learn more robust features. This is the best course of action because it addresses the root cause without further exacerbating the problem.

Exam trap

Google Cloud often tests the distinction between underfitting and overfitting, and the trap here is that candidates may confuse increasing validation loss with underfitting and incorrectly choose to increase epochs or learning rate, rather than recognizing the hallmark divergence of overfitting.

How to eliminate wrong answers

Option A is wrong because increasing the number of training epochs would further overfit the model to the training data, worsening the validation loss. Option C is wrong because increasing the learning rate can cause the model to overshoot minima and destabilize training, potentially increasing both training and validation loss, and does not address overfitting. Option D is wrong because switching to a smaller batch size introduces more noise in gradient estimates, which can sometimes help generalization but is not a direct or reliable remedy for overfitting; it may also slow convergence and is not the primary solution for the described loss divergence.

9
MCQmedium

A global retailer's customer service team wants to deploy a generative AI chatbot that answers questions about order status, return policies, and product availability. The chatbot must always reflect the latest policies and inventory data without requiring frequent model retraining. Which approach should they use?

A.Fine-tune a foundation model on historical customer service transcripts and redeploy it weekly.
B.Use retrieval-augmented generation to ground responses in live policy documents and inventory APIs.
C.Pre-train a custom foundation model from scratch using the retailer's historical order database.
D.Increase the model's temperature setting so it can generate more varied and current answers.
AnswerB

Retrieval-augmented generation separates knowledge from the model by fetching current documents and API data at inference time. This keeps answers accurate as policies and inventory change, with no retraining required. It directly satisfies the requirement for up-to-date responses while reducing operational overhead. This is the recommended Google Cloud pattern for dynamic, grounded enterprise assistants.

Why this answer

Grounding responses in live enterprise data through retrieval-augmented generation keeps a generative AI assistant accurate as policies and inventory change, without repeated model retraining. Fine-tuning and pre-training embed static knowledge that goes stale, and temperature adjustments affect style rather than factual currency. The retrieval pattern is the standard Google Cloud approach for assistants that must answer from authoritative, frequently updated sources.

Exam trap

The trap here is assuming that fine-tuning or a larger model will keep answers current, when freshness actually depends on retrieving live source data at inference time.

10
MCQmedium

A financial institution wants to ensure compliance with GDPR when using a generative AI service that processes EU user data. Which measure is most directly required?

A.Disable logging of all interactions to minimize data retention
B.Implement a mechanism to obtain user consent before processing data
C.Store all prompts and responses in a US-based data center
D.Use a model trained only on non-EU data
AnswerB

GDPR requires a lawful basis for processing personal data, and consent is one such basis. Obtaining user consent before processing EU user data directly satisfies that obligation, making it the measure most directly required for the described generative AI service.

Why this answer

GDPR requires a lawful basis for processing personal data, and consent is one of the most common bases, especially for marketing or non-contractual processing. Implementing a mechanism to obtain user consent before processing EU user data directly satisfies GDPR Article 6 and Article 7 requirements for lawful processing. This is the most directly required measure among the options.

Exam trap

The trap is focusing on data residency or training data origin as the primary GDPR requirement, when the exam expects recognition that a lawful basis — most commonly consent — is the foundational obligation for processing EU personal data.

How to eliminate wrong answers

Option A is wrong because disabling logging is not a GDPR requirement and may conflict with accountability obligations; GDPR mandates data minimization but not elimination of all logs. Option C is wrong because storing EU data in a US data center can violate GDPR transfer restrictions (Schrems II) unless adequate safeguards like SCCs are in place — it is not a required measure and is often prohibited. Option D is wrong because GDPR does not require models trained only on non-EU data; it governs processing of EU data subjects regardless of training data origin.

11
MCQhard

A government agency is deploying a generative AI chatbot to answer citizen questions about public services. The chatbot must provide accurate and consistent information, scale to handle peak loads during tax season, and comply with strict data sovereignty laws that require all data to stay within the country. The agency has a moderate budget and in-house IT team but limited AI expertise. Which deployment architecture should they choose?

A.Build and host the model on-premises using open-source tools
B.Deploy a pre-trained model on Vertex AI in the required region with auto-scaling
C.Deploy the model on Vertex AI across multiple regions for availability
D.Use a third-party managed generative AI service that guarantees data residency
AnswerB

Keeps data within region, auto-scales, and requires minimal AI expertise.

Why this answer

Deploying a pre-trained model on Vertex AI in the required region is the best architecture because it keeps data in-country, auto-scales for peak demand, and uses a managed service that minimizes AI operations burden. Option A is not suitable because on-premises hosting requires significant AI expertise and does not scale easily. Option C is not suitable because multi-region deployment can move data outside the required country and violate data sovereignty.

Option D is not the best choice: while a third-party managed service may claim data residency, it adds a separate vendor relationship, potential integration and cost overhead, and less direct control over regional infrastructure and compliance than Vertex AI's explicit in-region deployment.

12
MCQmedium

A media company uses a generative AI model to create marketing images. They want to ensure that AI-generated images can be identified as synthetic. Which Google Cloud capability should they use?

A.SynthID
B.Cloud Vision API
C.Data Loss Prevention (DLP) API
D.Vertex AI Model Registry
AnswerA

SynthID embeds imperceptible digital watermarks directly into AI-generated image pixels, enabling later detection that content is synthetic. This satisfies the media company's requirement to identify generated marketing images as artificial, since the watermark survives common edits and transformations that would defeat metadata-based labelling approaches.

Why this answer

SynthID is Google DeepMind's watermarking technology that embeds imperceptible digital watermarks into AI-generated content, including images, and provides a detector to identify them as synthetic. It is purpose-built for provenance and identification of generative AI output, matching the media company's requirement.

Exam trap

The trap here is confusing content-analysis services (Cloud Vision, DLP) with provenance/watermarking capabilities — candidates often assume any Google Cloud AI API can 'detect AI images,' but only SynthID is designed for that purpose.

How to eliminate wrong answers

Option B is wrong because Cloud Vision API performs image analysis (labels, OCR, faces) but does not embed or detect AI-generation watermarks. Option C is wrong because DLP API is for discovering and redacting sensitive data such as PII, not for marking synthetic media. Option D is wrong because Vertex AI Model Registry is a catalog for managing model versions and metadata, not a watermarking or provenance capability.

13
Multi-Selecthard

A government agency is procuring a generative AI solution for public service information. They require that the system can provide explanations for its answers, protect citizen privacy, and comply with data residency laws. Which three capabilities should they mandate? (Choose three.)

Select 3 answers
A.Integration with social media APIs
B.Ability to generate images from text descriptions
C.Support for grounding with citations to source documents
D.Differential privacy during model training or fine-tuning
E.Configurable data residency controls to keep data within specific regions
AnswersC, D, E

Grounding with citations ties each generated answer to verifiable source documents, giving the explanation capability the agency demands. Reviewers can trace claims back to authoritative public-service content, satisfying the transparency requirement without exposing personal data.

Why this answer

Option C (Support for grounding with citations to source documents) is correct because grounding ties generated answers to verifiable source documents and returns citations, which directly satisfies the requirement that the system can explain the basis of its answers and lets citizens and auditors trace claims back to authoritative public-service content. Option D (Differential privacy during model training or fine-tuning) is correct because differential privacy adds calibrated noise so that the model learns statistical patterns without memorizing or leaking any individual citizen's data, which is the standard technical mechanism for protecting citizen privacy in AI systems. Option E (Configurable data residency controls to keep data within specific regions) is correct because data residency laws require that data be stored and processed within designated jurisdictions, and configurable regional controls (for example, pinning storage, training, and inference to in-country regions) are how that legal constraint is enforced.

Option A (Integration with social media APIs) is not required for explainability, privacy, or residency and would in fact broaden data collection and privacy risk, so it does not belong. Option B (Ability to generate images from text descriptions) is an unrelated multimodal feature that addresses none of the stated requirements for explanations, privacy, or data residency.

Exam trap

The trap is selecting irrelevant capabilities like social media integration or image generation, which do not address the core requirements of explainability, privacy, and data residency.

14
MCQhard

A healthcare startup wants to generate synthetic patient notes for training medical residents. They need the output to follow a strict template with sections: Chief Complaint, History, Assessment, Plan. Which prompt engineering strategy should they use to ensure consistent structure?

A.Use a zero-shot prompt asking for a patient note in bullet points
B.Provide a few-shot example of the template filled out and instruct the model to follow that format for new cases
C.Set temperature to 0 and max tokens to a high value
D.Use a system prompt that lists the sections as instructions
AnswerB

A few-shot example demonstrates the exact section headings and ordering, anchoring generation to the required template. This constrains structure more reliably than zero-shot instructions alone, satisfying the stem's demand for consistent Chief Complaint, History, Assessment and Plan sections.

Why this answer

Few-shot prompting provides the model with concrete examples of the desired output structure, which is the most reliable way to enforce a strict multi-section template like Chief Complaint, History, Assessment, and Plan. By showing one or more filled-out examples, the model learns the exact section headers, ordering, and formatting conventions to replicate for new cases. This pattern-based conditioning is far more effective than abstract instructions alone for template adherence.

Exam trap

The trap here is assuming that listing sections in a system prompt is equivalent to few-shot examples — the exam tests whether you understand that concrete demonstrations outperform abstract instructions for strict structural adherence.

How to eliminate wrong answers

Option A is wrong because a zero-shot bullet-point request gives the model no structural anchor — it may produce inconsistent sections, reorder them, or omit headings entirely. Option C is wrong because temperature and max tokens control randomness and length, not structural format; low temperature makes output more deterministic but does not teach the model the template. Option D is wrong because a system prompt listing section names is weaker than few-shot examples — the model may still vary ordering, add extra sections, or change heading styles without concrete demonstrations.

15
MCQeasy

A startup is using Vertex AI to build a generative AI application. They need to ensure that the AI-generated content does not contain hate speech or violence. Which service should they use?

A.Google's safety filters in Vertex AI
B.Datasheets for Datasets
C.SynthID
D.Model Cards
AnswerA

Google's safety filters in Vertex AI apply configurable content thresholds that block hate speech and violence before responses reach users, directly satisfying the requirement to prevent harmful generated content. They integrate natively with Vertex AI models, so no separate moderation pipeline is needed.

Why this answer

Google's safety filters in Vertex AI are specifically designed to block harmful content such as hate speech and violence by evaluating model inputs and outputs against predefined safety categories. These filters operate at the API level, allowing developers to configure thresholds for blocking sensitive content, making them the direct solution for the startup's requirement to prevent AI-generated hate speech or violence.

Exam trap

Google often tests the distinction between documentation tools (Datasheets for Datasets, Model Cards) and active runtime safety mechanisms (safety filters), leading candidates to confuse transparency artifacts with operational guardrails.

How to eliminate wrong answers

Option B (Datasheets for Datasets) is wrong because it is a documentation framework for dataset transparency, not a runtime content moderation tool; it describes dataset characteristics but does not filter generated outputs. Option C (SynthID) is wrong because it is a watermarking technique for AI-generated images, not a safety filter for text content; it identifies synthetic media but does not block hate speech or violence. Option D (Model Cards) is wrong because they are standardized model documentation sheets that report model performance and limitations, not active filtering mechanisms; they inform users about model behavior but do not enforce content safety at inference time.

16
MCQeasy

A developer wants to experiment with Gemini Nano for on-device inference in a mobile app. Which Gemini API tier or environment provides access to Gemini Nano?

A.Vertex AI
B.Google AI Studio
C.Gemini API directly
D.MediaPipe and Android AICore
AnswerD

Gemini Nano runs on-device rather than through a cloud endpoint, so it is accessed via MediaPipe and Android AICore on supported Android hardware. This satisfies the stem's constraint of on-device inference within a mobile app, which the cloud Gemini API tiers cannot provide.

Why this answer

Gemini Nano is the smallest Gemini model designed specifically for on-device inference on mobile devices. Access to Gemini Nano is provided through MediaPipe and Android AICore, which are the frameworks that enable running the model locally on Android devices without requiring a network connection to Google's cloud servers.

Exam trap

A common pitfall in this exam is assuming all Gemini models are accessed through the same cloud API (Vertex AI, AI Studio, or Gemini API). Gemini Nano is specifically designed for on-device inference and is accessed via MediaPipe and Android AICore on mobile devices.

How to eliminate wrong answers

Option A is wrong because Vertex AI is a cloud-based platform for deploying and managing machine learning models on Google Cloud, not for on-device inference on mobile apps. Option B is wrong because Google AI Studio is a web-based tool for prototyping and testing prompts with Gemini models in the cloud, not for on-device execution. Option C is wrong because the Gemini API is a cloud API that requires network connectivity to send requests to Google's servers, whereas Gemini Nano runs entirely on the device.

17
Multi-Selectmedium

A company is fine‑tuning a large language model for a domain‑specific task. They have a limited budget and want to minimize the cost of fine‑tuning. Which TWO approaches are most cost‑effective?

Select 2 answers
A.Use a smaller base model like Gemini Flash
B.Use full fine‑tuning for better quality
C.Increase the number of training epochs
D.Use adapter‑based fine‑tuning (LoRA)
E.Use a larger base model like Gemini Ultra
AnswersA, D

Choosing a smaller base model such as Gemini Flash directly reduces fine-tuning cost, since compute and memory scale with parameter count. It satisfies the stem's limited-budget constraint by training fewer weights, while still adapting adequately to a narrow domain-specific task where a larger model's extra capacity adds little value.

Why this answer

Option A is correct because using a smaller base model such as Gemini Flash reduces compute and memory requirements, which directly lowers the cost of fine-tuning while still being adequate for many domain-specific tasks. Option D is correct because adapter-based fine-tuning with LoRA freezes the base model weights and trains only small low-rank adapter matrices, dramatically reducing the number of trainable parameters, GPU memory, and training time compared to full fine-tuning. Option B is not cost-effective because full fine-tuning updates all model parameters, requiring substantially more compute, memory, and storage.

Option C is not cost-effective because increasing training epochs raises compute time and cost without guaranteeing better results, and can even cause overfitting. Option E is not cost-effective because a larger base model like Gemini Ultra increases training and inference costs significantly, which conflicts with the limited-budget goal.

Exam trap

Candidates often think that larger models or full fine-tuning always yield better results, but the trap here is that cost-effectiveness prioritizes resource efficiency over raw quality, and adapter methods like LoRA provide a practical trade-off that they overlook.

18
MCQmedium

A healthcare organization is deploying a generative AI application that processes Protected Health Information (PHI). They must ensure compliance with HIPAA. Which Google Cloud offering should they use?

A.Google Workspace with Duet AI
B.Model Garden open-source models deployed on Compute Engine
C.Gemini API with default settings
D.Vertex AI APIs with data residency and HIPAA compliance enabled
AnswerD

Vertex AI APIs with data residency and HIPAA compliance enabled provide a Business Associate Agreement and keep PHI within approved regions, satisfying HIPAA's safeguards for protected health information. Generic AI services lacking these contractual and residency controls cannot lawfully process PHI.

Why this answer

Vertex AI APIs with data residency and HIPAA compliance enabled is the correct choice because it meets regulatory requirements for PHI. Other options either lack HIPAA coverage or introduce unnecessary complexity.

19
MCQeasy

A retail bank uses a generative AI assistant to answer customer questions about account policies. Compliance requires that every response cite the specific internal policy document section it used. Which approach best enforces this requirement?

A.Fine-tune the model on the bank's policy manuals so it memorizes the sections.
B.Use retrieval-augmented generation to fetch policy passages and instruct the model to cite them.
C.Add a system instruction telling the model to never guess and to be accurate.
D.Lower the model's temperature to zero so responses are deterministic.
AnswerB

Retrieval-augmented generation grounds responses in retrieved internal documents and lets the prompt require a citation to the fetched section. This directly satisfies the compliance rule because the model answers from authoritative policy text rather than parametric memory. It also makes citations verifiable, since each claim maps to a retrievable passage the bank controls.

Why this answer

A citation mandate is a grounding problem, not a sampling or behavior problem. Retrieval-augmented generation supplies authoritative policy passages at inference time and lets the prompt demand a citation for each. Temperature tuning, generic instructions, and fine-tuning either do not provide sources or cannot guarantee verifiable, current section references.

Exam trap

The trap here is believing that a strong instruction to be accurate or a low temperature setting produces verifiable citations, when only grounding the response in retrieved source documents can do that.

20
MCQmedium

A media company wants to automatically generate captions for video content in multiple languages. The captions should be synced with the audio timeline. Which combination of Google Cloud services is most appropriate?

A.Cloud Video Intelligence API and Translation API
B.Document AI and Translation API
C.Speech-to-Text and Translation API
D.Text-to-Speech and Translation API
AnswerC

Speech-to-Text transcribes the audio with timestamps, and the Translation API renders that text into other languages, satisfying the multilingual captioning requirement. The timestamps preserve audio-timeline sync, which a generic translation service alone could not provide.

Why this answer

The workflow requires first transcribing the audio track into text using Speech-to-Text API, which provides timestamps for each word or phrase to ensure synchronization with the video timeline. The resulting transcript is then passed to the Translation API to generate captions in the target languages, preserving the original timing data for alignment.

Exam trap

Candidates often mistakenly choose Cloud Video Intelligence API for captioning tasks because it can analyze video content, but caption generation requires audio transcription via Speech-to-Text API, not just visual analysis.

How to eliminate wrong answers

Option A is wrong because Cloud Video Intelligence API analyzes video content (objects, scenes, explicit content) but does not transcribe audio; it cannot generate captions from speech. Option B is wrong because Document AI is designed for extracting and processing text from documents (PDFs, invoices), not for transcribing audio or video content. Option D is wrong because Text-to-Speech API converts text into spoken audio, which is the reverse of the required workflow; it cannot transcribe existing audio into text for captioning.

21
MCQhard

A financial services firm is developing a GenAI application for investment advice. They need to ensure regulatory compliance. Which business strategy should they prioritize?

A.Rapidly deploy an MVP and iterate based on user feedback
B.Implement strict human-in-the-loop review for all investment recommendations
C.Open-source the model to gain community trust
D.Partner with a cloud provider that offers indemnification for model outputs
AnswerB

Human-in-the-loop review ensures qualified staff verify every recommendation before it reaches clients, satisfying the regulatory compliance constraint for investment advice. This oversight mitigates the risk of unverified GenAI output causing unsuitable or non-compliant financial guidance.

Why this answer

In regulated industries like financial services, GenAI applications must prioritize compliance over speed. Option B is correct because a human-in-the-loop (HITL) review ensures that every investment recommendation is auditable and meets regulatory standards (e.g., SEC or FINRA rules), mitigating risks of hallucinated or non-compliant outputs. This strategy directly addresses the need for accountability and transparency in high-stakes decision-making.

Exam trap

Google Cloud often tests the misconception that speed or technical features (like open-sourcing or indemnification) can substitute for regulatory compliance, but in regulated domains, human oversight and auditability are non-negotiable.

How to eliminate wrong answers

Option A is wrong because rapidly deploying an MVP without rigorous compliance checks risks generating non-compliant or misleading investment advice, which could lead to severe regulatory penalties and loss of client trust. Option C is wrong because open-sourcing the model does not inherently ensure regulatory compliance; it may expose proprietary data or create liability if the model produces biased or inaccurate outputs, and community trust does not substitute for legal adherence. Option D is wrong because cloud provider indemnification covers legal costs for model outputs but does not prevent the generation of non-compliant advice; it is a risk transfer mechanism, not a compliance strategy.

22
MCQeasy

A data scientist is using Vertex AI generative AI studio to create a chatbot. The chatbot gives inconsistent answers to similar questions. Which parameter should they adjust to make responses more consistent?

A.Decrease temperature to 0.2
B.Increase top-p to 0.9
C.Increase presence penalty to 0.5
D.Decrease frequency penalty to 0.0
AnswerA

Lowering temperature to 0.2 sharpens the probability distribution over the next token, so the model favours its highest-likelihood continuations rather than sampling broadly. That directly addresses the inconsistent answers described in the stem, since similar prompts then converge on near-identical outputs.

Why this answer

Decreasing the temperature to 0.2 reduces the randomness of the model's token sampling, making the output more deterministic and consistent. Temperature controls the probability distribution over tokens; lower values make the model more likely to choose the highest-probability token, reducing variability in responses to similar questions.

Exam trap

Google Cloud often tests the misconception that increasing top-p or adjusting penalties improves consistency, when in fact temperature is the primary parameter for controlling output determinism.

How to eliminate wrong answers

Option B is wrong because increasing top-p to 0.9 increases the cumulative probability threshold for token sampling, which actually introduces more diversity and randomness, making responses less consistent. Option C is wrong because increasing presence penalty to 0.5 penalizes tokens that have already appeared, encouraging the model to introduce new topics and variability, which reduces consistency. Option D is wrong because decreasing frequency penalty to 0.0 removes the penalty for token repetition, which can lead to repetitive or stuck responses but does not directly control the randomness of token selection; consistency is primarily governed by temperature, not frequency penalty.

23
Multi-Selecthard

A team is fine-tuning a large language model for medical advice. Which TWO techniques are most effective for improving the safety and reliability of the model's outputs?

Select 2 answers
A.Constitutional AI
B.Lowering the temperature to 0.0
C.Increasing training data size
D.Increasing top_p to 1.0
E.Reinforcement learning from human feedback (RLHF)
AnswersA, E

Constitutional AI uses predefined rules to guide model behavior.

Why this answer

Constitutional AI (A) is correct because it embeds a set of ethical principles directly into the model's training process, allowing the model to self-critique and revise its outputs to avoid harmful or unsafe medical advice. This technique proactively enforces safety constraints without requiring extensive human labeling, making it highly effective for high-stakes domains like healthcare.

Exam trap

Google Cloud often tests the misconception that hyperparameter tuning (temperature, top_p) or data scaling alone can solve safety issues, when in fact alignment techniques like Constitutional AI and RLHF are specifically designed for that purpose.

24
Multi-Selecteasy

Which TWO of the following are key differences between generative AI and discriminative AI? (Choose two.)

Select 2 answers
A.Generative models can create new data samples, while discriminative models only assign labels to existing data.
B.Generative models require less training data than discriminative models.
C.Generative models cannot be used for supervised learning tasks like classification.
D.Generative models model the joint probability distribution of inputs and labels, whereas discriminative models model the conditional probability of labels given inputs.
E.Discriminative models always outperform generative models on tasks like image classification.
AnswersA, D

The defining axis is output behaviour: generative models learn the joint distribution to produce novel samples, whereas discriminative models learn decision boundaries mapping inputs to labels. This distinction separates creation of new data from classification of existing data.

Why this answer

Option A is correct because generative AI learns the underlying data distribution so it can produce novel samples (e.g., images, text, audio), whereas discriminative AI learns only a decision boundary or mapping from inputs to labels and therefore just classifies or predicts labels for existing data. Option D is correct because, mathematically, generative models estimate the joint probability P(x, y) (or P(x) for unsupervised generation), allowing them to sample new data, while discriminative models estimate the conditional probability P(y | x) directly to separate classes. Option B is wrong because generative models typically require large amounts of training data, often more than discriminative models, to capture the full data distribution.

Option C is wrong because generative models can be adapted to supervised tasks such as classification (e.g., using class-conditional likelihoods or fine-tuning), so they are not inherently unusable for classification. Option E is wrong because no model class universally outperforms the other; performance depends on the task, data size, and architecture, and discriminative models often excel at classification while generative models excel at synthesis.

Exam trap

Google Cloud often tests the misconception that generative models are only for unsupervised tasks and cannot perform classification, leading candidates to incorrectly select Option C, while also testing the false assumption that discriminative models are universally superior, as in Option E.

25
MCQmedium

Despite applying safety filters, a generative AI model still produces toxic outputs in some cases. Which additional technique should be applied?

A.Add more examples of toxic content to training
B.Increase the filter threshold
C.Use RLHF with human feedback to reduce toxicity
D.Decrease the model's temperature
AnswerC

RLHF fine-tunes the model using human rankings of outputs, directly penalising toxic responses and steering generation toward safer behaviour. This addresses residual toxicity that static safety filters miss, satisfying the stem's need for an additional mitigation beyond filtering.

Why this answer

RLHF (Reinforcement Learning from Human Feedback) directly addresses toxicity by using human evaluators to rank model outputs, then fine-tuning the model to prefer less toxic responses. This technique teaches the model to avoid harmful patterns that safety filters might miss, as filters are static and can be bypassed by adversarial prompts or nuanced toxicity.

Exam trap

A common misconception is that adjusting static parameters (like temperature or filter thresholds) can solve alignment problems, when in fact dynamic human-in-the-loop methods like RLHF are required for nuanced safety issues.

How to eliminate wrong answers

Option A is wrong because adding more examples of toxic content to training would likely reinforce those patterns, increasing rather than decreasing toxicity. Option B is wrong because increasing the filter threshold would make the filter less sensitive, allowing more toxic content through instead of blocking it. Option D is wrong because decreasing the model's temperature reduces randomness in output but does not specifically target or mitigate toxic content generation.

26
MCQmedium

A company is building a generative AI chatbot for customer support using Vertex AI. They want to ground the model responses with their internal knowledge base stored in Cloud Storage and BigQuery. Which feature should they use to ensure the model only answers from the provided data and avoids hallucination?

A.Vertex AI Grounding with Vertex AI Search
B.Vertex AI Prediction
C.Vertex AI Pipelines
D.Cloud Functions
AnswerA

Vertex AI Grounding with Vertex AI Search connects the model to your Cloud Storage and BigQuery repositories, retrieving relevant passages before generation. This retrieval-augmented approach constrains responses to indexed enterprise content, satisfying the requirement that answers derive solely from the internal knowledge base and reducing hallucination.

Why this answer

Vertex AI Grounding with Vertex AI Search is the correct feature because it allows the model to retrieve and cite information from a specified data source (such as Cloud Storage and BigQuery) to generate responses. This process, known as grounding, ensures the model's output is based solely on the provided authoritative data, effectively reducing hallucinations by constraining the model to factual, retrieved content rather than relying on its internal parametric knowledge.

Exam trap

The trap here is that candidates may confuse Vertex AI Prediction (a general model serving endpoint) with the grounding feature, mistakenly thinking that simply deploying a model with Vertex AI Prediction will automatically restrict its answers to a specific knowledge base, when in fact grounding requires explicit integration with Vertex AI Search and a configured data store.

How to eliminate wrong answers

Option B is wrong because Vertex AI Prediction is a service for deploying and serving models to generate predictions or responses, but it does not inherently include grounding capabilities to restrict answers to a specific knowledge base; it would require additional integration with a retrieval system. Option C is wrong because Vertex AI Pipelines is an orchestration service for building and managing ML workflows, not a feature for grounding model responses or preventing hallucinations. Option D is wrong because Cloud Functions is a serverless compute service for running event-driven code, and while it could be used to build a custom retrieval pipeline, it is not a native Vertex AI feature for grounding and does not provide the built-in retrieval and citation mechanisms needed to ensure answers come only from the provided data.

27
Multi-Selecteasy

Which TWO features are available in Vertex AI Studio for prompt engineering? (Choose two.)

Select 2 answers
A.Side-by-side comparison of model outputs
B.One-click deployment to a Vertex AI endpoint
C.Ability to test prompts with different model parameters (temperature, top_p)
D.Fine-tuning models directly in the interface
E.Building conversational agents with drag-and-drop
AnswersA, C

Allows output comparison.

Why this answer

Vertex AI Studio provides a side-by-side comparison feature that allows prompt engineers to evaluate outputs from multiple model configurations or parameter settings simultaneously. This enables direct visual comparison of responses, helping to identify the most effective prompt phrasing or parameter combination without manual switching.

Exam trap

The trap here is that candidates may confuse Vertex AI Studio's prompt engineering features with those of Vertex AI Agent Builder or Vertex AI Model Registry, leading them to select options like one-click deployment or drag-and-drop agent building that belong to separate services.

28
Multi-Selecteasy

A team is selecting a foundation model for a text summarization use case. They need to consider factors that affect both model performance and production deployment. Which THREE factors are most critical? (Choose three.)

Select 3 answers
A.Model parameter count (billions of parameters).
B.Inference latency and throughput capabilities.
C.Context window length (maximum input tokens).
D.Training data provenance and licensing.
E.Pricing per token (input + output).
AnswersB, C, E

Inference latency and throughput are critical for production deployment because they directly determine user experience and operational cost. Low latency is essential for real-time summarization, and high throughput enables handling of concurrent requests efficiently.

Why this answer

Inference latency and throughput are critical for production deployment because they directly determine the user experience and operational cost. A model with high latency may be unsuitable for real-time summarization, while low throughput limits the number of concurrent requests the system can handle, affecting scalability and cost-efficiency.

Exam trap

Google Cloud often tests the distinction between model-centric factors (like parameter count) and deployment-centric factors (like latency and pricing), trapping candidates who assume bigger models are always better without considering operational constraints.

29
MCQmedium

A company wants to scale their generative AI application globally with low latency. Which infrastructure configuration is most suitable?

A.Use a CDN to cache responses.
B.Multiple regional endpoints with traffic routing to the nearest region.
C.On-premises deployment for all regions.
D.Single endpoint in us-central1 with high max replicas.
AnswerB

Deploying the model to multiple regional endpoints and routing each user to the nearest region reduces network distance, satisfying the low-latency requirement for global users. A single-region endpoint would force distant users to traverse long network paths.

Why this answer

Deploying multiple regional endpoints with traffic routing to the nearest region minimizes latency by directing user requests to the geographically closest inference endpoint. This architecture leverages global load balancing (e.g., using Anycast DNS or HTTP(S) load balancers with backend services in multiple regions) to reduce round-trip time (RTT) and meet latency SLAs for real-time generative AI applications.

Exam trap

The trap here is that candidates often confuse CDN caching with real-time inference, assuming caching can accelerate dynamic AI responses, but generative AI outputs are unique per request and cannot be pre-cached.

How to eliminate wrong answers

Option A is wrong because a CDN caches static content (e.g., images, CSS) but cannot cache dynamic, context-dependent generative AI responses, which require real-time model inference; thus, it does not reduce latency for API calls. Option C is wrong because on-premises deployment lacks global scalability and introduces high latency for users outside the local region, defeating the purpose of global low-latency access. Option D is wrong because a single endpoint in us-central1 forces all global traffic to traverse long distances, causing high latency for users far from that region, regardless of the number of replicas.

30
MCQeasy

Which Google Cloud product provides access to pre-trained foundation models like Gemini?

A.Dataflow
B.Vertex AI Generative AI Studio
C.Cloud Translation
D.Vertex AI Model Registry
AnswerB

Vertex AI Generative AI Studio exposes pre-trained foundation models, including Gemini, through a managed Google Cloud interface for prompting, tuning and deployment. It directly satisfies the requirement for a Google Cloud product offering access to such models.

Why this answer

Vertex AI Generative AI Studio is the correct answer because it is the Google Cloud service specifically designed to provide access to pre-trained foundation models like Gemini, allowing users to test, customize, and deploy them via a managed interface. Unlike other services, Generative AI Studio directly integrates with Gemini's API and offers prompt engineering, tuning, and model evaluation capabilities.

Exam trap

The trap here is that candidates confuse Vertex AI Model Registry (a model management tool) with Generative AI Studio (the actual interface for accessing and experimenting with foundation models), leading them to pick D instead of B.

How to eliminate wrong answers

Option A is wrong because Dataflow is a fully managed stream and batch data processing service based on Apache Beam, not a platform for accessing or interacting with pre-trained foundation models. Option C is wrong because Cloud Translation is a specialized service for language translation using pre-trained models, but it does not provide access to general-purpose foundation models like Gemini or support for multimodal tasks. Option D is wrong because Vertex AI Model Registry is a metadata management service for storing and versioning models, not a tool for directly accessing or experimenting with pre-trained foundation models like Gemini.

31
MCQmedium

A marketing team is using Vertex AI's text generation model to create product descriptions. They want to control the randomness of the output to ensure consistent, focused messaging for a new product line. Which parameter should they adjust to reduce randomness and make the output more deterministic?

A.Max output tokens
B.Temperature
C.Top-k
D.Top-p
AnswerB

Temperature controls the randomness of predictions by scaling the logits before applying softmax. A lower temperature (e.g., 0.2) makes the model more confident and deterministic, producing focused outputs. This directly addresses the need for consistent messaging. Higher temperatures increase diversity but reduce consistency.

Why this answer

Temperature is the parameter that directly controls the randomness of the model's output. Lowering the temperature makes the model more likely to choose high-probability tokens, resulting in more deterministic and consistent text. This is exactly what the marketing team needs for focused product descriptions.

Exam trap

The trap here is confusing top-k or top-p with temperature; while they influence sampling, temperature is the primary control for randomness.

32
MCQmedium

A financial services firm is using a generative AI model to answer customer queries about account balances. The model sometimes provides outdated information because it relies on its training data. The firm wants to ensure the model always uses the most current account data. Which technique should they use?

A.Fine-tune the model daily with the latest account data.
B.Use retrieval-augmented generation (RAG) to fetch current account data from a database.
C.Increase the model's context window to include more historical data.
D.Apply prompt engineering to instruct the model to use the latest data.
AnswerB

RAG combines a generative model with a retrieval system that fetches relevant, up-to-date information from external sources, such as a database. This ensures the model's responses are based on the most current data without retraining. It is ideal for scenarios requiring real-time or frequently updated information.

Why this answer

Retrieval-augmented generation (RAG) is designed to ground model outputs in external, up-to-date knowledge. By integrating a retriever that queries the firm's database for current account balances, the model can generate responses that reflect the latest information. This approach avoids the cost and latency of frequent fine-tuning and ensures accuracy for time-sensitive queries.

Exam trap

The trap here is assuming that the model can access real-time data through its training or by simply instructing it to do so, when in fact external retrieval is necessary to provide current information.

33
MCQmedium

A compliance officer requires that all AI-generated content in Google Workspace be reviewed before sharing externally. Which change management approach BEST supports this requirement while maintaining user adoption?

A.Roll out the feature to a pilot group with training and a feedback loop before company-wide deployment
B.Allow all sharing but audit logs after the fact
C.Disable AI features in Workspace for all users until a review tool is built
D.Immediately block all external sharing of AI-generated content
AnswerA

A pilot with training and a feedback loop surfaces review-workflow friction on a small scale, letting the compliance requirement for pre-sharing review be validated before company-wide rollout. This preserves user adoption by incorporating feedback rather than imposing controls abruptly.

Why this answer

A phased pilot with training and a feedback loop lets the compliance team validate that AI-generated content is reviewed before external sharing while giving users time to adapt, which preserves adoption. This change management approach balances governance with user enablement, reducing resistance and surfacing issues before broad rollout.

Exam trap

The trap here is conflating 'preventive control' with 'post-hoc audit' — candidates pick auditing or hard blocking because they sound strict, but the question requires both compliance enforcement and maintained user adoption, which only a phased pilot with training achieves.

How to eliminate wrong answers

Option B is wrong because post-hoc auditing does not satisfy the requirement that content be reviewed before external sharing — it is reactive, not preventive. Option C is wrong because disabling AI features entirely blocks productivity and destroys user adoption, which the question explicitly says must be maintained. Option D is wrong because immediately blocking all external sharing of AI-generated content is a blunt control that disrupts legitimate workflows and does not include the training or feedback needed for adoption.

34
MCQhard

A research team is using a large language model to analyze medical research papers and generate summaries. They need to minimize hallucinations while retaining key details. They have access to a curated database of paper abstracts. Which approach is best?

A.Fine-tune the model on the entire database of papers.
B.Use chain-of-thought prompting to reason step-by-step.
C.Use few-shot prompting with examples of accurate summaries and set temperature=0.0.
D.Implement RAG to retrieve relevant abstracts and incorporate them into the prompt.
AnswerD

RAG grounds generation in the curated abstracts by retrieving relevant passages and inserting them into the prompt, so the model summarises supplied evidence rather than relying on parametric memory. This directly minimises hallucination while retaining key details, satisfying the accuracy constraint.

Why this answer

Retrieval-Augmented Generation (RAG) directly addresses hallucination by grounding the model's output in a curated database of paper abstracts. By retrieving relevant abstracts and injecting them into the prompt, the model generates summaries based on verified facts rather than relying solely on its parametric knowledge, which is the most effective way to minimize hallucinations while retaining key details.

Exam trap

Many candidates mistakenly think that fine-tuning or low temperature alone can solve hallucination, but the trap here is that without external retrieval (RAG), the model has no mechanism to verify facts against a trusted source, so it will still generate plausible-sounding but incorrect details.

How to eliminate wrong answers

Option A is wrong because fine-tuning on the entire database of papers does not prevent hallucinations; it can cause catastrophic forgetting and the model may still fabricate details when asked to summarize unseen or edge-case content. Option B is wrong because chain-of-thought prompting improves reasoning but does not provide external factual grounding, so the model can still hallucinate based on its internal knowledge. Option C is wrong because few-shot prompting with temperature=0.0 reduces randomness but does not supply the model with the actual abstracts to reference; it relies on the model's memory of the examples, which can lead to hallucinated details not present in the source papers.

35
MCQhard

A financial services company needs to deploy an AI model that handles highly sensitive transaction data. They require that the model's predictions cannot be inspected by any third party, and the data must remain encrypted at all times, including during inference. Which Google Cloud feature should they use?

A.Customer-Managed Encryption Keys (CMEK)
B.Access Transparency logs
C.VPC Service Controls
D.Confidential VMs
AnswerD

Confidential VMs encrypt memory with hardware-based keys while data is in use, so transaction data and model predictions stay inaccessible to Google or any third party during inference. This satisfies the requirement for always-encrypted data and non-inspectable predictions.

Why this answer

Confidential VMs (D) are the correct choice because they provide hardware-based memory encryption using AMD Secure Encrypted Virtualization (SEV), ensuring that data remains encrypted while in use (during inference). This meets the requirement that the model's predictions cannot be inspected by any third party, including Google Cloud operators, and that data stays encrypted at all times.

Exam trap

The trap here is that candidates often confuse encryption at rest/in transit with encryption in use, and mistakenly choose CMEK or VPC Service Controls, not realizing that only Confidential VMs protect data during active computation.

How to eliminate wrong answers

Option A is wrong because Customer-Managed Encryption Keys (CMEK) protect data at rest and in transit but do not encrypt data during processing (in use), leaving it exposed in memory during inference. Option B is wrong because Access Transparency logs provide audit logs of Google Cloud administrator access but do not encrypt data or prevent third-party inspection of predictions. Option C is wrong because VPC Service Controls create a security perimeter to prevent data exfiltration but do not encrypt data in use; they control network access, not memory-level encryption.

36
MCQeasy

A company wants to generate marketing images for a new product launch using GenAI. Which Google Cloud service should they use?

A.Gemini for Google Workspace
B.Vertex AI Agent Builder
C.Vertex AI Model Garden
D.Vertex AI Imagen
AnswerD

Vertex AI Imagen generates photorealistic images from text prompts, directly satisfying the marketing image requirement. Unlike Vertex AI's text or code models, Imagen is purpose-built for image synthesis, so it produces launch-ready visuals without custom training. This matches the scenario's need for GenAI-generated product marketing imagery.

Why this answer

Imagen on Vertex AI is Google Cloud's image generation model, available through Vertex AI APIs or Model Garden.

37
MCQhard

A company is migrating a GenAI proof-of-concept to production. During the pilot, they used a large model (e.g., Gemini 1.5 Pro) and incurred high costs. The use case is simple: generating short product descriptions from structured data. Which cost optimization strategy should they implement first?

A.Fine-tune a smaller model on the specific task
B.Implement batch processing to group requests
C.Reduce the model's temperature to 0.0
D.Switch to a smaller model like Gemini 1.5 Flash and use structured prompts
AnswerD

Gemini 1.5 Flash handles short structured generation at far lower cost per token than Pro, and structured prompts reduce token overhead further. Since the task is simple, this preserves output quality while directly addressing the high-cost constraint identified during the pilot.

Why this answer

The primary cost driver in this scenario is the model size itself. Since the use case is simple (generating short product descriptions from structured data), a smaller model like Gemini 1.5 Flash can handle the task with significantly lower inference cost per token. Structured prompts further optimize by reducing token waste and ensuring consistent output, making this the most direct and impactful first step.

Exam trap

Google often tests the misconception that fine-tuning is the first step for any production optimization, when in reality, model selection and prompt engineering are cheaper and faster to implement for simple tasks.

How to eliminate wrong answers

Option A is wrong because fine-tuning a smaller model is a valid long-term optimization but introduces upfront cost and complexity (data preparation, training compute, and validation) that is unnecessary for a simple task that can be handled by a smaller model with prompt engineering. Option B is wrong because batch processing reduces per-request overhead but does not address the fundamental cost per token of the large model; the savings are marginal compared to switching to a cheaper model. Option C is wrong because reducing temperature to 0.0 only affects output randomness and token selection, not the model's size or inference cost; it may improve determinism but does not reduce the number of parameters or compute required per request.

38
MCQeasy

A retail company wants to use generative AI to generate product descriptions for thousands of items. They need to ensure that the descriptions are consistent with their brand voice and do not contain factual inaccuracies. What is the most effective strategy?

A.Use a rule-based system to generate descriptions from product attributes.
B.Fine-tune a model on historical product descriptions and use prompt engineering with brand guidelines.
C.Use a large language model with no safety filters to maximize output variety.
D.Use a pre-trained model without any customization and rely on post-processing filters.
AnswerB

Fine-tuning on historical descriptions embeds the brand voice directly in the model's weights, while prompt engineering injects brand guidelines at inference time to constrain tone. This combination satisfies the consistency requirement across thousands of items, and grounding prompts in approved copy reduces fabricated product claims.

Why this answer

Fine-tuning a model on historical product descriptions aligns the model with the company's specific brand voice and domain language, while prompt engineering with brand guidelines provides explicit guardrails for each generation. This combination ensures consistency and reduces factual inaccuracies by grounding the model in verified examples and structured instructions, which is more effective than rule-based systems or post-processing alone.

Exam trap

Google Gen AI Leader often tests the misconception that post-processing filters or rule-based systems can fully substitute for model customization, when in fact fine-tuning is required to embed brand-specific knowledge into the model's parameters for reliable, consistent generation.

How to eliminate wrong answers

Option A is wrong because rule-based systems lack the linguistic flexibility and contextual understanding of generative AI, often producing rigid, unnatural descriptions that fail to capture nuanced brand voice. Option C is wrong because using a large language model with no safety filters increases the risk of generating factually inaccurate or off-brand content, as there are no constraints to enforce accuracy or style. Option D is wrong because a pre-trained model without customization cannot reliably adhere to a specific brand voice, and relying solely on post-processing filters is insufficient to correct deep-seated factual errors or stylistic inconsistencies.

39
MCQmedium

A marketing team is using a generative AI model to create ad copy. They want to control the randomness of the output so that the same prompt produces consistent results for A/B testing. Which parameter should they adjust?

A.Top-p
B.Temperature
C.Top-k
D.Max output tokens
AnswerB

Temperature scales the probability distribution of the next token. Setting it to a low value, such as 0, makes the model choose the most likely token almost deterministically, leading to consistent outputs for the same prompt. This is ideal for A/B testing where reproducibility is needed. Higher temperatures increase randomness.

Why this answer

Temperature directly controls the randomness of token selection. A temperature of 0 (or very close to 0) makes the model deterministically pick the highest-probability token, ensuring the same prompt yields the same output. This is essential for A/B testing where consistent ad copy variants are required.

Other parameters affect diversity but do not guarantee reproducibility.

Exam trap

The trap here is confusing top-k or top-p with determinism; those parameters narrow the sampling pool but still allow random selection within it, unlike temperature set to zero.

40
MCQhard

A large e-commerce company deploys a generative AI chatbot on Vertex AI for customer service. The chatbot is powered by a fine-tuned model on the company's historical support tickets. Despite high accuracy on training topics, the chatbot frequently gives irrelevant or off-topic answers when customers ask about new products or promotions. The company maintains a comprehensive product catalog and a knowledge base of current promotions. The chatbot's prompts include a system instruction to 'Answer based on your knowledge' and no other retrieval mechanism. The response time requirement is under 3 seconds. Which course of action should the team take?

A.Implement a RAG pipeline that retrieves relevant product and promotion data from the knowledge base and injects it into the prompt.
B.Increase the temperature to encourage the model to generate more diverse answers.
C.Add additional safety filters to block irrelevant responses.
D.Fine-tune the model again on a larger dataset that includes recent support tickets.
AnswerA

Retrieval-augmented generation fetches current product and promotion content from the knowledge base and injects it into the prompt, grounding answers in up-to-date facts the fine-tuned model never saw. This addresses the off-topic responses while keeping latency within the three-second requirement.

Why this answer

Implementing a RAG (Retrieval-Augmented Generation) pipeline directly addresses the chatbot's inability to answer questions about new products or promotions. By retrieving relevant, up-to-date information from the company's product catalog and knowledge base and injecting it into the prompt, the model gains access to current data beyond its training cutoff. This approach keeps response times under 3 seconds (as retrieval is fast) and avoids the need for costly retraining, while the system instruction 'Answer based on your knowledge' is replaced with grounded context.

Exam trap

Google often tests the misconception that fine-tuning alone can solve knowledge gaps for dynamic or time-sensitive data, when in reality RAG is the appropriate technique for incorporating external, frequently updated information without retraining.

How to eliminate wrong answers

Option B is wrong because increasing the temperature would make the model generate more random and diverse outputs, which would worsen the problem of irrelevant or off-topic answers rather than fix it. Option C is wrong because adding safety filters blocks harmful or inappropriate content but does not solve the core issue of the model lacking knowledge about new products or promotions; it would merely suppress irrelevant responses without providing correct information. Option D is wrong because fine-tuning again on a larger dataset that includes recent support tickets is time-consuming, expensive, and still cannot keep up with rapidly changing promotions or new products; the model would remain static after training, whereas RAG provides dynamic, real-time retrieval.

41
MCQmedium

A global retailer wants to deploy a generative AI assistant that answers employee questions about HR policies in English, Spanish, and Japanese. The HR policy documents are updated monthly, and the company wants to avoid retraining the underlying large language model. Which approach should the retailer use?

A.Prepend the entire HR policy corpus to every prompt and rely on the model's context window for accuracy.
B.Fine-tune a foundation model on the translated HR policy documents each month and deploy the tuned model.
C.Increase the model's temperature setting so it can infer policy changes from patterns in previous answers.
D.Use retrieval-augmented generation with a vector index of the current HR policy documents and a foundation model.
AnswerD

Retrieval-augmented generation grounds responses in documents fetched at query time, so monthly policy changes are reflected as soon as the index is refreshed, without retraining the model. It also supports multilingual answers when the model is instructed to respond in the user's language, and it makes source attribution easier for HR compliance reviews.

Why this answer

Retrieval-augmented generation keeps the language model unchanged while supplying current, relevant HR policy passages at query time. That matches the need to avoid retraining, support multiple languages through prompting, and reflect monthly document updates after re-indexing. Fine-tuning, temperature changes, and full-corpus prompting either add recurring training work or fail to guarantee fresh, grounded answers.

Exam trap

The trap here is assuming that fine-tuning is required whenever a generative AI application must reflect organization-specific knowledge, when retrieval is the appropriate pattern for changing factual content.

42
MCQhard

A global corporation with 50,000 employees has seen rapid adoption of GenAI across marketing, product, and engineering teams. Each team selected its own models and cloud accounts, resulting in fragmented governance, unexpected costs, and varying output quality. The CFO demands a unified strategy to control costs and ensure consistency. The Chief AI Officer proposes several solutions. Which course of action best balances control with innovation?

A.Migrate all GenAI workloads to a single on-premises server to reduce cloud costs
B.Establish a GenAI Center of Excellence (CoE) that provides approved models, shared APIs, and best practices, while allowing team-specific customizations
C.Mandate all teams use a single model (e.g., Gemini) via a centralized Vertex AI endpoint with usage quotas
D.Allow teams to continue using their own models but require them to submit monthly cost reports
AnswerB

A CoE supplies approved models, shared APIs and best practices, centralising cost governance and output consistency across all 50,000 employees. Allowing team-specific customisations preserves the innovation autonomy that fragmented adoption previously delivered, satisfying the CFO's control demand without stifling each team.

Why this answer

A GenAI Center of Excellence (CoE) provides centralized governance through approved models and shared APIs, enabling cost control and quality consistency while preserving team-level flexibility for innovation. This balances the CFO's need for unified strategy with the CAIO's goal of avoiding rigid mandates that stifle experimentation.

Exam trap

The tension between centralization and flexibility is a frequent topic in Google Gen AI exams. Candidates often mistakenly choose Option C (single model mandate) because it appears to enforce strict control, but the trap is that it ignores the need for team-specific innovation and risks shadow AI adoption.

How to eliminate wrong answers

Option A is wrong because migrating all GenAI workloads to a single on-premises server ignores the scalability and elasticity requirements of 50,000 employees, leading to high capital expenditure, limited GPU availability, and potential performance bottlenecks—cloud-based GenAI models require dynamic resource allocation. Option C is wrong because mandating a single model (e.g., Gemini) via a centralized endpoint with usage quotas eliminates team-specific customizations and may not suit diverse use cases (e.g., marketing vs. engineering), reducing innovation and causing shadow IT workarounds. Option D is wrong because allowing teams to continue using their own models with only monthly cost reports provides no proactive governance—costs can spiral out of control before reports are reviewed, and output quality remains inconsistent without enforced standards.

43
Multi-Selecthard

Which THREE approaches are effective for reducing bias in generative model outputs? (Choose three.)

Select 3 answers
A.Set temperature to a very high value.
B.Use adversarial training.
C.Use a balanced training dataset.
D.Use prompt engineering to specify neutral tone.
E.Fine-tune on a debiased dataset.
AnswersC, D, E

Balanced data reduces representation bias.

Why this answer

A balanced training dataset reduces the risk of the model learning spurious correlations or skewed distributions that lead to biased outputs. By ensuring that all demographic groups, topics, or perspectives are represented proportionally, the model's learned probability distribution is less likely to favor one group over another, directly mitigating representation bias at the data level.

Exam trap

The trap here is that candidates confuse randomness (high temperature) with fairness, or mistake adversarial training (a robustness technique) for a bias mitigation method, when in fact bias reduction requires data-level or fine-tuning interventions like balanced datasets, debiased fine-tuning, or prompt engineering.

44
MCQmedium

A company is building a GenAI chatbot that needs to answer questions using real-time data from their CRM and inventory systems. They want to ensure the model can access external data on demand. Which approach should they use?

A.Fine-tune the model on historical CRM and inventory data
B.Prompt the model to guess the data based on general knowledge
C.Use Vertex AI Extensions to connect to CRM and inventory APIs
D.Export CRM data to BigQuery and use that static snapshot
AnswerC

Vertex AI Extensions let the model invoke external APIs at inference time, so the chatbot retrieves live CRM and inventory records on demand rather than relying on stale training data. This satisfies the requirement for real-time access to external systems that grounding alone cannot provide.

Why this answer

Vertex AI Extensions allow the GenAI chatbot to connect to external APIs (like CRM and inventory systems) in real time, enabling on-demand data retrieval without retraining the model. This approach uses a retrieval-augmented generation (RAG) pattern where the model queries live data sources via API calls, ensuring responses are based on current information rather than static snapshots.

Exam trap

This question tests the distinction between fine-tuning (which changes model weights for static knowledge) and real-time data access via extensions or RAG, where candidates mistakenly think fine-tuning can provide live data when it only captures historical patterns.

How to eliminate wrong answers

Option A is wrong because fine-tuning on historical CRM and inventory data would only embed past patterns into the model, not provide real-time access to current data; the model would still be unable to answer questions about live inventory levels or recent customer interactions. Option B is wrong because prompting the model to guess data based on general knowledge would produce hallucinated or outdated responses, as the model has no inherent access to proprietary, real-time business data. Option D is wrong because exporting CRM data to BigQuery as a static snapshot would create a fixed dataset that becomes stale over time, failing the requirement for on-demand, real-time data access.

45
MCQmedium

A financial institution wants to use Gemini to analyze customer support transcripts and generate summaries. They need to ensure that personally identifiable information (PII) is not included in the summaries. Which approach should they take?

A.Preprocess the transcripts with Cloud DLP API to redact PII before sending to Gemini
B.Use a carefully engineered prompt instructing Gemini not to include PII
C.Post‑process the generated summaries with a regex filter to remove PII
D.Fine‑tune Gemini to avoid generating PII
AnswerA

Cloud DLP API performs deterministic de-identification — infoType detectors plus redaction or replacement — on the transcript text before any prompt reaches Gemini, so PII never enters the model's context and cannot surface in generated summaries. This satisfies the stem's requirement that summaries exclude PII, rather than relying on post-hoc filtering.

Why this answer

The Cloud Data Loss Prevention (DLP) API provides a purpose-built, scalable service to detect and redact PII from text before it reaches the Gemini model. This ensures that sensitive data is removed at the source, preventing any possibility of leakage in the generated summary, regardless of model behavior or prompt engineering.

Exam trap

Google often tests the misconception that prompt engineering or post-processing can reliably handle security requirements, when in fact a dedicated data loss prevention service like Cloud DLP is the only robust approach for guaranteed PII redaction before model inference.

How to eliminate wrong answers

Option B is wrong because prompt engineering alone cannot guarantee PII removal; Gemini may still inadvertently include PII due to model hallucinations, context misinterpretation, or adversarial inputs. Option C is wrong because post-processing with a regex filter is brittle and cannot reliably catch all PII formats (e.g., context-dependent identifiers, non-standard patterns), and PII may have already been exposed in the model's output. Option D is wrong because fine-tuning Gemini to avoid generating PII is impractical and risky; it requires extensive labeled data, may degrade model performance, and cannot guarantee complete PII avoidance across all edge cases.

46
MCQmedium

A retail bank uses a Gemini model on Vertex AI to answer customer questions about its mortgage products. Testers report that the model sometimes invents interest rates that do not exist in the bank's rate sheet. The team wants the model to ground its answers only in an approved corpus of product documents and cite the passages it used. Which approach should they implement?

A.Lower the temperature parameter to 0 and increase the topK value so the model samples only from the most probable tokens.
B.Configure Vertex AI Search grounding with the bank's document store and enable citations in the generateContent request.
C.Fine-tune the Gemini model on the bank's historical customer support transcripts so it learns the correct rates.
D.Increase the model's output token limit so it has more room to state the correct rate before answering.
AnswerB

Grounding with Vertex AI Search retrieves passages from the indexed rate sheets and product documents, injects them into the prompt, and returns grounding metadata so the response can cite its supporting sources. This directly constrains the model to approved content and satisfies the citation requirement, which sampling parameters alone cannot do.

Why this answer

Grounding with Vertex AI Search connects the model to the bank's authoritative document store, so answers are conditioned on retrieved passages rather than on parametric memory, and the grounding metadata enables citations. Sampling controls only shape randomness, fine-tuning bakes in a snapshot of knowledge without provenance, and token limits only change length.

Exam trap

The trap here is assuming that deterministic sampling settings such as temperature 0 make a model factually accurate, when they only reduce randomness and never supply missing source data.

47
MCQmedium

A media company wants its editors to summarize long internal research documents. Legal requires that no document content be used to train or improve any model, and that data stays within the company's Google Cloud project. Which capability should the company verify before adopting a Gemini-based solution on Vertex AI?

A.That prompts and responses are excluded from model training and covered by Google Cloud's data processing terms.
B.That the summarization endpoint supports a 1-million-token context window for the longest documents.
C.That the company's editors can access the Gemini web app with their personal Google accounts for convenience.
D.That the model can be fine-tuned on the research documents so summaries match the editors' preferred style.
AnswerA

Vertex AI's terms state that customer prompts and responses are not used to train or improve Google's foundation models, and customer data is processed under the Google Cloud Data Processing Addendum. Verifying this directly satisfies legal's two requirements: no training use and residency within the company's Google Cloud project boundary, making it the decisive check before adoption.

Why this answer

Legal's requirements are about how data is governed, not about model features. Confirming that prompts and responses are excluded from training and processed under enterprise data terms addresses both the no-training condition and the residency condition. Fine-tuning contradicts the no-training rule, context window size is irrelevant to governance, and personal accounts bypass the enterprise controls the company depends on.

Exam trap

The trap here is treating a technical capability such as a large context window or fine-tuning as if it answered a data-governance question, when the real issue is contractual training exclusion and project-bound processing.

48
MCQhard

An organization is using Vertex AI to fine-tune a large language model. They notice training is taking longer than expected and cost is increasing. Which action is most likely to reduce training time and cost without significantly impacting model quality?

A.Increase the number of training steps
B.Increase the batch size
C.Use a higher learning rate
D.Enable mixed-precision training (bfloat16)
AnswerD

bfloat16 mixed-precision training halves memory bandwidth and uses tensor cores for faster matrix operations, cutting training time and compute cost. It preserves numerical range well enough that model quality remains largely unaffected, unlike aggressive precision reduction.

Why this answer

Mixed-precision training with bfloat16 reduces memory usage and accelerates computation by using half the bits of standard float32, which directly decreases training time and cost on TPUs and modern GPUs. Vertex AI supports bfloat16 natively on TPU v3+ and A100 GPUs, and for many large language models, this precision preserves model quality because bfloat16 retains the same exponent range as float32, avoiding underflow issues common with float16.

Exam trap

A common pitfall is assuming increasing batch size or learning rate will speed up training, but this can destabilize training or degrade quality. Mixed-precision training (bfloat16) directly reduces computation without the risk of underflow, making it the optimal choice for Vertex AI.

How to eliminate wrong answers

Option A is wrong because increasing the number of training steps would lengthen training time and increase cost, directly opposing the goal. Option B is wrong because increasing batch size can reduce the number of weight updates per epoch, but it often requires tuning the learning rate and may degrade convergence or model quality if pushed too high, and it does not directly address the computational bottleneck of precision. Option C is wrong because using a higher learning rate risks training instability, divergence, or poor convergence, which can significantly harm model quality and may require additional tuning steps, negating any time savings.

49
MCQhard

A company is using Gemini Pro for code generation. They want to ensure that the generated code does not contain security vulnerabilities. Which approach should they implement?

A.Enable grounding with security scanning tools
B.Use the Vertex AI Codey API with safety settings
C.Implement a human-in-the-loop review with automated scanning
D.Use a custom safety attribute filter
AnswerC

Automated scanning catches known vulnerability patterns such as injection flaws, while human review judges context and logic that scanners miss. Combining both satisfies the requirement that generated code contain no security vulnerabilities, since neither control alone covers the full risk surface.

Why this answer

Combining human-in-the-loop review with automated scanning directly addresses the need to catch security vulnerabilities in AI-generated code. Human reviewers can identify logic flaws and context-specific risks that automated tools miss, while automated scanners provide consistent, rapid detection of known vulnerability patterns (e.g., OWASP Top 10). This layered approach is a best practice for production-grade code generation with Gemini Pro, as it mitigates the inherent limitations of relying solely on AI safety filters or static analysis.

Exam trap

The trap here is that candidates confuse content safety filters (which block toxic or harmful text) with code security vulnerability scanning, leading them to incorrectly select options that rely on Vertex AI safety settings or grounding, which are not designed to detect code-level security flaws.

How to eliminate wrong answers

Option A is wrong because grounding with security scanning tools typically refers to grounding model outputs against external data sources (e.g., enterprise databases) to improve factual accuracy, not to scan generated code for vulnerabilities; it does not replace dedicated vulnerability scanning. Option B is wrong because the Vertex AI Codey API's safety settings are designed to filter harmful or toxic content (e.g., hate speech, violence) in the generated text, not to detect security vulnerabilities like SQL injection or buffer overflows in code. Option D is wrong because a custom safety attribute filter in Vertex AI is used to block outputs based on categories like toxicity or harassment, not to perform code-specific security analysis; it lacks the semantic understanding needed to identify vulnerabilities in code logic.

50
MCQmedium

A software company wants to build a generative AI application that can answer questions based on its internal documentation. The documentation is stored in Google Drive and Confluence. They need a managed solution that can index these sources, provide relevant answers with citations, and integrate with their existing identity provider for access control. Which Google Cloud service should they use?

A.Vertex AI Search
B.Vertex AI Feature Store
C.Cloud SQL
D.Vertex AI Pipelines
AnswerA

Vertex AI Search is a managed service that can index data from various sources including Google Drive and Confluence, and it provides generative AI-powered answers with citations. It supports integration with identity providers for access control, ensuring users only see documents they are authorized to access. This matches all the requirements.

Why this answer

Vertex AI Search is a fully managed service that can connect to Google Drive and Confluence, index the content, and provide generative AI answers with citations. It also integrates with identity providers to enforce document-level access control, making it the ideal solution for the software company's needs.

Exam trap

The trap here is assuming that any Vertex AI service can handle search; only Vertex AI Search is designed for enterprise document indexing and question answering.

51
Multi-Selecteasy

A marketing team wants to generate social media posts using generative AI. They need the tone to be consistent with their brand voice. Which two prompt engineering techniques should they use? (Choose TWO)

Select 2 answers
A.Set maximum output tokens to a low value
B.Use few-shot examples of approved posts
C.Use negative prompts like 'do not be casual'
D.Set high temperature to encourage creativity
E.Include a detailed brand style guide in the system prompt
AnswersB, E

Few-shot examples of approved posts demonstrate the exact tone, vocabulary and structure the brand expects. Providing these in-context lets the model imitate the desired voice, satisfying the consistency requirement through pattern matching rather than abstract description alone.

Why this answer

Providing a style guide in the system prompt and using few-shot examples are effective techniques to enforce brand voice. Random examples and negative phrasing are not recommended. Maximum tokens does not affect tone.

52
MCQmedium

A team is tuning a large language model for a question-answering task. They notice the model gives high confidence scores to answers that are factually incorrect. Which evaluation metric should they primarily use to detect this overconfidence problem?

A.Perplexity
B.Expected Calibration Error (ECE)
C.BLEU score
D.ROUGE-L
AnswerB

Expected Calibration Error directly measures the gap between predicted confidence and actual accuracy across probability bins, so systematically high confidence on wrong answers produces a large ECE value. This satisfies the stem's overconfidence constraint, unlike accuracy or F1, which ignore confidence entirely and cannot expose miscalibration.

Why this answer

Expected Calibration Error (ECE) directly measures the alignment between a model's predicted confidence and its actual accuracy. In this scenario, high confidence on incorrect answers indicates miscalibration, and ECE quantifies this mismatch by binning predictions by confidence and computing the average absolute difference between accuracy and confidence per bin.

Exam trap

Google Cloud often tests the distinction between intrinsic evaluation metrics (like perplexity) and calibration metrics, leading candidates to mistakenly choose perplexity when the core issue is confidence miscalibration rather than general model uncertainty.

How to eliminate wrong answers

Option A is wrong because Perplexity measures how well a probability distribution predicts a sample, reflecting model uncertainty over token sequences, but it does not assess calibration of confidence scores against factual correctness. Option C is wrong because BLEU score evaluates n-gram overlap between generated and reference texts for translation quality, not confidence calibration or factual accuracy. Option D is wrong because ROUGE-L measures longest common subsequence recall for summarization tasks, and is unrelated to detecting overconfidence in model predictions.

53
MCQhard

A financial services firm needs to deploy a generative AI model that generates reports from structured and unstructured data. The solution must ensure that outputs never contain sensitive customer information. Which combination of Google Cloud services should they use?

A.Vertex AI Gemini API with DLP integration + IAM roles + VPC Service Controls
B.Vertex AI Gemini API + Cloud Key Management Service (KMS) + Secret Manager
C.Vertex AI Gemini API + Cloud Data Loss Prevention (DLP) standalone + Cloud NAT
D.Vertex AI Gemini API + Cloud DLP + Cloud Armor
AnswerA

Gemini on Vertex AI generates the reports, while DLP integration inspects and redacts sensitive customer data in prompts and outputs. IAM roles and VPC Service Controls restrict access and exfiltration, collectively satisfying the requirement that outputs never contain sensitive information.

Why this answer

The requirement is twofold: generate reports from mixed data and guarantee that outputs never contain sensitive customer information. Vertex AI Gemini API provides the generative capability, Cloud DLP integration (via the DLP API or Vertex AI's built-in DLP integration) inspects and de-identifies sensitive data in prompts and responses, IAM roles enforce least-privilege access to the model and data, and VPC Service Controls create a security perimeter that prevents data exfiltration. Together these address generation, data protection, access control, and network isolation — the complete combination.

Exam trap

Generative AI Leader often tests the misconception that encryption (KMS) or network controls (Cloud NAT/Armor) protect against sensitive data leakage in AI outputs — only DLP-style inspection and de-identification actually address that requirement.

How to eliminate wrong answers

Option B is wrong because KMS and Secret Manager handle encryption key management and secret storage respectively; neither inspects or redacts sensitive data in model inputs/outputs, so outputs could still contain PII. Option C is wrong because Cloud NAT provides outbound internet address translation for private instances — it has nothing to do with data loss prevention or access control, and standalone DLP without IAM/VPC-SC leaves exfiltration paths open. Option D is wrong because Cloud Armor is a WAF/DDoS protection service for external HTTP(S) load balancers; it does not inspect generative AI outputs for sensitive data, so it fails the core requirement.

54
MCQmedium

A company is fine-tuning a generative AI model on proprietary customer data. They are concerned about copyright and IP issues when using the model commercially. What is the BEST practice to mitigate these risks?

A.Apply SynthID to all generated content to prove origin
B.Only use data that is explicitly licensed for commercial use and document its provenance
C.Include a disclaimer on all outputs that the company is not liable for IP infringement
D.Use a model trained on publicly available data only
AnswerB

Licensed commercial-use data with documented provenance directly satisfies the copyright and IP constraint: each training example carries traceable rights, so commercial output cannot infringe unlicensed material. Provenance records also evidence due diligence if a claim arises.

Why this answer

The most reliable way to mitigate copyright and IP risk when fine-tuning on proprietary data is to ensure every piece of training data is explicitly licensed for commercial use and to maintain clear documentation of its provenance. This creates an auditable chain of custody that supports both legal defensibility and compliance. It addresses the root cause—data rights—rather than trying to patch symptoms after generation.

Exam trap

The trap is assuming that technical measures like watermarking or disclaimers can substitute for proper data licensing; the exam expects candidates to recognize that IP risk is fundamentally a data-rights and documentation problem.

How to eliminate wrong answers

Option A is wrong because SynthID is a watermarking technique for identifying AI-generated content, not a mechanism for resolving copyright or IP ownership of training data. Option C is wrong because a disclaimer does not eliminate legal liability for IP infringement and may not be enforceable. Option D is wrong because publicly available data is not necessarily free of copyright restrictions; much public web content is still protected and may have license terms that prohibit commercial training use.

55
MCQmedium

A retail company is building a product description generator using a large language model on Vertex AI. They need to ensure the generated descriptions do not contain offensive language. Which strategy should they implement?

A.Fine-tune the model on a dataset of clean product descriptions
B.Implement a content moderation filter (e.g., Perspective API) as a post-processing step
C.Use Vertex AI Model Monitoring to detect anomalies in model predictions
D.Include explicit instructions in the prompt to avoid offensive language
AnswerB

A post-processing moderation filter satisfies the requirement that outputs contain no offensive language, because the filter inspects each generated description and blocks or rewrites flagged text. This catches toxicity the base model may emit regardless of prompt design, giving a deterministic safety layer before content reaches customers.

Why this answer

Content moderation filters like Perspective API act as a post-processing safeguard that can catch offensive language the model might generate despite prompt engineering or fine-tuning. This approach provides a deterministic, rule-based or ML-based check that is independent of the model's training, ensuring compliance with content policies in production. It is a standard practice for deploying LLMs in customer-facing applications where safety is critical.

Exam trap

Google Cloud often tests the misconception that prompt engineering or fine-tuning alone can guarantee safety, when in practice a dedicated post-processing filter is required for reliable content moderation in production.

How to eliminate wrong answers

Option A is wrong because fine-tuning on clean product descriptions reduces but does not eliminate the risk of generating offensive language; the model can still hallucinate or produce harmful outputs due to biases in the base model or adversarial inputs. Option C is wrong because Vertex AI Model Monitoring detects anomalies in prediction distributions (e.g., drift, data skew) but does not inspect individual outputs for offensive content; it is a monitoring tool, not a content filter. Option D is wrong because including explicit instructions in the prompt is a weak safeguard; LLMs can ignore or misinterpret instructions, especially under prompt injection or when generating long descriptions, making it unreliable as a sole defense.

56
MCQeasy

A developer is using Vertex AI Studio to prototype a chat application. They want to provide the model with a system instruction to set the tone and style. How should they configure this in the Vertex AI Studio interface?

A.Add the instruction as part of the prompt text
B.Set the temperature parameter to a high value
C.Use the 'System Instruction' field in the model configuration
D.Add the instruction in the 'Context' parameter
AnswerC

The System Instruction field accepts a dedicated prompt that persists across turns, defining the assistant's persona, tone and style independently of user messages. Entering the desired tone there satisfies the requirement to set behaviour at the model configuration level rather than restating it in every user prompt.

Why this answer

Vertex AI Studio provides a dedicated 'System Instruction' field in the model configuration panel, which allows developers to set the tone, style, and behavioral guidelines for the model without mixing them into the user prompt. This field is specifically designed to hold system-level instructions that are prepended to the conversation context, ensuring consistent behavior across multiple turns.

Exam trap

The trap here is that candidates often confuse the 'System Instruction' field with the 'Context' parameter, mistakenly thinking both serve the same purpose, but the 'Context' parameter is designed for providing background knowledge or few-shot examples, not for setting persistent behavioral instructions.

How to eliminate wrong answers

Option A is wrong because adding the instruction as part of the prompt text would mix system-level guidance with user input, making it harder to maintain consistency and potentially causing the model to treat the instruction as part of the conversation rather than a persistent directive. Option B is wrong because the temperature parameter controls randomness in output generation, not the tone or style; a high temperature increases creativity and variability but does not enforce a specific behavioral instruction. Option D is wrong because the 'Context' parameter in Vertex AI Studio is used to provide background information or examples for grounding the model, not for setting system-level behavioral instructions like tone or style.

57
MCQhard

Which of the following is a best practice when using Vertex AI for prompt engineering?

A.Always set temperature to 0
B.Use consistent formatting and delimiters
C.Avoid using examples in the prompt
D.Use very long prompts to include all possible instructions
AnswerB

Consistent formatting and delimiters give the model an unambiguous structure, separating instructions from input data. This reduces parsing ambiguity and variance in outputs, satisfying the prompt engineering best practice of reproducible, predictable responses across repeated runs.

Why this answer

Consistent formatting and delimiters (e.g., using triple backticks, XML tags, or clear section headers) help the model parse instructions and context reliably, reducing ambiguity and improving output quality. This is a core best practice in prompt engineering on Vertex AI because it leverages the model's attention mechanisms to focus on distinct prompt segments, leading to more predictable and accurate responses.

Exam trap

Google Cloud often tests the misconception that 'more is better' in prompts or that deterministic settings like temperature=0 are universally optimal, leading candidates to overlook the importance of structured, concise formatting.

How to eliminate wrong answers

Option A is wrong because setting temperature to 0 always is not a best practice; temperature controls randomness, and while 0 yields deterministic outputs, many tasks benefit from slight variability (e.g., creative generation or diverse suggestions), and Vertex AI supports a range of 0.0 to 1.0. Option C is wrong because including examples (few-shot prompting) is a powerful technique to guide the model's behavior and improve performance, especially for complex or nuanced tasks; avoiding them would reduce effectiveness. Option D is wrong because very long prompts can exceed context windows, dilute key instructions, and increase latency or cost; Vertex AI models have token limits (e.g., 8,192 tokens for Gemini), and concise, well-structured prompts are more efficient.

58
Multi-Selecthard

A team is fine-tuning a model for a legal document summarization task. They need to ensure high accuracy and avoid hallucinations. Which TWO approaches should they combine? (Choose two.)

Select 2 answers
A.Use Retrieval-Augmented Generation to retrieve relevant legal texts
B.Increase temperature to 1.5 during inference
C.Implement early stopping during fine-tuning
D.Incorporate a human-in-the-loop review process
E.Use character-level tokenization to improve spelling
AnswersA, D

RAG grounds the summary in actual documents, reducing hallucination.

Why this answer

Retrieval-Augmented Generation (RAG) is correct because it grounds the model's output in retrieved, authoritative legal texts, directly reducing hallucination by providing factual context during generation. This is critical for legal summarization where accuracy is paramount, as RAG ensures the model references specific statutes or case law rather than relying solely on its parametric memory.

Exam trap

A common misconception is that increasing temperature or using training-time techniques like early stopping can improve inference accuracy, when in fact they either increase randomness or address overfitting, not factual grounding. This trap is frequently tested in Google certification exams.

59
MCQmedium

A marketing team is using a generative AI model to create ad copy. They notice that the outputs sometimes include made-up statistics and false claims about their products. They want to reduce these hallucinations without retraining the model. What should they do?

A.Use grounding with a source of truth such as a product database.
B.Decrease the maximum output tokens.
C.Increase the model's temperature parameter.
D.Fine-tune the model on a small set of correct ad copies.
AnswerA

Grounding allows the model to reference an external, authoritative data source (like a product database) when generating responses. This reduces hallucinations by anchoring outputs in verified facts. It does not require retraining and can be implemented via Vertex AI's grounding features, such as using a corpus or connecting to a database.

Why this answer

Grounding connects the model to verified external data, ensuring responses are based on facts rather than model assumptions. It is a no-retraining approach that directly targets hallucinations by providing context. Other methods like temperature or token limits do not reliably improve factual accuracy.

Exam trap

The trap here is assuming that fine-tuning or parameter tweaks can eliminate hallucinations without external data grounding.

60
MCQhard

An e-commerce company fine-tunes a model on customer reviews to generate product feedback summaries. They want to ensure the model does not reproduce toxic language from the training data. Besides filtering the training data, which additional technique is most effective at inference time?

A.Set temperature to 0.0 to reduce variance
B.Set top-k to 10 to limit token choices
C.Pass the model output through a toxicity detection model and conditionally regenerate or block
D.Use beam search with a high beam width
AnswerC

Filtering training data alone cannot guarantee toxic-free output, so a post-generation toxicity classifier screens each summary and triggers regeneration or blocking when thresholds are breached. This directly satisfies the inference-time constraint, catching residual toxicity the fine-tuning absorbed. It operates on the model's actual output rather than inputs, making it the most reliable safeguard.

Why this answer

It directly addresses the safety requirement at inference time by introducing a secondary guardrail. A toxicity detection model (e.g., a classifier trained on the Jigsaw Toxic Comment dataset) can score the generated output in real time; if the score exceeds a threshold, the system can either block the response or trigger a regeneration with adjusted parameters. This is the only technique that actively filters for toxic language after generation, rather than merely reducing output variance or exploring alternative sequences.

Exam trap

Google often tests the misconception that controlling randomness (temperature, top-k) or search strategy (beam search) can prevent toxic outputs, when in fact these techniques only affect token probability distributions and do not perform any semantic safety filtering.

How to eliminate wrong answers

Option A is wrong because setting temperature to 0.0 makes the model deterministic (always picks the highest-probability token), which reduces randomness but does not prevent the model from reproducing toxic phrases that were present in the training data. Option B is wrong because top-k sampling limits the token pool to the k most likely tokens, which can still include toxic tokens if they rank highly; it does not perform any semantic or safety filtering. Option D is wrong because beam search with a high beam width explores multiple candidate sequences to find a high-probability output, but it does not incorporate any toxicity detection or safety constraint, so it may still select a toxic sequence if it has high likelihood.

61
MCQmedium

A startup with limited ML expertise wants to add a GenAI feature to their SaaS application that can generate personalized email drafts for users. They need fast time-to-market and low maintenance. Which build-vs-buy decision is BEST?

A.Select a model from Model Garden and deploy it on Vertex AI
B.Fine-tune an open-source model on a corpus of email drafts to create a custom model
C.Buy a pre-built API such as the Gemini API and integrate it with prompt engineering for personalization
D.Build a custom transformer model from scratch
AnswerC

A pre-built API delivers generative capability immediately, with no model training, hosting or tuning burden, meeting the fast time-to-market and low-maintenance constraints. Prompt engineering supplies the personalisation, so the startup avoids building and operating its own model infrastructure.

Why this answer

Buying a pre-built API like the Gemini API and using prompt engineering for personalization gives the startup fast time-to-market with minimal ML expertise and low maintenance, since Google manages the model infrastructure. Prompt engineering can tailor email drafts without training or fine-tuning, aligning with the need for speed and low operational overhead.

Exam trap

Generative AI Leader often tests the trade-off between customization and speed — candidates overvalue fine-tuning or custom builds for personalization, ignoring that prompt engineering on a managed API meets most personalization needs faster and cheaper.

How to eliminate wrong answers

Option A is wrong because deploying a model from Model Garden on Vertex AI still requires ML expertise to select, deploy, and manage the model, and it adds infrastructure maintenance overhead compared to a managed API. Option B is wrong because fine-tuning an open-source model requires a corpus, ML expertise, training infrastructure, and ongoing model maintenance — contradicting the need for fast time-to-market and low maintenance. Option D is wrong because building a custom transformer from scratch is the most expensive, time-consuming, and expertise-intensive option, completely misaligned with the startup's constraints.

62
Multi-Selectmedium

A company is using Vertex AI RAG Engine to ground a chatbot in internal documents. The chatbot sometimes returns outdated information. Which TWO steps should they take to improve freshness?

Select 2 answers
A.Set up automated re-indexing on a schedule (e.g., daily)
B.Reduce the chunk size to 50 tokens
C.Increase the model's temperature to 1.0
D.Implement document chunking with metadata such as version or timestamp
E.Disable grounding and rely on the model's pre-training data
AnswersA, D

Scheduled automated re-indexing periodically ingests updated source documents into the RAG corpus, so newly revised content replaces outdated embeddings. This directly satisfies the freshness requirement by preventing the index from serving stale chunks after source documents change.

Why this answer

Option A is correct because RAG Engine only retrieves from the indexed corpus, so scheduling automated re-indexing (e.g., daily via a pipeline or scheduled job) ensures newly updated or replaced documents are reflected in the vector index and stale chunks are removed. Option D is correct because attaching metadata such as version or timestamp to chunks lets retrieval filter or rank by recency, so the chatbot prefers the latest document version instead of matching an outdated chunk with similar embeddings. Option B is not appropriate because shrinking chunks to 50 tokens fragments context and does not address index staleness.

Option C is wrong because raising temperature to 1.0 increases randomness and does not improve factual freshness. Option E is wrong because disabling grounding removes the internal-document retrieval entirely and falls back to potentially outdated pre-training knowledge.

Exam trap

The trap is confusing 'freshness' with other RAG tuning knobs (chunk size, temperature) — candidates who don't distinguish retrieval-index currency from generation parameters pick B or C.

63
MCQeasy

A user provides a long document as context for a question-answering task, but the model outputs irrelevant answers. What is the most likely cause?

A.The document exceeds the model's context window, truncating important details.
B.Safety filters are blocking the relevant response.
C.The model's temperature is too low, making it deterministic.
D.The model is not generating any tokens.
AnswerA

Transformer models process a fixed token budget; text beyond the context window is silently truncated. With a long document, the crucial passages answering the question may fall outside that window, so the model responds from incomplete context, producing irrelevant output.

Why this answer

The most common cause of irrelevant answers when a long document is provided is that the document exceeds the model's fixed context window (e.g., 8K tokens for PaLM 2, 128K for Gemini 1.0 Pro, or up to 1M for Gemini 1.5 Pro). When the input is truncated, critical details needed for accurate retrieval and generation are lost, leading to off-target responses. This is a fundamental limitation of transformer architectures, which cannot attend to tokens beyond their maximum sequence length.

Exam trap

Google often tests the misconception that safety filters or temperature settings are the primary cause of irrelevant outputs, when in fact the context window limit is the most direct and common technical constraint in long-document QA tasks.

How to eliminate wrong answers

Option B is wrong because safety filters block harmful or policy-violating content, not relevant factual answers; they would either suppress the response entirely or return a refusal, not produce irrelevant answers. Option C is wrong because a low temperature (e.g., 0.0) makes the model more deterministic and repetitive, but it does not cause irrelevance—it would still generate answers based on the available context, albeit with less creativity. Option D is wrong because if the model were not generating any tokens, the output would be empty or a failure, not irrelevant answers; the question explicitly states the model outputs irrelevant answers, meaning tokens are being generated.

64
MCQmedium

A company is using Google Cloud's generative AI offerings to build a customer-facing application. They need to ensure that the AI-generated content complies with their brand guidelines and does not produce harmful or inappropriate responses. They also want to monitor and filter content in real-time. Which Google Cloud feature should they use?

A.Vertex AI safety filters and attributes
B.Cloud Armor
C.Identity and Access Management (IAM)
D.Cloud Data Loss Prevention (DLP)
AnswerA

Vertex AI provides configurable safety filters and attributes that can block or flag harmful content in real-time. These can be customized to align with brand guidelines and compliance requirements. They are integrated into the model API, allowing real-time monitoring and filtering of generated content.

Why this answer

Vertex AI safety filters and attributes are the correct choice because they are specifically designed to detect and block harmful content in generative AI outputs. They can be configured to enforce brand-specific guidelines and are applied in real-time during model inference. The other options are security or data protection services that do not address content moderation for generative AI.

Exam trap

The trap here is confusing general security services like Cloud Armor or IAM with content moderation features tailored for generative AI outputs.

65
MCQhard

A financial analytics firm is building a generative AI application that must analyze long earnings call transcripts and produce summaries. The transcripts often exceed 200,000 tokens, and the firm wants to minimize cost while maintaining high accuracy. They plan to use Gemini models on Vertex AI. Which approach should they take?

A.Use gemini-1.0-pro and rely on its built-in automatic long-document summarization feature.
B.Split the transcript into 1,000-token chunks and summarize each chunk separately, then concatenate the summaries.
C.Fine-tune a smaller Gemini model on the firm's past transcripts and use it for summarization.
D.Use gemini-1.5-pro with its long context window and pass the entire transcript in a single request.
AnswerD

Gemini 1.5 Pro supports a context window of up to 2 million tokens, which can accommodate transcripts exceeding 200,000 tokens in a single request. This avoids the complexity and potential accuracy loss of chunking and summarization pipelines, and it often reduces overall cost by eliminating multiple inference calls and preprocessing steps.

Why this answer

Gemini 1.5 Pro's extended context window allows the entire transcript to be processed in one request, preserving cross-references and reducing the need for complex chunking pipelines. Chunking loses context, gemini-1.0-pro lacks the required context length, and fine-tuning does not increase context capacity, so the long-context model is the most accurate and cost-effective choice.

Exam trap

The trap here is believing that fine-tuning can overcome a model's context window limitation, when fine-tuning only adapts behavior and does not expand the maximum input length.

66
MCQmedium

A company is evaluating the ROI of implementing GenAI for code generation. Which metric BEST captures the productivity improvement of developers?

A.Percentage of code that passes unit tests on the first attempt
B.Time saved per development task (e.g., from 2 hours to 30 minutes)
C.Number of lines of code generated per day
D.Number of bugs found in production after code review
AnswerB

Time saved per development task directly quantifies productivity gain by comparing task duration before and after GenAI assistance, such as two hours reduced to thirty minutes. This measurable delta captures developer efficiency improvement, unlike adoption rates or satisfaction scores, making it the strongest ROI productivity metric.

Why this answer

Time saved per development task directly quantifies productivity improvement by measuring the reduction in effort required to complete a unit of work. If a task that previously took 2 hours now takes 30 minutes, that is a concrete, measurable 75% reduction in time — the most direct and interpretable ROI metric for GenAI code generation. It captures the actual efficiency gain in terms of developer hours, which translates directly to cost savings and capacity for more work.

Exam trap

Generative AI Leader often tests the confusion between output volume metrics (lines of code) and outcome-based productivity metrics (time saved) — candidates pick lines of code because it seems quantifiable, missing that it does not measure actual productivity or value.

How to eliminate wrong answers

Option A is wrong because first-attempt unit test pass rate measures code quality, not productivity — code could pass tests quickly but still require significant developer time for design, integration, and debugging. Option C is wrong because lines of code generated per day is a vanity metric that does not correlate with value delivered — more lines of code can mean more complexity, more bugs, and more maintenance burden, not more productivity. Option D is wrong because bugs found in production measures quality outcomes, not productivity improvement — it is a lagging indicator of code quality and does not capture time savings or efficiency gains from GenAI assistance.

67
MCQmedium

A company is building a customer support chatbot using Vertex AI Agent Builder. They want the agent to answer questions based on internal knowledge base documents stored in Cloud Storage. Which feature should they configure to ensure the agent can retrieve relevant information from these documents?

A.Deploy the agent to a Vertex AI endpoint
B.Fine-tune a Gemini model on the knowledge base
C.Enable grounding with a data store
D.Configure a safety filter to block irrelevant queries
AnswerC

Grounding with a data store indexes the Cloud Storage documents and retrieves relevant passages at query time, anchoring responses in that content. This satisfies the requirement to answer from internal knowledge base documents rather than the model's parametric memory.

Why this answer

Vertex AI Agent Builder uses grounding to connect the agent to external data sources, such as documents stored in Cloud Storage. By enabling grounding with a data store, the agent can retrieve and reference relevant information from the knowledge base documents in real time, ensuring accurate and context-aware responses without requiring model retraining.

Exam trap

Candidates often confuse fine-tuning with grounding. They mistakenly choose fine-tuning (Option B) assuming the model must be retrained on the knowledge base, but the correct approach for retrieval-based Q&A is grounding with a data store.

How to eliminate wrong answers

Option A is wrong because deploying the agent to a Vertex AI endpoint is about making the agent accessible for inference, not about connecting it to a knowledge base for retrieval. Option B is wrong because fine-tuning a Gemini model on the knowledge base would adapt the model's weights to the specific data, which is unnecessary and inefficient for retrieval-based tasks; Vertex AI Agent Builder uses retrieval-augmented generation (RAG) via grounding instead. Option D is wrong because configuring a safety filter blocks harmful or irrelevant queries but does not enable the agent to retrieve information from the knowledge base documents.

68
MCQeasy

A marketing team wants to use generative AI to create ad copy that matches their brand voice. They have several examples of previous high-performing ads. Which Vertex AI Studio feature would best help them achieve consistent tone and style without custom model training?

A.Use of a pre-built template in Vertex AI Studio
B.Supervised fine-tuning on the ad examples
C.Few-shot prompting with examples of previous ads
D.Model evaluation to compare outputs
AnswerC

Few-shot prompting supplies several previous high-performing ads as in-context examples, steering the model toward the brand's tone and style without fine-tuning. This achieves consistency using Vertex AI Studio's prompt design, avoiding custom model training entirely.

Why this answer

Few-shot prompting in Vertex AI Studio allows the model to infer the desired tone and style from a small set of example ads without requiring custom model training. This approach leverages the model's in-context learning capability, making it ideal for quickly adapting to a brand voice while avoiding the cost and complexity of fine-tuning.

Exam trap

The trap here is that candidates often confuse few-shot prompting with fine-tuning, assuming that any use of examples requires model retraining, when in fact few-shot prompting achieves style transfer through in-context learning without modifying model parameters.

How to eliminate wrong answers

Option A is wrong because pre-built templates provide generic structures and do not adapt to a specific brand voice or learn from provided examples. Option B is wrong because supervised fine-tuning requires custom model training, which the question explicitly states should be avoided. Option D is wrong because model evaluation is a post-generation step used to assess output quality, not a method for guiding the model to produce consistent tone and style.

69
MCQeasy

An employee wants to use GenAI to assist with writing formulas in Google Sheets. Which Gemini for Google Workspace feature should they use?

A.Formula assistance in Google Sheets
B.Help me write in Google Docs
C.Image generation in Google Slides
D.Smart Compose in Gmail
AnswerA

Formula assistance in Google Sheets generates and explains spreadsheet formulas directly within the sheet, matching the employee's need to write formulas. Gemini for Google Workspace surfaces this contextual help where the data lives, rather than in a separate chat interface.

Why this answer

Gemini for Sheets provides formula assistance, helping users generate, explain, or debug formulas.

70
MCQeasy

A social media company uses a generative AI model to moderate user posts. The model occasionally allows offensive content. Which safety technique should be implemented?

A.Use a different tokenizer to avoid offensive words.
B.Configure safety filters on the model endpoint in Vertex AI.
C.Add few-shot examples of safe posts in the prompt.
D.Reduce the temperature to 0.
AnswerB

Configuring safety filters on the Vertex AI model endpoint applies configurable thresholds that block harmful categories before responses reach users, catching offensive content the base model permits. This is a platform-level control applied at inference time, complementing prompt design or fine-tuning.

Why this answer

Configuring safety filters on the model endpoint in Vertex AI directly blocks offensive content at inference time by applying predefined or custom safety thresholds (e.g., toxicity, harassment categories). This is the most reliable technique for real-time moderation, as it prevents harmful outputs regardless of prompt engineering or tokenization changes.

Exam trap

A common misconception tested in the Google Gen AI Leader exam is that prompt engineering (few-shot examples) or parameter tuning (temperature) can substitute for dedicated safety mechanisms, but these do not provide hard guarantees against offensive content.

How to eliminate wrong answers

Option A is wrong because changing the tokenizer does not prevent the model from generating offensive content; tokenizers only split text into tokens and do not understand or filter semantics. Option C is wrong because few-shot examples in the prompt can guide the model but are not a safety mechanism—they can be overridden by the model's training data or adversarial inputs, and they do not enforce hard safety constraints. Option D is wrong because reducing temperature to 0 makes the model deterministic but does not eliminate offensive content; it may even amplify biased or toxic patterns from the training data by always choosing the most likely token.

71
MCQeasy

What is the primary purpose of the temperature parameter in a generative language model?

A.It defines the context window size for the prompt
B.It controls the tradeoff between creativity and determinism
C.It limits the vocabulary to the top K tokens
D.It sets the maximum number of tokens in the output
AnswerB

Temperature scales the sampling distribution before token selection. Lower values sharpen probabilities toward the highest-scoring tokens, producing deterministic output; higher values flatten the distribution, letting lower-probability tokens be chosen and increasing creative variation. This directly governs the creativity-versus-determinism tradeoff.

Why this answer

The temperature parameter directly controls the probability distribution over the next token. A lower temperature (e.g., 0.1) makes the model more deterministic by favoring high-probability tokens, while a higher temperature (e.g., 1.5) flattens the distribution, increasing randomness and creative output. This tradeoff is fundamental to balancing coherence with novelty in generative models.

Exam trap

Candidates often confuse temperature with other sampling parameters (top-k, top-p) — they mistakenly think temperature controls output length or vocabulary size, when it strictly governs the randomness of token selection.

How to eliminate wrong answers

Option A is wrong because the context window size is determined by the model's architecture (e.g., 2048 tokens for GPT-2, 8192 for GPT-4), not by the temperature parameter. Option C is wrong because limiting vocabulary to the top K tokens is the function of the top-k sampling parameter, not temperature. Option D is wrong because the maximum number of output tokens is set by the max_tokens parameter in the API call, not by temperature.

72
MCQeasy

A developer is using Vertex AI PaLM API to generate code snippets. The responses sometimes contain security vulnerabilities. What is the best practice to mitigate this?

A.Implement input validation and output filtering with safety attributes
B.Disable safety filters to allow more output
C.Increase the max output tokens
D.Set safety settings to block all categories
AnswerA

Input validation rejects malicious prompts, while output filtering with safety attributes screens generated code for insecure patterns before delivery. This layered control mitigates vulnerable snippets at both ends, satisfying the requirement to reduce security flaws in Vertex AI PaLM responses.

Why this answer

Input validation and output filtering with safety attributes directly address security vulnerabilities by sanitizing user inputs and filtering model outputs for harmful content. The Vertex AI PaLM API provides safety attribute scores (e.g., toxicity, harassment) that allow developers to programmatically block or flag responses that exceed defined thresholds, reducing the risk of generating insecure code snippets.

Exam trap

The trap here is that candidates may think increasing token limits or disabling filters improves output quality, when in fact the core issue is controlling content safety through validation and filtering, not adjusting generation parameters.

How to eliminate wrong answers

Option B is wrong because disabling safety filters removes all guardrails, allowing the model to generate potentially harmful or insecure code without any mitigation, which increases security risks. Option C is wrong because increasing max output tokens does not affect the content's security; it only allows longer responses, which could include more vulnerabilities. Option D is wrong because setting safety settings to block all categories is overly restrictive and may prevent legitimate code generation, but more importantly, it does not address the root cause of vulnerabilities—input validation and output filtering are needed to catch context-specific issues like insecure code patterns.

73
MCQmedium

A fintech startup is building a generative AI application that generates personalized investment advice based on user profiles and market data. They are using Vertex AI Agent Builder to create an agent that retrieves information from a BigQuery table containing user data and from a real-time market data API. The agent needs to ensure that responses comply with financial regulations, meaning the model must not give specific stock recommendations unless the user explicitly requests them after disclaimers. The team has implemented grounding with both sources. During testing, the agent sometimes spontaneously suggests buying a particular stock without being asked, which could lead to regulatory issues. The team wants to enforce strict control over the agent's behavior. What should the team do?

A.Increase the safety filter sensitivity to block any financial recommendations
B.Add more historical data to the BigQuery table to improve grounding accuracy
C.Implement a custom system instruction that explicitly prohibits unsolicited stock recommendations and requires a disclaimer before any advice
D.Fine-tune the model on a dataset of compliant conversations
AnswerC

A custom system instruction directly constrains the model's generative behaviour, prohibiting unsolicited stock recommendations and mandating a disclaimer before advice. This satisfies the stem's regulatory constraint by enforcing deterministic policy at inference time, unlike grounding, which only supplies factual context and cannot govern what the agent chooses to say.

Why this answer

System instructions in Vertex AI Agent Builder allow you to define strict behavioral rules that the agent must follow, such as prohibiting unsolicited stock recommendations and requiring a disclaimer before any advice. This directly addresses the regulatory compliance issue by enforcing a policy at the agent's instruction layer, which overrides any learned or grounded behavior. Unlike other options, this approach provides explicit, enforceable control without altering data sources or model training.

Exam trap

The trap here is that candidates often confuse grounding (data retrieval) with behavioral control (system instructions), assuming that better data or safety filters can enforce compliance, when in fact only explicit instructions in the agent's configuration can enforce such nuanced policies.

How to eliminate wrong answers

Option A is wrong because increasing safety filter sensitivity would block all financial recommendations, including compliant ones after disclaimers, which breaks the required user-requested flow and is too blunt for nuanced regulatory compliance. Option B is wrong because adding more historical data to the BigQuery table improves grounding accuracy but does not prevent the agent from spontaneously generating unsolicited stock recommendations; grounding only ensures factual retrieval, not behavioral constraints. Option D is wrong because fine-tuning the model on a dataset of compliant conversations may reduce but not eliminate unsolicited recommendations, as the model can still generalize or hallucinate; it also requires significant effort and does not provide a deterministic, enforceable rule like system instructions do.

74
MCQmedium

A large enterprise has deployed generative AI assistants in three separate departments (HR, Marketing, and Customer Support) using different tools and models. Over the past quarter, the company has observed escalating cloud costs, inconsistent user experiences, and reports of data leakage in Customer Support logs. The CTO wants to address these issues while maintaining innovation velocity. As the Generative AI Leader, what course of action should you recommend?

A.Standardize on a single model and tool across all departments, restricting usage to one platform.
B.Implement a centralized AI governance platform with cost monitoring, model registry, and security guardrails.
C.Discontinue the Customer Support assistant to eliminate data leakage risk and reduce costs.
D.Allow each department to continue independently but require monthly cost and compliance reports.
AnswerB

A centralised governance platform directly addresses the three stated problems: cost monitoring curbs escalating cloud spend, a model registry standardises experiences across departments, and security guardrails contain the Customer Support data leakage, while preserving innovation velocity through shared controls.

Why this answer

A centralized AI governance platform directly addresses the CTO's concerns by providing cost monitoring to control escalating cloud costs, a model registry to ensure consistent user experiences across departments, and security guardrails to prevent data leakage. This approach maintains innovation velocity by allowing departments to continue using different tools and models while enforcing enterprise-wide policies, rather than restricting them to a single platform or eliminating valuable services.

Exam trap

Common misconception: standardization or elimination is often seen as the only way to solve governance issues, when in fact a centralized governance platform provides the necessary control without sacrificing flexibility or innovation velocity.

How to eliminate wrong answers

Option A is wrong because standardizing on a single model and tool restricts innovation velocity and ignores the fact that different departments (HR, Marketing, Customer Support) have unique requirements that are best served by specialized models; it also does not inherently solve data leakage or cost issues without governance. Option C is wrong because discontinuing the Customer Support assistant eliminates a valuable business function and fails to address the root cause of data leakage, which requires security guardrails and proper configuration rather than outright removal. Option D is wrong because allowing independent operation with only monthly reports provides no real-time enforcement of security or cost controls, leaving the enterprise vulnerable to continued data leakage and uncontrolled cloud spend.

75
MCQhard

A multinational corporation is using Vertex AI to generate multilingual customer support responses. They have fine-tuned the Gemini model on support tickets in English and now want to extend to 10 additional languages. The fine-tuning dataset for new languages is small (1000 tickets each). During evaluation, the model performs well for common languages (Spanish, French) but poorly for languages like Finnish and Thai. The team needs to improve performance for low-resource languages. They have budget constraints and cannot collect more data quickly. Which approach should they take?

A.Switch to Vertex AI Codey API for generating responses in all languages.
B.Use a multilingual foundation model and fine-tune with cross-lingual transfer learning techniques.
C.Deploy separate fine-tuned models for each language.
D.Collect more training data for low-resource languages via crowdsourcing.
AnswerB

Cross-lingual transfer learning leverages the multilingual foundation model's shared representations, so the 1,000-ticket datasets per language transfer knowledge from the English fine-tune. This lifts Finnish and Thai performance without collecting more data, satisfying the budget constraint.

Why this answer

Using a multilingual foundation model (like Gemini's multilingual variant) with cross-lingual transfer learning leverages the model's pre-trained knowledge across languages, allowing it to generalize from high-resource languages (Spanish, French) to low-resource ones (Finnish, Thai) even with small fine-tuning datasets. This approach is budget-friendly as it avoids separate models or costly data collection, and it directly addresses the performance gap by sharing linguistic patterns across languages.

Exam trap

The trap here is that candidates often assume more data (Option D) or separate models (Option C) are the only solutions, ignoring that cross-lingual transfer learning can effectively bootstrap low-resource languages from high-resource ones without additional data collection.

How to eliminate wrong answers

Option A is wrong because the Vertex AI Codey API is designed for code generation, not multilingual customer support responses, and switching to it would not improve performance for low-resource languages. Option C is wrong because deploying separate fine-tuned models for each language multiplies cost and maintenance overhead, and with only 1000 tickets per language, each model would suffer from the same data scarcity issue without cross-lingual benefits. Option D is wrong because the team has budget constraints and cannot collect more data quickly, making crowdsourcing infeasible in the short term, and it does not address the underlying need for transfer learning.

Page 1 of 14

Page 2