Courseiva

Google Cloud Generative AI Leader Generative AI Leader (Generative AI Leader) — Questions 526–600

1008 questions total · 14pages · All types, answers revealed

Page 7

Page 8 of 14

Page 9
526
MCQhard

A logistics company wants to forecast delivery delays and also generate plain-language explanations for dispatchers. Leadership asks the GenAI Leader to choose an approach that keeps the numeric forecast auditable while adding generative explanations. Which design should the GenAI Leader propose?

A.Use a single foundation model prompt that returns both the predicted delay in minutes and a narrative explanation in one response.
B.Fine-tune Gemini on historical delay records so it learns to output delay minutes directly, and have dispatchers read its responses.
C.Deploy a retrieval-augmented generation pipeline over past dispatch notes and ask the model to infer the expected delay from similar historical cases.
D.Train a Vertex AI forecasting model on historical shipment data, then pass its output and key features to Gemini to generate the dispatcher explanation.
AnswerD

A dedicated forecasting model produces a repeatable numeric prediction that can be backtested and audited, satisfying the accuracy requirement. Feeding that prediction plus the contributing features into Gemini lets the generative layer translate model output into readable guidance for dispatchers. Each component can be evaluated and improved independently, which is the cleanest separation of concerns for this scenario.

Why this answer

Keeping prediction and narration in separate components lets the forecasting model be validated with standard regression metrics while the language model handles communication. The numeric output remains deterministic and traceable, and the generative layer receives grounded inputs rather than inventing figures, which satisfies both leadership requirements.

Exam trap

The trap here is treating a generative model as a substitute for a purpose-built predictive model whenever the output can be phrased in natural language.

527
Multi-Selecthard

A team is designing prompts for a generative AI application that summarizes long legal contracts. They want outputs that are accurate, consistently formatted, and safe from leaking instructions. Which two prompt engineering practices should they apply? (Choose two.)

Select 2 answers
A.Set temperature to the maximum value to encourage more varied legal wording.
B.Instruct the model to ignore any instructions embedded inside the contract text.
C.Insert the entire contract text without delimiters and ask the model to figure out the important parts.
D.Provide a clear task instruction with an example of the desired output format.
E.Ask the model to reveal its system instructions so the team can verify them.
AnswersB, D

Contracts may contain clauses that look like instructions, and a malicious or accidental phrase could steer the model. Explicitly telling the model to treat document content as data and not as commands reduces prompt-injection risk and helps keep the summary faithful to the intended task, which is a recommended safety practice.

Why this answer

Giving a clear task plus a formatting example improves output consistency, and explicitly telling the model to ignore instructions inside the contract text reduces prompt-injection risk. These two practices together address accuracy, formatting stability, and safety when summarizing long legal documents, which were the team's stated goals.

Exam trap

The trap here is treating any instruction about instructions as helpful, when asking the model to reveal or obey embedded text can weaken safety instead of strengthening it.

528
MCQmedium

A healthcare startup needs to process medical claims and extract structured data (e.g., patient name, procedure codes, amounts) from scanned PDF forms. They require HIPAA compliance and prefer a pre-built Google Cloud service. Which service should they use?

A.Cloud Vision API with OCR
B.Vertex AI AutoML Tables
C.Document AI Healthcare API with the Claims processor
D.Natural Language AI with entity extraction
AnswerC

The Document AI Healthcare API's Claims processor is purpose-built to extract structured fields such as procedure codes and amounts from scanned claim forms, and it operates under HIPAA-compliant terms. That satisfies both the pre-built service and compliance constraints.

Why this answer

The Document AI Healthcare API with the Claims processor is a pre-built, HIPAA-compliant Google Cloud service specifically designed to extract structured data (e.g., patient name, procedure codes, amounts) from scanned medical claim forms. It leverages specialized machine learning models trained on healthcare documents, ensuring accurate parsing of fields like CPT codes and ICD-10 codes while meeting regulatory requirements.

Exam trap

The trap here is that candidates often confuse general OCR services (like Cloud Vision API) with domain-specific document processors, overlooking that HIPAA compliance and pre-built healthcare form parsing require a specialized service like Document AI Healthcare API rather than a generic text extraction tool.

How to eliminate wrong answers

Option A is wrong because Cloud Vision API with OCR only extracts raw text from images without understanding the document structure or healthcare-specific fields, and it lacks built-in HIPAA compliance for protected health information (PHI). Option B is wrong because Vertex AI AutoML Tables is for tabular data prediction (e.g., regression or classification on structured datasets), not for extracting structured data from scanned PDF forms. Option D is wrong because Natural Language AI with entity extraction is designed for analyzing unstructured text (e.g., clinical notes) and cannot process scanned PDF forms or extract structured fields like procedure codes from document layouts.

529
MCQhard

A media company is using a generative AI model to create video captions. The model is deployed on Vertex AI with autoscaling. During peak hours, they observe high latency and request timeouts. Which action would most effectively address this issue?

A.Optimize the prompt to reduce output length
B.Reduce the maximum number of replicas to limit resource usage
C.Switch to a GPU-based machine type for faster inference
D.Increase the minimum number of replicas in the autoscaling configuration
AnswerD

Autoscaling adds replicas only after load rises, so cold-start latency and timeouts persist during sudden peaks. Raising the minimum replica count keeps capacity warm, directly satisfying the peak-hour latency and timeout constraint by ensuring baseline throughput is always available.

Why this answer

Increasing the minimum number of replicas ensures that during peak hours, the model already has a baseline of warm instances ready to handle requests, reducing cold-start latency and preventing timeouts. Autoscaling can take time to spin up new replicas, so a higher minimum replica count directly mitigates the latency spike by pre-provisioning capacity.

Exam trap

The trap here is that candidates confuse performance optimization (faster inference per request) with capacity planning (ensuring enough concurrent replicas), leading them to choose GPU upgrades or prompt tweaks instead of addressing the autoscaling configuration.

How to eliminate wrong answers

Option A is wrong because optimizing the prompt to reduce output length may lower per-request compute time but does not address the root cause of insufficient concurrent serving capacity during traffic spikes. Option B is wrong because reducing the maximum number of replicas would cap the autoscaler's ability to add instances, worsening the bottleneck and increasing timeouts. Option C is wrong because switching to a GPU-based machine type can accelerate inference per request but does not solve the scaling issue; it may even increase cold-start time and cost without guaranteeing enough replicas to handle peak load.

530
MCQmedium

A financial analyst needs to quickly extract key figures and summarize insights from a 200-page earnings report PDF. They want to use a Google Cloud generative AI model that can process long documents and answer questions. Which Gemini model capability should they leverage?

A.Gemini's multimodal input
B.Gemini's function calling
C.Gemini's long context window
D.Gemini's grounding with Google Search
AnswerC

Gemini models offer a long context window, allowing them to ingest very large documents like a 200-page PDF in a single prompt. This enables the analyst to ask questions and extract figures without chunking the document manually. The model can reason across the entire report, providing accurate summaries and answers. This directly addresses the need for processing long documents efficiently.

Why this answer

Gemini's long context window is designed to handle very large inputs, such as a 200-page PDF, in a single request. This allows the analyst to ask questions and extract insights without manual segmentation. Other features like function calling, multimodality, or grounding address different needs and do not solve the core challenge of processing a long document.

Exam trap

The trap here is assuming that multimodal input alone solves long document processing, when the key enabler is the context window size.

531
MCQmedium

A financial analytics team needs a managed Google Cloud environment to ground Gemini responses in their proprietary market reports and to evaluate model outputs before releasing an internal research assistant. They want minimal infrastructure management and native integration with BigQuery. Which Google Cloud offering should they choose?

A.Gemini for Google Workspace
B.Vertex AI
C.Vertex AI Agent Builder
D.Vertex AI Studio
AnswerB

Vertex AI is Google Cloud's managed end-to-end platform for building, grounding, evaluating, and deploying generative AI. It provides grounding with Vertex AI Search and your own data, evaluation tooling, and direct BigQuery integration, so the team can ground Gemini in proprietary reports and evaluate outputs without managing infrastructure. This matches the requirement for a managed environment with native BigQuery connectivity.

Why this answer

Vertex AI is the managed Google Cloud platform that unifies model access, grounding with enterprise data, evaluation, and deployment. It natively connects to BigQuery and supports grounding Gemini in proprietary content, which directly addresses the need to ground responses and evaluate outputs before release. The other services are either prototyping surfaces, agent-building tools, or productivity assistants that do not deliver the required governed, end-to-end environment.

Exam trap

The trap here is assuming any Gemini-branded tool can ground and evaluate models, when only the managed Vertex AI platform provides the governed grounding, evaluation, and BigQuery integration described.

532
Multi-Selectmedium

A healthcare company is building a generative AI application using Vertex AI. They need to ensure that the application adheres to responsible AI principles, such as avoiding harmful outputs and protecting patient data. Which two Google Cloud features or practices should they implement? (Choose two.)

Select 2 answers
A.Configure safety filters and thresholds in Vertex AI
B.Use Customer-Managed Encryption Keys (CMEK) for data at rest
C.Increase the model's temperature to generate more diverse responses
D.Disable logging of all API requests to protect patient data
E.Fine-tune the model on a dataset of patient interactions without anonymization
AnswersA, B

Safety filters in Vertex AI allow you to set thresholds for categories like hate speech, harassment, and dangerous content. By configuring these, the company can block or reduce harmful outputs from the generative model. This directly addresses the responsible AI principle of avoiding harmful content, making it a correct choice for the scenario.

Why this answer

Configuring safety filters helps prevent harmful outputs, directly supporting the principle of avoiding harm. Using CMEK ensures that patient data is encrypted with keys managed by the company, enhancing data protection and compliance. Together, these features address both output safety and data privacy, which are critical for responsible AI in healthcare.

The other options either increase risk or violate privacy.

Exam trap

The trap here is assuming that any data protection measure, like disabling logging, is good, but responsible AI requires a balance of safety, privacy, and accountability.

533
MCQmedium

After deploying a text-to-image model, the output images often contain distorted objects. The team suspects the prompt is too complex. Which prompt engineering technique should they try first?

A.Increase the guidance scale.
B.Add more descriptive adjectives.
C.Use a negative prompt to exclude distortions.
D.Break the prompt into simpler, separate steps.
AnswerD

Complex prompts overload the model's attention, producing distorted objects. Decomposing the prompt into simpler sequential steps reduces the conditioning burden per generation, letting the model resolve each element correctly before combining them into the final image.

Why this answer

Breaking a complex prompt into simpler, separate steps reduces the cognitive load on the diffusion model, allowing it to focus on generating each element sequentially. This technique, often called 'prompt decomposition' or 'step-by-step prompting,' directly addresses the root cause of distorted objects when the model struggles to attend to multiple conflicting details simultaneously in a single pass.

Exam trap

Candidates often mistakenly think that increasing guidance scale or adding more descriptive details always improves output quality, when in fact these actions can worsen distortions by over-constraining the model's latent space.

How to eliminate wrong answers

Option A is wrong because increasing the guidance scale forces the model to adhere more strictly to the prompt, which can amplify artifacts and distortions rather than reduce them, especially when the prompt is already too complex. Option B is wrong because adding more descriptive adjectives increases prompt complexity, making it harder for the model to disentangle attributes, often leading to more distorted or merged objects. Option C is wrong because using a negative prompt to exclude distortions is a reactive fix that does not address the underlying issue of prompt complexity; it may suppress some artifacts but cannot resolve the model's inability to handle too many simultaneous constraints.

534
MCQmedium

A logistics company wants its operations analysts to ask natural-language questions such as 'Which routes had the most delays last quarter?' and receive answers grounded in data stored in BigQuery, without analysts writing SQL. The team has no plans to build or train custom models. Which Google Cloud offering should they adopt?

A.Vertex AI Pipelines
B.Vertex AI Model Garden
C.Cloud Data Fusion
D.BigQuery data canvas
AnswerD

BigQuery data canvas is a Gemini-powered experience inside BigQuery that lets users explore and analyze data using natural-language prompts, generating SQL and visualizations grounded in the queried tables. Because the analysts' data already lives in BigQuery and they need conversational, no-code analysis without training any model, this is the purpose-built fit. It leverages the underlying Gemini models in BigQuery rather than requiring a custom model deployment.

Why this answer

The analysts need conversational analytics over data already in BigQuery with no model training, which is exactly what BigQuery data canvas delivers through Gemini-assisted natural-language exploration that generates SQL and visualizations. The other services target model cataloging, pipeline orchestration, or data integration, none of which provide prompt-driven querying of BigQuery tables for non-technical users.

Exam trap

The trap here is assuming that any Gemini-powered capability must come from Vertex AI, when BigQuery itself embeds Gemini through data canvas for conversational analytics.

535
MCQeasy

A product team uses Gemini via the Vertex AI API to draft customer emails. The drafts are accurate but often too long and include unnecessary background. The team wants shorter, more direct outputs while keeping the same model. Which approach should they take?

A.Fine-tune the model on a dataset of short emails to permanently change its verbosity.
B.Set the max output tokens parameter to a very low value and leave the prompt unchanged.
C.Revise the prompt to explicitly instruct concise, direct language and specify a target length or format.
D.Increase the top-p value so the model selects from a narrower set of likely tokens.
AnswerC

Prompt instructions are the primary control for output style. Telling the model to be concise, avoid background, and follow a target length or bullet format directly shapes the response. This preserves completeness of key information while meeting the brevity requirement. It works without changing model parameters or retraining.

Why this answer

The most direct and efficient fix for verbose outputs is to change the prompt to request concise, direct language with a specified length or format. Prompt engineering shapes style immediately without retraining or infrastructure changes. Token limits truncate rather than summarize, fine-tuning is overkill for a style tweak, and top-p affects randomness rather than length.

Exam trap

The trap here is reaching for parameter changes like max output tokens or top-p when the real issue is that the prompt never asked for brevity.

536
MCQmedium

A company runs a pilot for a GenAI-powered internal knowledge base assistant. They want to measure adoption. Which metric is BEST for this purpose?

A.Number of follow-up questions asked per session
B.User satisfaction score from surveys
C.Average response time of the assistant
D.Number of unique users per week divided by total employees
AnswerD

Weekly active users divided by total employees yields an adoption rate normalised to headcount, showing what proportion of staff actually use the assistant. This satisfies the requirement to measure adoption, unlike raw query counts or satisfaction scores.

Why this answer

The best metric for measuring adoption because it directly captures the breadth of usage across the organization. Adoption is defined as the proportion of the target user base that actively uses the system, and dividing unique weekly users by total employees provides a clear percentage of uptake, which is the standard measure for adoption in enterprise GenAI deployments.

Exam trap

The trap here is that candidates confuse adoption with engagement or satisfaction, picking metrics like follow-up questions or survey scores, but Google specifically tests that adoption is about the proportion of the target population using the system, not how deeply or happily they use it.

How to eliminate wrong answers

Option A is wrong because the number of follow-up questions per session measures engagement depth or conversational complexity, not adoption; a single power user could inflate this metric while overall adoption remains low. Option B is wrong because user satisfaction scores measure quality or user experience, not adoption; a system can have high satisfaction but low usage if few employees try it. Option C is wrong because average response time measures performance latency, not adoption; a fast assistant is irrelevant if no one uses it.

537
MCQhard

A company is deploying a generative AI model on Vertex AI for real-time chat. They observe that the model sometimes generates toxic or biased responses. They want to implement a safety mechanism to filter out such content before it reaches users. Which Google Cloud feature should they use?

A.Sensitive Data Protection
B.Vertex AI Model Monitoring
C.Cloud Armor
D.Vertex AI Safety Filters
AnswerD

Vertex AI Safety Filters are designed to detect and block harmful content, including toxicity, bias, and explicit material. They can be configured with thresholds and applied to both prompts and responses in real time. This directly addresses the need to filter out toxic or biased outputs before they reach users.

Why this answer

Vertex AI Safety Filters are specifically built to detect and block harmful content in generative AI applications. They can be applied to both input and output, providing real-time protection. Other services like Cloud Armor or Sensitive Data Protection serve different purposes and cannot effectively filter toxic or biased language.

Exam trap

The trap here is confusing security or data privacy services with content safety filters designed for generative AI.

538
MCQmedium

A hospital network wants to build a generative AI search experience over its clinical guideline PDFs so clinicians can ask natural-language questions and receive answers with citations to the source documents. The solution must run on Google Cloud and keep data within the network's project. Which offering is purpose-built for this requirement?

A.Gemini for Google Workspace
B.Cloud Vision API with OCR on the PDFs
C.Cloud Translation API
D.Vertex AI Search with grounding on the guideline corpus
AnswerD

Vertex AI Search is designed to index enterprise documents and return grounded, cited answers to natural-language queries. Ingesting the guideline PDFs into a data store and enabling grounding lets the hospital build the clinician-facing search experience while keeping data inside its own Google Cloud project with enterprise access controls.

Why this answer

Vertex AI Search indexes the guideline documents and returns grounded answers with citations, which is exactly the retrieval-augmented experience described. Vision, Workspace, and Translation services perform extraction, productivity assistance, or language conversion and cannot deliver cited question answering over the hospital's corpus.

Exam trap

The trap here is equating text extraction from PDFs with a complete retrieval and answer-generation solution, when extraction alone produces no answers or citations.

539
MCQeasy

Which Google DeepMind technology can be used to embed an invisible watermark into AI-generated images to help identify their origin?

A.SynthID
B.TensorFlow Privacy
C.PAIR Explorables
D.What-If Tool
AnswerA

SynthID embeds an imperceptible digital watermark directly into generated image pixels, surviving common edits and transformations. This satisfies the stem's requirement to identify AI-generated origin, unlike metadata tagging, which is easily stripped or lost during processing.

Why this answer

SynthID is Google DeepMind's tool for watermarking AI-generated content. The other options are unrelated.

540
MCQeasy

A small marketing team wants to use a Google Cloud generative AI model to generate creative text for social media posts. They prefer a fully managed, ready-to-use API that requires minimal setup and does not need model training. Which Google Cloud service should they use?

A.Vertex AI Gemini API
B.AutoML Natural Language
C.Vision AI
D.Dialogflow CX
AnswerA

The Vertex AI Gemini API provides access to Google's Gemini foundation models through a fully managed API. It requires no model training or infrastructure setup, and developers can start generating text immediately by sending prompts. This matches the team's need for a ready-to-use generative AI service with minimal setup for creative text generation.

Why this answer

The Vertex AI Gemini API offers direct access to powerful generative models without any training or infrastructure management. It is ideal for tasks like creative writing, where the team can simply send prompts and receive generated text. The other services are either for custom model training, image analysis, or conversational agents, and do not provide the same ease of use for text generation.

Exam trap

The trap here is assuming that any AI service can generate text, but many Google Cloud AI services are specialized for vision, language understanding, or conversation, not open-ended text generation.

541
MCQeasy

A data science team wants to run a machine learning model directly on data stored in BigQuery without moving the data to a separate environment. Which Google Cloud service should they use?

A.Vertex AI Training
B.Cloud Dataproc
C.Google AI Studio
D.BigQuery ML
AnswerD

BigQuery ML lets the team create and run machine learning models using SQL directly against data held in BigQuery, so no extraction or movement to a separate environment occurs. This satisfies the constraint of training without relocating the data.

Why this answer

BigQuery ML enables creating and executing ML models using standard SQL queries directly on data in BigQuery, eliminating data movement.

542
Multi-Selectmedium

Which THREE are best practices for responsible deployment of generative AI in a customer-facing application?

Select 3 answers
A.Implement human-in-the-loop review for sensitive outputs
B.Train the model on all available data to maximize coverage
C.Implement content filters to block inappropriate outputs
D.Use only small models to reduce risk
E.Conduct regular bias and fairness audits
AnswersA, C, E

Human review adds accountability and error correction.

Why this answer

Human-in-the-loop (HITL) review ensures that sensitive outputs—such as those involving protected health information (PHI), personally identifiable information (PII), or high-stakes decisions—are vetted by a human before reaching the customer. This mitigates the risk of harmful or biased generations that automated guardrails might miss, aligning with responsible AI principles like accountability and safety.

Exam trap

Google Cloud often tests the misconception that 'more data is always better' or that 'smaller models are safer,' when in fact responsible deployment hinges on data quality, continuous monitoring, and layered safeguards rather than model size or data volume alone.

543
Multi-Selecthard

A logistics company plans to embed a generative AI assistant into its dispatch workflow, where it will draft driver instructions and answer operations questions. The executive sponsor insists that the initiative include mechanisms to measure whether the assistant is delivering business value and to keep its behavior within approved limits as usage grows. (Choose two.)

Select 2 answers
A.Increase the foundation model's temperature setting so the assistant produces more varied dispatch instructions over time.
B.Establish an operating model with human review thresholds, escalation paths, and periodic prompt and policy updates as usage expands.
C.Move all dispatch assistant traffic to a single region to simplify network routing and reduce egress charges.
D.Purchase the largest available context window model so the assistant can accept longer operations queries.
E.Define business-aligned success metrics such as dispatch time saved and reduction in manual corrections, and review them on a recurring cadence.
AnswersB, E

An operating model that specifies when a human must review output, how operators escalate questionable responses, and how prompts and policies are revised over time keeps the assistant inside approved limits as volume grows. This is the governance counterpart to value measurement and directly satisfies the sponsor's requirement to control behavior during scaling.

Why this answer

Value measurement and behavioral control are distinct governance obligations. Business-aligned metrics with recurring reviews show whether the dispatch assistant is producing tangible benefit, while an operating model with human review thresholds, escalation paths, and periodic prompt and policy updates keeps outputs within approved limits as volume grows. Together they satisfy the sponsor's two requirements.

Exam trap

The trap here is treating model configuration choices such as temperature or context window as governance or value mechanisms, when those are tuning parameters that neither prove business benefit nor constrain behavior.

544
MCQeasy

A marketing agency wants to quickly generate original images for social media campaigns without deep technical expertise. They need a fully managed Google Cloud service that provides an API for text-to-image generation. Which service should they use?

A.Cloud Vision API
B.Vertex AI Imagen
C.Dialogflow CX
D.Vertex AI Vision
AnswerB

Vertex AI Imagen is a fully managed text-to-image generation service available through Vertex AI. It provides an API that accepts text prompts and returns generated images, making it accessible for users without deep ML expertise. It is designed for creative tasks like social media content, aligning perfectly with the agency's needs.

Why this answer

Vertex AI Imagen is Google Cloud's managed text-to-image generation service, offering an API that turns text prompts into images. It is designed for creative use cases and requires no ML expertise, making it the right choice. The other services are for vision analysis or conversational AI and do not generate images.

Exam trap

The trap here is confusing image analysis services like Cloud Vision API or Vertex AI Vision with image generation services, when only Imagen provides text-to-image synthesis.

545
MCQhard

A legal team wants to use GenAI to review contracts and highlight risky clauses. They need the AI to consistently follow a specific classification taxonomy. The team has a small set of labeled examples (500 contracts). Which approach yields the BEST accuracy for this use case?

A.Prompt engineer a large foundation model with few-shot examples in Vertex AI Studio
B.Use the RAG Engine to retrieve similar clauses and ask the model to classify
C.Fine-tune a base model using the labeled examples in Vertex AI
D.Use a larger foundation model without fine-tuning and rely on its pre-trained knowledge
AnswerC

Fine-tuning adjusts the base model's weights on the 500 labelled contracts, embedding the specific taxonomy into the model itself. This yields higher, more consistent classification accuracy than prompt engineering or retrieval alone, satisfying the small-labelled-dataset constraint.

Why this answer

Fine-tuning a base model on the 500 labeled contracts teaches the model the specific classification taxonomy directly in its weights, producing consistent, high-accuracy outputs that follow the exact label set. With a small but well-labeled dataset, supervised fine-tuning is the most reliable way to enforce a fixed taxonomy rather than relying on prompt instructions.

Exam trap

Generative AI Leader often tests the confusion between RAG (good for grounding on external knowledge) and fine-tuning (good for teaching a specific output format or taxonomy), causing candidates to pick RAG when consistency of classification is the requirement.

How to eliminate wrong answers

Option A is wrong because few-shot prompting, while useful, is less consistent for a strict taxonomy and is limited by context window and prompt sensitivity — it does not internalize the label definitions. Option B is wrong because RAG retrieves similar clauses for context but does not train the model on the taxonomy; the model may still misclassify or drift from the required labels. Option D is wrong because a larger base model without fine-tuning has no knowledge of the organization's specific taxonomy and will produce inconsistent labels.

546
MCQeasy

Which tool from Google's Responsible AI toolkit is designed to document the intended use, performance, and limitations of a machine learning model?

A.PAIR Explorables
B.Datasheets for Datasets
C.People + AI Guidebook
D.Model Cards
AnswerD

Model Cards provide structured documentation covering a model's intended use, performance metrics, and known limitations, directly satisfying the stem's requirement to record these three attributes. Unlike data-focused tools such as Datasheets, Model Cards address the model itself, making them the designated artefact within Google's Responsible AI toolkit for this purpose.

Why this answer

Model Cards are the correct tool because they are specifically designed to document the intended use, performance metrics, and limitations of a machine learning model. This standardized documentation format, introduced by Google, provides transparency by detailing evaluation results across different conditions, intended use cases, and known biases, which is essential for responsible AI deployment.

Exam trap

The Generative AI Leader exam often tests the distinction between tools that document datasets (Datasheets for Datasets) versus tools that document models (Model Cards), leading candidates to confuse the two when the question specifically asks about documenting a machine learning model.

How to eliminate wrong answers

Option A is wrong because PAIR Explorables are interactive articles and visualizations designed to help people understand and explore concepts in machine learning and AI, not to document a model's intended use, performance, and limitations. Option B is wrong because Datasheets for Datasets are focused on documenting the characteristics, collection process, and intended uses of datasets, not the machine learning model itself. Option C is wrong because the People + AI Guidebook is a set of design guidelines and patterns for building human-centered AI products, not a documentation tool for model specifications and limitations.

547
Multi-Selecthard

A telecommunications provider is preparing a business case for a generative AI virtual assistant that handles billing inquiries. Executives want the proposal to address financial viability, not just technical feasibility. Which two elements should the business case include to demonstrate responsible financial planning? (Choose two.)

Select 2 answers
A.The number of prompt variants tested during internal prototyping.
B.A model of cost per resolved inquiry that includes inference, grounding, and human escalation.
C.A list of every foundation model available in Vertex AI Model Garden.
D.A diagram of the virtual private cloud network topology used by the assistant.
E.A projected reduction in average handling cost with assumptions and sensitivity ranges.
AnswersB, E

A cost-per-resolved-inquiry model captures the true unit economics of the assistant by combining inference charges, retrieval or grounding costs, and the expense of escalations to human agents. Because not every conversation is fully automated, ignoring escalations would overstate savings. This element gives executives a defensible basis for comparing the assistant against current billing support costs and for setting a target automation rate that keeps the program financially viable.

Why this answer

A credible financial case for the billing assistant needs unit economics and a forward-looking benefit estimate. Cost per resolved inquiry ties inference, grounding, and human escalation into one comparable figure, while projected handling-cost reduction with sensitivity ranges shows how savings vary under different automation and volume assumptions. Together they let executives judge viability and set targets, whereas model inventories, prompt counts, and network diagrams do not quantify financial impact.

Exam trap

The trap here is filling a business case with technical artifacts such as model lists, prompt counts, or network diagrams instead of quantified unit costs and sensitivity-tested savings.

548
MCQmedium

A company is building a search application that requires grounding answers in their internal knowledge base. They want to use Vertex AI Search and Conversation with a custom datastore. Which configuration is essential to ensure the model only answers based on their documents?

A.Enable streaming responses to get real-time answers.
B.Fine-tune the model on the company's documents.
C.Configure the answer generation to use grounding with the enterprise datastore as the source.
D.Set the model's temperature to 0 to make responses deterministic.
AnswerC

Grounding binds generated answers to retrieved passages from the specified enterprise datastore, so the model cites and answers only from those documents rather than its pretrained knowledge. Pointing answer generation at that datastore as the grounding source is therefore essential to constrain responses to internal content.

Why this answer

Vertex AI Search and Conversation provides a built-in grounding capability that explicitly ties answer generation to a specified enterprise datastore. By configuring grounding with the custom datastore as the source, the model is constrained to retrieve and synthesize answers exclusively from the indexed documents, preventing reliance on its parametric knowledge or external sources.

Exam trap

Google Cloud often tests the distinction between techniques that influence output style (temperature, streaming) versus those that control knowledge sources (grounding), leading candidates to confuse deterministic generation with factual grounding.

How to eliminate wrong answers

Option A is wrong because enabling streaming responses controls the delivery mechanism (real-time token-by-token output) but does not restrict the model's knowledge source; it can still generate answers from its training data. Option B is wrong because fine-tuning adapts the model's weights to the company's documents, which can improve relevance but does not guarantee grounding—the model may still hallucinate or use pre-training knowledge, and Vertex AI Search does not require fine-tuning for retrieval-augmented generation. Option D is wrong because setting temperature to 0 makes responses deterministic (low randomness) but does not enforce grounding; the model can still confidently produce incorrect answers from its internal knowledge.

549
MCQmedium

A healthcare company needs to generate synthetic medical images for research while ensuring compliance with patient privacy regulations. Which Google Cloud generative AI service should they use?

A.Codey for code generation
B.Chirp for speech recognition
C.Imagen on Vertex AI
D.Gemini 1.5 Pro with multimodal prompting
AnswerC

Imagen on Vertex AI generates synthetic images from text prompts, so the company can create research imagery without exposing real patient data. Vertex AI's enterprise controls and data-handling commitments support the privacy compliance constraint, unlike general-purpose image tools lacking healthcare-grade governance.

Why this answer

Imagen on Vertex AI is Google's image generation service that can create synthetic images and offers controls for responsible AI and data governance.

550
MCQhard

A legal firm wants to automate contract analysis. They need to extract key clauses (e.g., termination, indemnification) from scanned PDFs. The team expects high accuracy and must maintain data privacy. Which combination of services is most suitable?

A.Use AutoML Tables to train a classification model on text features
B.Use Document AI for OCR and Vertex AI with a custom fine-tuned model for clause extraction
C.Use Gemini API directly with a prompt to analyze PDFs
D.Use AppSheet to create a form for manual entry and then use BigQuery ML
AnswerB

Document AI performs OCR on scanned PDFs, converting them to structured text, while Vertex AI trains a fine-tuned model on the firm's clause examples for accurate extraction. Processing stays within the firm's Google Cloud project, satisfying the data privacy constraint.

Why this answer

Document AI performs OCR and extracts text from scanned PDFs; Vertex AI with a custom fine-tuned model provides high accuracy for clause extraction while keeping data within the customer's project.

551
MCQeasy

A retail company wants to build an internal assistant that answers employee questions using the company's own HR policy documents. The team has no machine learning engineers and wants a managed Google Cloud approach that grounds responses in those documents without training a new foundation model. Which Google Cloud capability best fits this need?

A.Deploying an open-source model on a Compute Engine GPU VM and fine-tuning it weekly
B.Vertex AI Search with grounding on a data store built from the HR documents
C.Using BigQuery ML to run a logistic regression over HR ticket categories
D.Training a custom foundation model from scratch on the HR documents
AnswerB

Vertex AI Search lets teams index enterprise documents into a data store and then ground generative responses on retrieved passages. It is a managed service requiring no model training, and it directly addresses the requirement to answer from the company's own HR policies. This matches the no-ML-engineer constraint and the grounding goal precisely.

Why this answer

Vertex AI Search provides managed retrieval over enterprise documents and grounds generative answers in that content, so employees receive responses based on current HR policy rather than the model's general knowledge. It requires no foundation-model training and minimal ML operations, matching the team's skills and timeline. The other choices either demand heavy ML work or address unrelated prediction tasks.

Exam trap

The trap here is assuming grounding requires fine-tuning a model, when retrieval-augmented grounding through a managed search service is the intended no-training approach.

552
Multi-Selectmedium

An enterprise is evaluating whether to build a custom fine-tuned model or use a pre-built API for code generation. Which three factors should they consider in the build vs. buy decision? (Choose THREE)

Select 3 answers
A.Data privacy and security requirements
B.Number of developer seats in the organization
C.Compatibility with on-premises legacy systems
D.Level of customization needed for the organization's coding standards
E.Availability of pre-built models for the specific programming language
AnswersA, D, E

Where source code cannot leave the tenant or reach third-party endpoints, a hosted API may be disqualified, pushing the decision toward a self-hosted or fine-tuned model. Data privacy and security requirements thus directly constrain the build-versus-buy choice.

Why this answer

Option A (Data privacy and security requirements) is correct because a build decision is often driven by the need to keep proprietary source code and sensitive data within the organization's own environment, whereas a pre-built API may send prompts and code to a third-party provider, raising data-residency, retention, and compliance concerns. Option D (Level of customization needed for the organization's coding standards) is correct because fine-tuning a custom model allows tailoring to internal style guides, frameworks, and domain-specific APIs, while a generic pre-built API may not align with those standards. Option E (Availability of pre-built models for the specific programming language) is correct because if a suitable pre-built model already supports the target language well, buying is more efficient, whereas a lack of adequate pre-built support pushes the decision toward building.

Option B (Number of developer seats) is not a primary build-vs-buy factor because seat count mainly affects licensing cost and scaling, not the fundamental capability or control trade-off. Option C (Compatibility with on-premises legacy systems) is not a core factor because integration can typically be handled through APIs or gateways regardless of whether the model is built or bought.

Exam trap

The trap is that candidates pick operational or cost-related factors (developer seats, legacy system compatibility) instead of the strategic factors (privacy, customization, model availability) that actually drive the build-vs-buy decision.

553
MCQmedium

During a proof-of-concept for a GenAI document summarization tool, the team wants to evaluate whether the summaries are accurate and retain key information before scaling. Which evaluation approach is most appropriate for this stage?

A.Measure latency and cost as the primary evaluation metrics
B.Deploy to all users and collect feedback via a survey
C.Run an A/B test with a small user group and have domain experts manually review a sample of summaries for accuracy
D.Use ROUGE scores exclusively to compare summaries against human-written ones
AnswerC

At proof-of-concept stage, accuracy and key-information retention matter more than scale metrics. Expert manual review of a sampled subset directly measures those qualities, whereas a small A/B test captures user preference, not factual fidelity, and is premature before quality is validated.

Why this answer

A/B testing with manual review by domain experts provides qualitative and quantitative feedback on accuracy and completeness, which is crucial for a pilot. Automated metrics alone may not capture business relevance.

554
Multi-Selectmedium

A team wants to reduce hallucinations in a question-answering model. Which THREE techniques should they consider?

Select 3 answers
A.Fine-tune the model on a curated factual dataset
B.Use retrieval-augmented generation (RAG)
C.Apply prompt engineering with specific instructions to cite sources
D.Reduce the number of tokens in output
E.Increase the temperature parameter
AnswersA, B, C

Fine-tuning on factual data improves accuracy.

Why this answer

Fine-tuning on a curated factual dataset directly adjusts the model's weights to prioritize accurate, domain-specific knowledge, reducing the likelihood of generating unsupported or hallucinated content. This technique anchors the model's output in verified data, making it more reliable for question-answering tasks.

Exam trap

Google Cloud often tests the misconception that reducing output length or increasing randomness (temperature) can improve factual accuracy, when in reality these parameters control style and creativity, not truthfulness.

555
MCQmedium

A financial technology company has deployed a custom-tuned PaLM 2 model on Vertex AI to generate personalized investment recommendations for retail clients. The model was fine-tuned on a corpus of historical market data and advisory transcripts. Recently, the compliance team flagged that several recommendations contradicted SEC guidelines, and the model sometimes repeated prohibited statements from outdated training materials. The team has already implemented safety filters (e.g., blocking toxic content) and adjusted the model's system instructions to be more conservative. However, the issues persist. The model's deployment parameters are: temperature=0.4, top_p=0.9, max_output_tokens=500, and no grounding. The company must maintain compliance without significantly increasing latency. What should they do next?

A.Increase temperature to 0.7 to allow more diverse responses, and add a second model to verify outputs
B.Perform an additional fine-tuning round exclusively on the most recent SEC regulatory filings and compliance-approved content
C.Implement a chain-of-thought prompting technique that requires the model to explain its reasoning step by step
D.Configure Vertex AI grounding using a curated data store of real-time SEC regulations and market data
AnswerD

Grounding with a curated Vertex AI data store anchors generation to current SEC regulations, directly addressing the prohibited statements and contradictions sourced from outdated training materials. Unlike fine-tuning, retrieval injects authoritative text at inference time, so compliance updates take effect immediately without retraining. This satisfies the compliance constraint while adding only modest retrieval latency.

Why this answer

Configuring Vertex AI grounding with a curated data store of real-time SEC regulations directly addresses the root cause: the model is generating outputs that contradict current compliance rules. Grounding forces the model to base its responses on authoritative, up-to-date sources, which is more effective than safety filters or system instructions alone, and it avoids the latency increase of a second model or the risk of catastrophic forgetting from additional fine-tuning.

Exam trap

This exam often tests the misconception that fine-tuning or prompt engineering alone can solve compliance issues, when in fact grounding with authoritative data sources is the only reliable method for ensuring outputs adhere to real-time, external regulations without sacrificing latency.

How to eliminate wrong answers

Option A is wrong because increasing temperature to 0.7 would make outputs more random and less deterministic, increasing the likelihood of generating non-compliant statements, and adding a second model for verification would significantly increase latency and cost without fixing the underlying data contamination. Option B is wrong because performing additional fine-tuning on recent SEC filings risks catastrophic forgetting of the original training data and does not guarantee real-time compliance, as fine-tuning is static and cannot adapt to rapidly changing regulations. Option C is wrong because chain-of-thought prompting only improves reasoning transparency but does not constrain the model to use compliant sources; the model could still generate prohibited statements from its outdated training data.

556
MCQmedium

A product team at a retailer is using Vertex AI Studio to build a Gemini-powered assistant that answers questions about their internal return policy. Early tests show the model invents policy details such as a 45-day return window, even though the official policy allows only 30 days. The team wants the assistant to answer strictly from a set of approved policy PDFs stored in a Cloud Storage bucket, and they want to avoid retraining the model. Which technique should they use?

A.Adding a system instruction telling the model to be accurate and to never hallucinate policy details.
B.Supervised fine-tuning of the Gemini model on a labeled dataset of correct policy answers.
C.Lowering the model's temperature to 0 and increasing the top-k value in the generation configuration.
D.Retrieval-augmented generation (RAG) with a Vertex AI Search data store grounded on the approved policy documents.
AnswerD

RAG retrieves relevant passages from the indexed policy PDFs at query time and passes them to Gemini as grounding context, so answers reflect the approved 30-day policy instead of the model's pretrained assumptions. Vertex AI Search provides managed indexing and grounding, and no model retraining is required, which matches the team's constraint.

Why this answer

Grounding with retrieval-augmented generation lets the assistant fetch the relevant passages from the approved policy PDFs and condition its answer on that retrieved text, so the 30-day window is stated correctly and the response can cite the source. It requires no retraining, updates automatically when documents change, and directly addresses hallucinated policy details.

Exam trap

The trap here is assuming that lowering temperature or adding a firm instruction eliminates hallucination, when only supplying authoritative source content through grounding actually fixes a missing factual detail.

557
Multi-Selectmedium

A retail company wants to generate personalized marketing content (emails, social posts) at scale using generative AI. They need consistent brand voice and the ability to review outputs before publishing. Which two Google Cloud capabilities should they use? (Choose TWO)

Select 2 answers
A.Google Workspace Duet AI in Docs for drafting, then manual review
B.AutoML Tables for predicting customer segments
C.Vertex AI Agent Builder with grounding and human-in-the-loop
D.Pre-trained Gemini model via Vertex AI API with no customization
E.Cloud Vision API for image analysis
AnswersA, C

Duet AI can assist in drafting content quickly, and manual review ensures brand alignment.

Why this answer

Google Workspace Duet AI in Docs allows marketers to draft personalized content using generative AI while maintaining control over brand voice through iterative editing. The manual review step ensures outputs meet quality and compliance standards before publishing, addressing the need for human oversight in content generation.

Exam trap

The trap here is that candidates may confuse AutoML Tables (a Google Cloud predictive modeling tool) with generative AI capabilities, or overlook that pre-trained Gemini models without customization fail to meet brand voice requirements, while Cloud Vision API is irrelevant to text generation tasks. Candidates might also mistakenly think that Duet AI in Docs is not suitable for personalized marketing at scale, but it allows iterative editing and manual review to maintain brand voice.

558
MCQmedium

A financial institution is deploying a generative AI chatbot to provide investment advice. According to regulatory requirements, high-stakes AI decisions must have human review. Which setup BEST satisfies this requirement?

A.Use a separate AI model to review the first model's outputs and flag issues
B.AI generates recommendations, but a human advisor must review and approve before any action, and can override the AI
C.AI provides advice directly to the customer with a disclaimer that it is not financial advice
D.Allow the AI to execute trades automatically, with an audit log for later review
AnswerB

Human-in-the-loop approval directly satisfies the regulatory mandate for human review of high-stakes decisions. The advisor gates every recommendation before action and retains override authority, so no investment advice reaches the client without documented human judgement, meeting the requirement for meaningful oversight rather than post-hoc monitoring.

Why this answer

It directly implements the regulatory requirement for human-in-the-loop (HITL) oversight in high-stakes AI decisions. In financial advisory contexts, regulations like the EU AI Act or SEC guidelines mandate that a qualified human advisor must review and approve AI-generated recommendations before any action is taken, ensuring accountability and the ability to override erroneous outputs.

Exam trap

Google often tests the distinction between 'human review' and 'human oversight' — candidates mistakenly think that an audit log or a disclaimer satisfies the requirement, but the trap is that regulators require proactive human approval before the action occurs, not after.

How to eliminate wrong answers

Option A is wrong because using a separate AI model to review outputs merely replaces one automated system with another, failing to satisfy the regulatory requirement for human review; this is a form of 'AI oversight of AI' that does not provide the necessary human accountability. Option C is wrong because a disclaimer does not constitute human review; the AI is still directly providing advice to the customer without any human intervention, which violates the requirement for high-stakes decisions. Option D is wrong because allowing the AI to execute trades automatically with only an audit log for later review is a 'human-out-of-the-loop' approach; post-hoc auditing does not prevent harm from occurring in real time, and regulators require proactive human approval before execution.

559
MCQeasy

A startup is building a customer service chatbot that generates responses in real-time. They want the model to have up-to-date information on the latest product catalog but cannot afford frequent fine-tuning. Which technique should they use to inject current data into the model without retraining?

A.Rely on the model's zero-shot capabilities to infer product details.
B.Use retrieval-augmented generation (RAG) to fetch relevant documents from a vector database at inference time.
C.Craft detailed system prompts that include the entire product catalog in the prompt.
D.Fine-tune the base model weekly on the latest product catalog.
AnswerB

RAG decouples knowledge from model weights: relevant catalogue documents are retrieved from a vector database and injected into the prompt at inference time. This satisfies the constraint of current data without the cost or delay of frequent fine-tuning.

Why this answer

Retrieval-Augmented Generation (RAG) is the correct technique because it allows the chatbot to fetch the most current product catalog entries from an external vector database at inference time, without requiring any model retraining. This keeps responses grounded in up-to-date information while avoiding the cost and latency of frequent fine-tuning.

Exam trap

Google Cloud often tests the distinction between in-context learning (via RAG or prompt engineering) and parametric knowledge (via fine-tuning), trapping candidates who think that simply adding more data to the prompt is scalable or that zero-shot inference can substitute for external retrieval.

How to eliminate wrong answers

Option A is wrong because zero-shot capabilities rely solely on the model's pre-existing knowledge, which cannot incorporate new or updated product catalog details without retraining. Option C is wrong because crafting detailed system prompts with the entire product catalog would exceed the model's context window limits and incur high token costs, making it impractical for real-time inference. Option D is wrong because fine-tuning weekly is expensive, time-consuming, and contradicts the requirement to avoid frequent retraining; it also risks catastrophic forgetting of previously learned information.

560
MCQmedium

An ML engineer sees the above deployment output. The business wants to reduce inference cost. Which action should they take?

A.Use a larger model
B.Change to a lower-cost machine type
C.Deploy to multiple regions
D.Increase traffic split
AnswerB

Inference cost scales with the compute SKU hosting the deployed model. Selecting a lower-cost machine type reduces the hourly rate charged for the endpoint while preserving the same model, directly satisfying the business goal of lowering inference cost.

Why this answer

Switching to a lower-cost machine type directly reduces the per-request compute cost without altering the model architecture or inference logic. This is a common cost-optimization strategy in cloud-based ML deployments, where instance types (e.g., from GPU to CPU or from a larger to a smaller GPU) can be selected based on latency and throughput requirements, provided the model fits within the machine's memory and compute constraints.

Exam trap

Google Cloud often tests the misconception that 'more resources' (larger model, more regions) always improves performance, but here the business goal is cost reduction, so the correct action is to downsize infrastructure while maintaining acceptable quality.

How to eliminate wrong answers

Option A is wrong because using a larger model increases both memory footprint and compute operations per inference, which raises cost and latency—the opposite of the business goal. Option C is wrong because deploying to multiple regions adds infrastructure overhead, data transfer costs, and management complexity, increasing rather than reducing inference cost. Option D is wrong because increasing traffic split (e.g., routing more requests to a shadow or canary deployment) does not reduce cost; it may increase resource utilization or require additional compute capacity.

561
MCQhard

A media company plans to add generative AI features to its video editing suite. Executives want to prove business value within one quarter but cannot predict which of three candidate features editors will actually adopt. Which strategy best balances fast validation against wasted investment?

A.Build all three features to production quality in parallel so the strongest one wins on merit.
B.Ship a thin prototype of the highest-uncertainty feature to a small editor group and measure adoption before committing further.
C.Select the feature with the lowest estimated engineering cost and take it directly to general availability.
D.Run a company-wide survey asking editors to rank the three features by preference before writing any code.
AnswerB

A thin prototype focused on the riskiest assumption generates real usage evidence quickly and cheaply. Measuring adoption with a small editor cohort tells the company whether to invest, pivot, or stop, so the quarter produces a decision rather than a large sunk cost. This is the core value of time-boxed experimentation in generative AI portfolios.

Why this answer

When demand is genuinely uncertain, the fastest path to business value is a small experiment against the riskiest assumption, measured with real users. A thin prototype of the most uncertain feature yields adoption data within the quarter and converts an unknown into a go, pivot, or stop decision. Parallel production builds, preference surveys, and cost-led launches all postpone or avoid that evidence.

Exam trap

The trap here is equating speed with shipping the cheapest or most complete build, when speed to validated learning is what actually de-risks the investment.

562
MCQeasy

A retail company wants to build a chatbot that answers product questions and provides personalized recommendations. They have a small labeled dataset and limited ML expertise. Which approach should they take?

A.Fine-tune Gemini with their product data using Vertex AI Generative AI Studio.
B.Build a custom transformer model using TensorFlow on Vertex AI Workbench.
C.Use BigQuery ML to train a classification model on customer queries.
D.Use Vertex AI Agent Builder with a pre-built agent and integrate their product catalog via Search and Conversation.
AnswerD

Vertex AI Agent Builder supplies pre-built agent scaffolding and Search and Conversation grounding, so the retailer needs no model training. This satisfies the small labelled dataset and limited ML expertise constraints, since retrieval over their product catalogue replaces fine-tuning.

Why this answer

Vertex AI Agent Builder provides a pre-built agent framework that integrates with Search and Conversation, allowing the company to quickly deploy a chatbot using their product catalog without needing extensive ML expertise. This approach leverages Google's foundation models and retrieval-augmented generation (RAG) to answer product questions and generate personalized recommendations, making it ideal for a small labeled dataset and limited ML resources.

Exam trap

Google Cloud often tests the misconception that fine-tuning or custom model building is necessary for domain-specific tasks, when in fact pre-built agent frameworks with RAG can achieve the same goal with far less data and expertise.

How to eliminate wrong answers

Option A is wrong because fine-tuning Gemini with a small labeled dataset risks overfitting and requires significant ML expertise to manage the fine-tuning pipeline, which the company lacks. Option B is wrong because building a custom transformer model from scratch using TensorFlow on Vertex AI Workbench demands deep ML expertise and large datasets, contradicting the company's constraints. Option C is wrong because BigQuery ML is designed for structured data classification (e.g., SQL-based models), not for building conversational chatbots that handle natural language queries and recommendations.

563
MCQeasy

A retailer wants to use generative AI to write product descriptions automatically. They have a large dataset of existing product descriptions and need to customize a foundation model for their brand voice. Which Vertex AI feature should they use?

A.Vertex AI Search with grounding
B.Prompt design with the Gemini API directly
C.Vertex AI Model Evaluation
D.Vertex AI custom model tuning
AnswerD

Custom model tuning adapts a foundation model's weights to a proprietary dataset, embedding the retailer's brand voice into generated descriptions. This supervised fine-tuning satisfies the requirement to customise the model rather than rely on prompt engineering alone.

Why this answer

Vertex AI custom model tuning (Option D) is correct because it allows the retailer to fine-tune a foundation model on their proprietary dataset of existing product descriptions, adapting the model's output to match their specific brand voice and style. This process adjusts the model's weights using supervised learning on the retailer's data, enabling personalized and consistent content generation that generic prompt engineering cannot achieve.

Exam trap

Google often tests the distinction between prompt engineering (which only changes input instructions) and model tuning (which modifies the model's internal parameters), leading candidates to mistakenly choose prompt design when deep customization is required.

How to eliminate wrong answers

Option A is wrong because Vertex AI Search with grounding is designed for enterprise search and retrieval-augmented generation (RAG) to ground responses in specific data sources, not for fine-tuning a model to adopt a brand voice. Option B is wrong because prompt design with the Gemini API directly only modifies the input instructions without altering the underlying model weights, which is insufficient for deeply customizing the model's writing style to a unique brand voice. Option C is wrong because Vertex AI Model Evaluation is a tool for assessing model performance and detecting issues like bias or drift, not for training or customizing a model's output behavior.

564
MCQeasy

A company is evaluating whether to build a custom generative AI solution from scratch or use a pre-built API from a cloud provider. Which factor most strongly supports the build-from-scratch approach?

A.The team has limited machine learning expertise.
B.Speed to market is the top priority.
C.Minimizing initial development cost is critical.
D.The solution requires deep integration with proprietary data and unique domain-specific outputs.
AnswerD

Building from scratch allows fine-tuning or training on proprietary datasets, giving the model access to unique domain vocabulary and logic that a generic pre-built API cannot replicate. This directly satisfies the stem's constraint of deep proprietary-data integration and unique domain-specific outputs.

Why this answer

Building a custom generative AI solution from scratch is most strongly supported when deep integration with proprietary data and unique domain-specific outputs is required. Pre-built APIs are typically trained on general data and may not capture the nuances of specialized domains, whereas a custom model can be fine-tuned or trained from scratch on proprietary datasets to achieve higher accuracy and relevance for unique business needs.

Exam trap

The trap here is that candidates may confuse 'minimizing cost' (Option C) with long-term total cost of ownership, but The Generative AI Leader exam specifically tests the immediate strategic driver for build vs. buy, which is the need for proprietary data integration and unique outputs.

How to eliminate wrong answers

Option A is wrong because limited ML expertise would favor using a pre-built API to avoid the complexity of model training, infrastructure management, and hyperparameter tuning. Option B is wrong because speed to market is a key advantage of pre-built APIs, which offer immediate access to generative capabilities without the months of development required for a custom solution. Option C is wrong because minimizing initial development cost typically favors pre-built APIs, which have lower upfront investment compared to the significant costs of data preparation, compute resources, and specialized talent needed for building from scratch.

565
MCQmedium

A developer wants to build a RAG application using Vertex AI. Which vector database is natively integrated with Vertex AI for storing embeddings?

A.Firestore
B.Vertex AI Vector Search
C.Cloud SQL
D.Bigtable
AnswerB

Vertex AI Vector Search is Google Cloud's native vector store, purpose-built for the Vertex AI ecosystem, so embeddings generated by Vertex AI models can be indexed and queried without custom integration work. It satisfies the stem's requirement for a natively integrated vector database, unlike third-party options such as Pinecone or open-source alternatives.

Why this answer

Vertex AI Vector Search is the native vector database integrated with Vertex AI for storing and querying embeddings. It is purpose-built for high-dimensional vector similarity search, enabling efficient retrieval in RAG applications without requiring external infrastructure.

Exam trap

Google Cloud often tests the misconception that any database can store embeddings equally well, but the key differentiator is native vector indexing and ANN search support, which only Vertex AI Vector Search provides among the listed options.

How to eliminate wrong answers

Option A is wrong because Firestore is a NoSQL document database designed for storing structured data, not optimized for vector similarity search or embedding storage. Option C is wrong because Cloud SQL is a relational database service (MySQL, PostgreSQL, SQL Server) that lacks native vector indexing and similarity search capabilities required for RAG. Option D is wrong because Bigtable is a wide-column NoSQL database for large-scale analytical workloads, not designed for low-latency vector similarity queries.

566
MCQmedium

A developer deployed a large language model on Vertex AI for real-time chat. Users report slow response times. The model generates sentences one word at a time. Which optimization should be applied to reduce latency?

A.Batch multiple user queries together.
B.Deploy the model with more accelerators.
C.Enable prompt caching to reuse previous queries.
D.Use streaming responses to start output earlier.
AnswerD

Autoregressive generation emits tokens sequentially, so the full response must complete before any output appears. Streaming returns tokens as they are produced, letting the client render the first words immediately and cutting perceived latency without changing the model.

Why this answer

Streaming responses allow the model to send tokens to the client as they are generated, rather than waiting for the full sequence to complete. This reduces perceived latency significantly in real-time chat, as users see the first word appear almost immediately, even though the total generation time remains similar.

Exam trap

The trap here is that candidates often confuse throughput optimization (batching or more accelerators) with latency reduction, failing to recognize that streaming directly minimizes the time users wait for the first visible output in real-time scenarios.

How to eliminate wrong answers

Option A is wrong because batching multiple user queries together increases latency for individual requests, as the system waits to accumulate enough queries before processing, which is counterproductive for real-time chat. Option B is wrong because deploying with more accelerators improves throughput and total generation speed, but does not address the fundamental issue of word-by-word generation latency; the model still outputs one token at a time, and the user must wait for the full response. Option C is wrong because prompt caching reuses previous queries to avoid recomputation, but this optimization targets repeated or similar prompts, not the latency of generating a new response token-by-token.

567
MCQmedium

A company wants to build a chatbot that answers questions based on internal documents. Which approach is most appropriate?

A.Use a pre-trained model without any customizations
B.Train a custom model from scratch
C.Fine-tune a model on the documents
D.Use a prompt with the documents in the context
AnswerD

Placing the internal documents directly in the prompt context lets the model condition its answer on that retrieved text, grounding responses in company content without retraining. This suits document question answering where the corpus changes and fine-tuning would be impractical.

Why this answer

Retrieval-Augmented Generation (RAG) allows the chatbot to dynamically include relevant internal documents in the prompt context without modifying the underlying model. This approach leverages the pre-trained model's language understanding while grounding answers in specific, up-to-date internal data, avoiding the cost and latency of fine-tuning or retraining.

Exam trap

Google Cloud often tests the misconception that fine-tuning is the only way to incorporate proprietary data, but RAG is the most appropriate for dynamic, retrieval-based Q&A because it avoids retraining and keeps the model's knowledge current.

How to eliminate wrong answers

Option A is wrong because a pre-trained model without customization lacks access to the company's internal documents, leading to hallucinated or generic answers not grounded in proprietary data. Option B is wrong because training a custom model from scratch is computationally prohibitive and unnecessary; it requires massive labeled datasets and resources, whereas RAG achieves the same goal with far less effort. Option C is wrong because fine-tuning on documents teaches the model to memorize specific content, which is inefficient for large, frequently updated document sets and risks catastrophic forgetting, whereas RAG keeps the model static and retrieves fresh context per query.

568
MCQhard

An enterprise customer needs to ensure that all data sent to the Gemini API is not used by Google for model improvement and must support a HIPAA BAA. Which access tier should they use?

A.Google AI Studio (free tier)
B.Vertex AI (enterprise tier)
C.Gemini API via Google Workspace
D.Gemini API via API key (developer tier)
AnswerB

Vertex AI operates under Google Cloud's enterprise terms, which exclude customer data from model improvement and are covered by a HIPAA BAA. This satisfies both the no-training-on-data and BAA constraints that the consumer Gemini API tier cannot meet.

Why this answer

Vertex AI (enterprise tier) is the Google Cloud platform that provides data governance commitments, including that customer data is not used for model improvement, and supports HIPAA BAA coverage. Google AI Studio and the developer-tier Gemini API via API key do not offer these enterprise guarantees. Therefore, the enterprise customer must use Vertex AI.

Exam trap

Generative AI Leader often tests the misconception that all Gemini API access tiers offer the same data governance — candidates must know that only Vertex AI provides enterprise no-training commitments and HIPAA BAA support.

How to eliminate wrong answers

Option A is wrong because Google AI Studio (free tier) is a prototyping environment where data may be used to improve Google models and no HIPAA BAA is available. Option C is wrong because Gemini API via Google Workspace is not the enterprise data-governance tier for custom applications and does not provide the BAA and no-training guarantees required. Option D is wrong because the Gemini API via API key (developer tier) is a self-serve tier without enterprise data protection commitments or HIPAA BAA coverage.

569
Multi-Selectmedium

A healthcare startup is building a generative AI application to draft patient education materials. They need to ensure the outputs are accurate, up-to-date, and tailored to each patient's condition. Which two techniques should they use to ground the model's responses in reliable medical knowledge? (Choose two.)

Select 2 answers
A.Retrieval-augmented generation (RAG) using a curated medical knowledge base.
B.Fine-tuning the model on a large dataset of historical patient education materials.
C.Using a model with a larger context window to include more patient history.
D.Prompt engineering with few-shot examples of desired outputs.
E.Implementing a fact-checking layer that queries a trusted medical API for validation.
AnswersA, E

RAG retrieves relevant documents from a trusted knowledge base and provides them as context to the model, ensuring responses are grounded in authoritative, current information. This reduces hallucinations and allows tailoring to specific conditions by retrieving patient-specific or condition-specific guidelines. It is a core technique for building accurate generative AI applications in specialized domains.

Why this answer

Retrieval-augmented generation (RAG) grounds the model by providing relevant, trusted documents as context, while a fact-checking layer validates outputs against a reliable API. Together, they ensure accuracy and currency. Fine-tuning, few-shot prompting, and larger context windows do not inherently guarantee factual correctness or access to the latest medical knowledge.

Exam trap

The trap here is assuming that fine-tuning or a larger context window alone can ground a model in reliable knowledge, when retrieval and validation are needed for accuracy and currency.

570
MCQhard

A data scientist wants to generate photorealistic images of products from text descriptions for an e-commerce catalog. The images must be brand-consistent and avoid generating distorted product features. Which Google Cloud generative AI service should they use?

A.Veo
B.Imagen
C.Chirp
D.Gemini Pro Vision
AnswerB

Imagen on Vertex AI generates photorealistic images from text prompts, directly satisfying the catalogue requirement. Its controls for brand consistency and product fidelity reduce distorted features, unlike general-purpose or conversational models. This makes it the appropriate Google Cloud service for producing reliable e-commerce product imagery at scale.

Why this answer

Imagen is Google's text-to-image model that produces high-quality, photorealistic images. It is designed for brand consistency and safe image generation.

571
MCQmedium

A company wants to build a customer support chatbot that answers based on internal documentation. They use Vertex AI Search and want to ensure the model only uses retrieved documents. What should they do?

A.Fine-tune the model on the documentation
B.Enable grounding with Vertex AI Search
C.Increase max output tokens
D.Set temperature to 0.0
AnswerB

Grounding with Vertex AI Search constrains generation to retrieved internal documents, satisfying the requirement that answers derive only from that corpus. The model cites source passages rather than relying on parametric memory, which reduces hallucination. This directly meets the stem's constraint that the chatbot use solely retrieved documentation.

Why this answer

Grounding with Vertex AI Search ensures the model's responses are strictly based on the retrieved documents from the internal documentation, preventing hallucination or reliance on pre-trained knowledge. Grounding works by providing the model with a search result context that it must use as the sole source for generating answers, effectively constraining the output to the provided documents.

Exam trap

The exam often tests the distinction between controlling model behavior (temperature, token limits) and controlling the source of information (grounding), leading candidates to mistakenly choose temperature or token adjustments as a solution for hallucination prevention.

How to eliminate wrong answers

Option A is wrong because fine-tuning the model on the documentation would embed the knowledge into the model's parameters, but it does not guarantee that the model will only use that knowledge during inference; the model could still generate responses from its pre-trained weights or hallucinate. Option C is wrong because increasing max output tokens only controls the length of the response, not the source of the information; the model could still generate content not found in the retrieved documents. Option D is wrong because setting temperature to 0.0 makes the model deterministic (greedy decoding) but does not restrict the model to only use retrieved documents; it can still produce answers based on its internal knowledge.

572
MCQmedium

A retail company wants to use a generative AI model to create personalized product descriptions for thousands of items. They need the descriptions to be consistent in style and format, but they also want to avoid the model inventing false features. Which approach best balances consistency and accuracy?

A.Use a high temperature to encourage creativity and let the model generate unique descriptions for each item.
B.Increase the max output tokens to allow the model to elaborate more on each product.
C.Fine-tune the model on a dataset of existing product descriptions to learn the style.
D.Set a low temperature and provide a structured prompt with clear guidelines and examples.
AnswerD

A low temperature reduces randomness, making outputs more deterministic and consistent. A structured prompt with guidelines and examples further constrains the model to follow the desired style and format. This combination minimizes the risk of hallucinated features while maintaining uniformity across thousands of descriptions, directly addressing both requirements.

Why this answer

Setting a low temperature makes the model's output more predictable and consistent, while a structured prompt with guidelines and examples steers the model to follow a specific style and format. This combination reduces the likelihood of hallucinated features because the model is constrained by explicit instructions and examples. Other approaches either increase variability or do not directly address accuracy.

Exam trap

The trap here is thinking that fine-tuning alone ensures accuracy, when in fact prompt design and temperature settings are more direct controls for consistency and hallucination reduction.

573
MCQhard

A team is building a medical diagnosis assistant using a foundation model. To comply with regulations, they need to ensure the model does not make up facts. What is the best approach?

A.Use a small model to hallucinate less
B.Use grounding with Vertex AI Search
C.Reduce temperature to 0
D.Fine-tune on medical journals
AnswerB

Grounding with Vertex AI Search anchors responses to retrieved, authoritative sources, so the model cites verified content instead of fabricating facts. This directly satisfies the regulatory requirement that the medical assistant must not make up information.

Why this answer

Grounding with Vertex AI Search is the best approach because it connects the foundation model's outputs to a verifiable, curated knowledge base, ensuring factual accuracy and compliance with regulations that prohibit hallucination. By retrieving information from a trusted source (e.g., medical databases) in real time, the model can cite evidence and avoid generating unverified claims.

Exam trap

Google Cloud often tests the misconception that reducing temperature or using a smaller model can eliminate hallucination, when in fact only grounding with external, verifiable data sources can reliably prevent fact fabrication in high-stakes domains.

How to eliminate wrong answers

Option A is wrong because using a smaller model does not inherently reduce hallucination; smaller models have less capacity and may actually hallucinate more due to limited training data and weaker reasoning. Option C is wrong because reducing temperature to 0 makes the model deterministic but does not prevent it from generating plausible-sounding but false information; it still relies on its parametric knowledge, which can be incomplete or outdated. Option D is wrong because fine-tuning on medical journals alone does not guarantee factual accuracy; the model may memorize and reproduce errors, and it cannot dynamically verify facts against a live, authoritative source.

574
MCQhard

A company wants to generate a video from a text description using Google Cloud. Which service is designed for this?

A.Codey
B.Chirp
C.Imagen
D.Veo
AnswerD

Veo is Google Cloud's generative video model, converting text prompts directly into video output. It satisfies the stem's requirement for text-to-video generation, unlike Vertex AI's text or image models. Veo handles prompt understanding, temporal consistency and rendering natively, making it the purpose-built service for this scenario.

Why this answer

Veo is Google Cloud's generative AI model specifically designed for creating high-quality videos from text or image prompts. It leverages advanced diffusion and transformer architectures to generate coherent video sequences, making it the correct choice for text-to-video generation.

Exam trap

The trap here is that candidates often confuse Imagen (text-to-image) with Veo (text-to-video), assuming any generative visual model can handle video, but Google Cloud explicitly separates these capabilities into distinct services.

How to eliminate wrong answers

Option A is wrong because Codey is Google's model for code generation and chat, not video creation. Option B is wrong because Chirp is a speech-to-text and text-to-speech model, focused on audio processing. Option C is wrong because Imagen is a text-to-image model, capable of generating static images but not video sequences.

575
MCQhard

A media company uses generative AI to produce personalized news summaries. They notice that summaries occasionally contain factual errors and biased language. What business strategy should they implement to address these issues while maintaining user engagement?

A.Disable personalization and serve generic summaries to all users.
B.Allow users to flag errors and manually correct summaries in real-time.
C.Implement a human review layer for high-risk topics and use automated fact-checking for all content, with a feedback loop for model improvement.
D.Replace AI with entirely human-written summaries.
AnswerC

A human review layer catches high-risk factual and bias errors that automation misses, while automated fact-checking scales across all content. The feedback loop retrains the model, reducing recurrence without suppressing the personalisation that sustains user engagement.

Why this answer

It balances accuracy and engagement by combining automated fact-checking with human review for high-risk topics. This hybrid approach reduces factual errors and biased language while maintaining the personalization that drives user engagement. The feedback loop continuously improves the model, addressing root causes rather than just symptoms.

Exam trap

Google Cloud often tests the misconception that either full automation or full human oversight is the only solution, when the correct answer is a hybrid approach that leverages the strengths of both AI and human judgment.

How to eliminate wrong answers

Option A is wrong because disabling personalization eliminates the core value proposition of generative AI for news summaries, likely reducing user engagement significantly without addressing the underlying model flaws. Option B is wrong because allowing real-time manual corrections by users is impractical at scale, introduces latency, and does not prevent errors from reaching users in the first place; it also lacks a systematic feedback mechanism for model improvement. Option D is wrong because replacing AI with entirely human-written summaries is cost-prohibitive, slow, and defeats the purpose of using generative AI for scalability and personalization.

576
MCQhard

A media company plans to launch a generative AI feature that creates personalized article summaries for subscribers. Before launch, the product team must choose an operating model for ongoing quality, cost, and safety oversight. Which approach best supports responsible scaling of the feature?

A.Assign a cross-functional team to own evaluation, monitoring, and incident response for the feature after launch.
B.Rely on the model provider's built-in safety filters as the complete control set for the feature.
C.Freeze the model version and prompt at launch so behavior remains stable and no further review is needed.
D.Delegate all quality decisions to the engineering team that built the integration and revisit them annually.
AnswerA

Generative AI features require continuing oversight because model behavior, content, and costs drift over time. A cross-functional team combining product, engineering, and editorial or legal perspectives can run evaluations, watch quality and safety signals, and respond to incidents. This creates clear accountability for the feature's operation rather than treating launch as the finish line.

Why this answer

Responsible scaling depends on continuing ownership after launch. A cross-functional team can evaluate summary quality against source articles, monitor safety and cost signals, and respond quickly when issues arise. Freezing the system, depending solely on provider filters, or leaving decisions to one function with annual reviews all leave gaps in quality, safety, and accountability for a live subscriber-facing feature.

Exam trap

The trap here is treating launch as the end of governance, when generative AI features need ongoing evaluation and incident response ownership.

577
MCQeasy

A marketing team wants to generate product descriptions using a text generation model on Vertex AI. They need consistent output style across all descriptions, including tone and length. They have a small set of 10 high-quality example descriptions that capture the desired style. The team has limited ML expertise and wants a quick solution that does not require model retraining. Which approach should they use?

A.Use a pre-built template with no model input.
B.Fine-tune the model on a large external dataset of product descriptions.
C.Use few-shot prompting with the examples in the prompt.
D.Set the temperature to 0.9 to maximize creativity.
AnswerC

Few-shot prompting embeds the ten exemplar descriptions directly in the prompt, steering tone and length without weight updates. This satisfies the no-retraining constraint and suits limited ML expertise, since no training pipeline or labelled dataset is needed. It reliably enforces the consistent style the marketing team requires.

Why this answer

Few-shot prompting is the correct approach because it allows the team to inject the desired style, tone, and length directly into the prompt using the 10 high-quality examples, without any model retraining. This technique leverages the in-context learning capability of large language models on Vertex AI, enabling consistent output from a small set of demonstrations. It is ideal for teams with limited ML expertise as it requires only prompt engineering, not fine-tuning or infrastructure changes.

Exam trap

Google Cloud often tests the misconception that higher temperature always improves output quality, but the trap here is that temperature controls randomness, not consistency, so candidates may incorrectly choose Option D without understanding that low temperature is required for reproducible style and length.

How to eliminate wrong answers

Option A is wrong because a pre-built template with no model input cannot generate dynamic, context-aware product descriptions; it produces static text that lacks the flexibility and nuance of a generative model. Option B is wrong because fine-tuning on a large external dataset would require significant ML expertise, data preparation, and compute resources, contradicting the requirement for a quick solution without model retraining. Option D is wrong because setting temperature to 0.9 maximizes randomness and creativity, which is the opposite of what is needed for consistent output style; a lower temperature (e.g., 0.2) would be more appropriate for deterministic, reproducible results.

578
MCQmedium

A marketing team wants to generate personalized email subject lines that match their brand voice. They need a Google Cloud service that allows them to fine-tune a generative model on their own email data. Which service should they use?

A.Cloud Natural Language API
B.Vertex AI
C.Dialogflow CX
D.AutoML Natural Language
AnswerB

Vertex AI provides a comprehensive platform for training and fine-tuning generative models, including Gemini, on custom datasets. The marketing team can use their email data to fine-tune a model to adopt their brand voice. Vertex AI offers tools for data preparation, training, and deployment, making it the right choice for this customization need. It supports supervised fine-tuning and reinforcement learning from human feedback.

Why this answer

Vertex AI is the Google Cloud platform that supports fine-tuning generative models like Gemini on custom datasets. It allows the marketing team to train a model on their email data to capture their brand voice and generate personalized subject lines. Other services either lack generative capabilities or are not intended for fine-tuning on custom text data.

Exam trap

The trap here is confusing AutoML Natural Language, which is for classification, with Vertex AI's generative fine-tuning capabilities.

579
MCQeasy

A data scientist wants to generate realistic product images for an online catalog using Google Cloud's generative AI. Which service should they use?

A.Imagen on Vertex AI
B.Codey API for code generation
C.Gemini API with text-to-text prompts
D.Vertex AI Model Garden without a specific model
AnswerA

Imagen on Vertex AI generates photorealistic images from text prompts, directly satisfying the requirement for realistic product imagery. Unlike text-only models such as Gemini, Imagen is purpose-built for image synthesis, offering controls over aspect ratio, resolution and style that suit catalogue photography.

Why this answer

Imagen on Vertex AI is Google Cloud's specialized service for generating high-quality, photorealistic images from text prompts. It is built on diffusion models and is directly designed for image generation tasks, making it the correct choice for creating product images for an online catalog.

Exam trap

The trap here is that candidates may confuse the general-purpose Gemini API (which can handle multimodal inputs) with a dedicated image generation service, overlooking that Gemini's text-to-text mode does not generate images, while Imagen is purpose-built for that task.

How to eliminate wrong answers

Option B is wrong because Codey API is designed for code generation, not image generation; it uses models specialized in programming languages and cannot produce visual outputs. Option C is wrong because Gemini API with text-to-text prompts is optimized for text-based tasks like summarization or question answering, not for generating images; while Gemini can process images, its primary text-to-text mode does not generate visual content. Option D is wrong because Vertex AI Model Garden is a repository of pre-trained models and frameworks, but without selecting a specific model like Imagen, it cannot directly generate images; it requires explicit model selection and configuration.

580
MCQhard

A healthcare company is using a generative AI model to draft patient education materials. The model sometimes generates content that includes specific medical advice, which could be harmful if inaccurate. The company wants to ensure that the model's outputs are safe and do not provide medical recommendations. Which technique should they implement?

A.Apply a safety filter that blocks any output containing medical terminology.
B.Reduce the model's temperature to make outputs more deterministic and less likely to include advice.
C.Implement prompt engineering with explicit instructions to avoid providing medical advice.
D.Use reinforcement learning from human feedback (RLHF) to fine-tune the model to avoid giving advice.
AnswerC

Prompt engineering with clear instructions, such as 'Do not provide medical advice; only provide general information,' can effectively steer the model away from generating harmful recommendations. This is a low-cost, immediate solution that can be refined iteratively. It leverages the model's ability to follow instructions when they are explicit and well-crafted.

Why this answer

Prompt engineering with explicit instructions is a direct and efficient way to constrain the model's output. By clearly stating that the model should not provide medical advice, the model is guided to generate only general educational content. This approach is flexible and can be updated as needed without retraining, making it suitable for the healthcare company's requirement to ensure safety.

Exam trap

The trap here is thinking that technical parameter adjustments like temperature will solve content safety issues, when in fact clear instructions in the prompt are often more effective for controlling what the model says.

581
Multi-Selectmedium

A marketing team is using a generative AI model to create ad copy. They want to ensure the outputs align with their brand voice and avoid offensive language. Which two practices should they adopt? (Choose two.)

Select 2 answers
A.Set the model's temperature to the maximum value to encourage creativity.
B.Deploy multiple model versions and manually select the best output for each ad.
C.Provide clear instructions and examples in the prompt to guide the model's tone and style.
D.Fine-tune the model on a large dataset of competitor ads.
E.Use Vertex AI Safety Filters to block harmful content.
AnswersC, E

Prompt engineering with explicit instructions and examples helps steer the model toward the desired brand voice and away from offensive language. By specifying tone, style, and constraints, the team can influence the output effectively. This is a direct and flexible method to align generated content with brand guidelines without retraining the model.

Why this answer

To align outputs with brand voice and avoid offensive language, the team should use prompt engineering to guide tone and style, and implement Vertex AI Safety Filters to block harmful content. These two practices work together to ensure both stylistic and safety requirements are met. Other options either introduce risks or are inefficient.

Exam trap

The trap here is assuming that fine-tuning on competitor data or maximizing creativity will achieve brand alignment, when they actually increase risk and inconsistency.

582
Multi-Selectmedium

Which TWO techniques are effective for reducing bias in generative AI model outputs?

Select 2 answers
A.Increasing model size to learn more patterns
B.Training on diverse and representative datasets
C.Relying solely on post-hoc filters
D.Using adversarial debiasing methods during fine-tuning
E.Limiting the model to only factual prompts
AnswersB, D

Correct: Diverse data helps reduce biased associations.

Why this answer

Training on diverse and representative datasets directly reduces sampling bias and coverage gaps in the training distribution, which are primary sources of stereotypical or skewed outputs. By ensuring the model sees balanced examples across demographics, contexts, and edge cases, it learns more equitable representations and reduces the likelihood of generating biased content.

Exam trap

Google Cloud often tests the misconception that increasing model size or adding post-hoc filters is sufficient to mitigate bias, when in reality these approaches fail to address the root causes of bias in training data and model representations.

583
MCQeasy

A university research group wants to experiment with Google's Gemini models through a simple web interface, without writing any code or provisioning cloud infrastructure. They need to upload PDFs, ask questions, and iterate on prompts interactively. Which Google Cloud offering should they use?

A.Vertex AI Studio
B.Vertex AI Model Registry
C.Vertex AI Feature Store
D.Vertex AI Pipelines
AnswerA

Vertex AI Studio provides a console-based playground where users can prompt Gemini models, upload files such as PDFs, tune parameters, and compare responses without writing code. It is designed exactly for interactive experimentation and prompt iteration, so it meets the research group's requirement without any infrastructure setup.

Why this answer

Vertex AI Studio is the console-based environment for prompting Gemini models, uploading files such as PDFs, adjusting parameters like temperature and token limits, and saving prompts. It requires no coding or infrastructure provisioning, which matches the research group's interactive experimentation needs. Pipelines, Feature Store, and Model Registry serve orchestration, feature management, and model cataloging respectively, none of which provide a prompt playground.

Exam trap

The trap here is confusing Vertex AI Studio's interactive prompt playground with infrastructure services like Pipelines or Model Registry that manage workflows and artifacts rather than model prompting.

584
MCQhard

A financial institution deploys a chatbot using Gemini Pro in Vertex AI. Compliance requires logging all user inputs and model outputs for audit. Which approach meets this requirement?

A.Capture logs via Cloud Monitoring
B.Enable Vertex AI Endpoint request-response logging
C.Use Cloud Logging sink with a filter for Vertex AI requests
D.Enable Vertex AI Model Registry logging
AnswerB

Endpoint request-response logging captures the full prompt and completion payloads for every prediction, storing them in Cloud Logging for audit retrieval. This directly satisfies the compliance constraint to log all user inputs and model outputs without altering application code.

Why this answer

Vertex AI Endpoint request-response logging captures both the user's input prompt and the model's generated output, which is precisely what compliance auditing requires. This feature logs the exact payloads sent to and received from the deployed model, ensuring a complete audit trail without additional configuration.

Exam trap

The trap here is that candidates confuse Cloud Logging sinks or Cloud Monitoring with the specific Vertex AI feature that must be explicitly enabled on the endpoint, assuming that default logging captures request-response payloads when it does not.

How to eliminate wrong answers

Option A is wrong because Cloud Monitoring is designed for metrics, alerts, and dashboards, not for capturing detailed request-response payloads for audit compliance. Option C is wrong because a Cloud Logging sink with a filter can only export logs that already exist; it does not enable the capture of Vertex AI request-response logs, which must be explicitly enabled on the endpoint. Option D is wrong because Vertex AI Model Registry logging tracks model version metadata and lifecycle events, not the user inputs and model outputs from inference calls.

585
Multi-Selecthard

A financial analyst uses generative AI to summarize earnings reports. The summaries vary in style. Which THREE methods can improve consistency? (Choose three.)

Select 3 answers
A.Set temperature to 0.2
B.Increase max output tokens
C.Enable citation mode
D.Use few-shot prompting with fixed examples
E.Fine-tune on a curated dataset of desired summaries
AnswersA, D, E

Reduces output randomness.

Why this answer

Setting temperature to 0.2 reduces randomness in token sampling, making the model more deterministic and less likely to produce stylistic variations. Lower temperatures (e.g., 0.1–0.3) narrow the probability distribution, forcing the model to select the most likely next token, which directly improves consistency across multiple summaries.

Exam trap

A common misconception in this Google exam is that increasing max output tokens or enabling citation mode improves consistency. In reality, these features control length and attribution respectively, not stylistic uniformity. The correct methods focus on reducing randomness (low temperature), providing consistent examples (few-shot), or fine-tuning.

586
Multi-Selecteasy

A developer is using the Vertex AI PaLM API to generate code. They want to ensure the output is safe and adheres to company policies. Which THREE attributes can they configure in the safety_settings parameter?

Select 3 answers
A.Language detection
B.Sentiment analysis
C.Toxicity
D.Harassment
E.Sexually explicit content
AnswersC, D, E

Toxicity is a configurable safety attribute within the Vertex AI PaLM API's safety_settings, letting the developer set thresholds that block harmful or offensive generated code. This directly satisfies the stem's requirement to keep output safe and aligned with company policies, alongside other harm categories.

Why this answer

The safety_settings parameter in the Vertex AI PaLM API accepts a list of safety categories, each with a configurable threshold, and the supported categories include Toxicity (C), Harassment (D), and Sexually explicit content (E), so these three are the attributes the developer can configure to filter unsafe output and enforce company policies. Toxicity (C) lets them block content that is rude, disrespectful, or otherwise harmful, Harassment (D) targets content that bullies or intimidates individuals or groups, and Sexually explicit content (E) filters sexually explicit material; each is a distinct safety category with its own threshold setting. Language detection (A) is not a safety category but a separate text-analysis capability, and sentiment analysis (B) is likewise an analytical feature rather than a configurable safety attribute, so neither belongs in safety_settings.

Exam trap

The trap here is that candidates may confuse general NLP features (like language detection or sentiment analysis) with the specific safety filtering attributes available in the safety_settings parameter, leading them to select options that are not part of the API's harm category configuration.

587
MCQhard

A data science team wants to build a custom model for generating product descriptions that adhere to specific brand guidelines. They have 5,000 high-quality examples. Which approach balances cost and accuracy?

A.Fine-tune a foundation model (e.g., PaLM 2) using Vertex AI Model Garden
B.Use a pre-built API with prompt engineering and few-shot examples
C.Use Vertex AI Agent Builder with a custom prompt
D.Train a model from scratch using TensorFlow on Vertex AI
AnswerA

Fine-tuning adapts a foundation model's weights to the brand's tone using the 5,000 labelled examples, achieving higher fidelity than prompting alone at far lower cost than training from scratch. Vertex AI Model Garden provides the managed pipeline for this.

Why this answer

Fine-tuning a foundation model on the examples yields high accuracy with moderate cost. Training from scratch is overkill; prompt engineering may not capture all nuances.

588
MCQmedium

A user reports that the model's response to the same prompt varies significantly across different calls. Which parameter change would most likely reduce variability?

A.Decrease topK to 10.
B.Decrease temperature to 0.2.
C.Increase candidateCount to 3.
D.Increase maxOutputTokens to 2000.
AnswerB

Lowering temperature to 0.2 sharpens the softmax probability distribution, so high-probability tokens dominate sampling and near-deterministic output replaces the wide variance seen at higher values. This directly satisfies the stem's requirement to reduce run-to-run variability for an identical prompt, since temperature is the parameter governing sampling randomness.

Why this answer

Temperature controls the randomness of token sampling. Lowering temperature (e.g., to 0.2) makes the model's output more deterministic by reducing the probability of low-likelihood tokens, thus decreasing variability across calls for the same prompt.

Exam trap

Candidates often mistake topK or candidateCount for the primary control of output variability, but in Google's Vertex AI and Gen AI models, temperature is the direct parameter that governs randomness in token selection.

How to eliminate wrong answers

Option A is wrong because decreasing topK to 10 still allows sampling from a limited set of tokens, which can introduce variability if temperature is not also reduced; topK alone does not control randomness as directly as temperature. Option C is wrong because increasing candidateCount to 3 generates multiple independent responses, which increases variability rather than reducing it. Option D is wrong because increasing maxOutputTokens to 2000 only extends the maximum length of the response, not the consistency of the output; it has no effect on token selection randomness.

589
MCQhard

A multimodal generative AI system processes both image and text inputs to produce captions. During inference, the image encoder sometimes produces noisy or missing features. Which architectural design decision best handles such input degradation without retraining?

A.Train a separate variational autoencoder to produce a clean latent representation from the noisy image.
B.Increase the image encoder’s capacity to better extract robust features.
C.Apply standard image preprocessing (e.g., denoising) to all inputs before feeding to the encoder.
D.Introduce a gating mechanism that learns to weigh image features based on confidence scores from the encoder.
AnswerD

A learned gating mechanism dynamically weights image features by encoder confidence, letting the decoder down-weight noisy or missing inputs at inference. This handles degradation without retraining, satisfying the stem's constraint of robustness to unreliable visual features.

Why this answer

A gating mechanism dynamically adjusts the contribution of image features based on confidence scores from the encoder, allowing the model to gracefully handle noisy or missing features without retraining. This architectural design learns to suppress unreliable image inputs and rely more on text or other modalities, ensuring robust caption generation under input degradation.

Exam trap

Google Cloud often tests the misconception that preprocessing or model capacity adjustments are the only ways to handle input noise, but the key insight is that architectural mechanisms like gating can adaptively handle degradation at inference time without retraining.

How to eliminate wrong answers

Option A is wrong because training a separate variational autoencoder (VAE) to produce clean latent representations requires additional training data and retraining, which contradicts the 'without retraining' constraint; it also adds complexity without addressing dynamic degradation during inference. Option B is wrong because increasing the image encoder’s capacity does not inherently handle noisy or missing features—it may overfit to training data and still produce unreliable outputs when inputs degrade, and it requires retraining to change capacity. Option C is wrong because standard image preprocessing like denoising is a fixed, non-adaptive approach that cannot compensate for missing features or varying noise levels, and it may discard useful information; it also does not leverage the model’s ability to learn confidence-based weighting.

590
MCQmedium

A company is developing a generative AI application that will be used by customers in multiple countries, including those with strict data residency laws. How should they approach data governance?

A.Store all data in a single central data center to simplify management
B.Use a VPN to route data through compliant regions
C.Use data residency controls to keep data in specified regions
D.Anonymize all data before processing to avoid residency issues
AnswerC

Data residency controls pin storage and processing to approved geographic regions, directly satisfying the strict residency laws named in the stem. This keeps training and inference data within mandated jurisdictions, unlike encryption or consent mechanisms, which address confidentiality and lawful basis rather than the physical location of data.

Why this answer

Data residency controls, such as those provided by Google Cloud Organization Policies, Assured Workloads, or Cloud Storage location constraints, allow the company to enforce that data is stored and processed only within specified geographic regions. This directly addresses strict data residency laws by preventing data from leaving the jurisdiction, which is a fundamental requirement for compliance with regulations like GDPR or Brazil's LGPD. Unlike workarounds, this approach provides native, auditable enforcement at the infrastructure level.

Exam trap

A common misconception is that technical workarounds like VPNs or anonymization can substitute for native data residency enforcement, when in fact only infrastructure-level controls provide the auditable, deterministic compliance required by law.

How to eliminate wrong answers

Option A is wrong because storing all data in a single central data center violates data residency laws that require data to remain within specific national or regional boundaries, and it does not provide any mechanism to segregate or control data flow based on user location. Option B is wrong because using a VPN to route data through compliant regions does not change the physical storage location of the data; it only masks the network path, and the data still resides in a non-compliant data center, which fails legal audits. Option D is wrong because anonymization is not a guaranteed solution for data residency; many regulations (e.g., GDPR) still apply to pseudonymized or anonymized data if re-identification is possible, and the data's physical location remains non-compliant unless stored in the required region.

591
MCQhard

After fine-tuning a model on customer support data, the model starts using profanity. What is the most effective mitigation?

A.Add profanity to training data as negative examples
B.Reduce learning rate and retrain
C.Increase temperature to reduce confidence
D.Enable a safety attribute filter
AnswerD

A safety attribute filter intercepts model outputs and blocks or redacts harmful content such as profanity before it reaches users, directly addressing the fine-tuning side effect. Retraining or prompt engineering may reduce but not reliably prevent toxic generations.

Why this answer

Enabling a safety attribute filter is the most effective mitigation because it acts as a post-processing guardrail that blocks profanity at inference time, regardless of the model's training data. This is a standard practice in production LLM deployments, where safety filters (e.g., using keyword matching or classifier models) intercept and redact harmful outputs before they reach the user, providing immediate and reliable control without requiring retraining.

Exam trap

Google often tests the misconception that modifying training parameters (like learning rate or temperature) can fix output quality issues, when in fact post-processing filters are the standard, immediate solution for content safety in production LLM systems.

How to eliminate wrong answers

Option A is wrong because adding profanity as negative examples in training data can inadvertently reinforce the behavior or cause the model to learn spurious correlations, and it does not guarantee removal of already learned profanity patterns. Option B is wrong because reducing the learning rate and retraining only adjusts the model's weights during fine-tuning, which does not address the root cause of profanity generation and may not eliminate learned toxic patterns without extensive data curation. Option C is wrong because increasing temperature increases randomness in token sampling, which can actually increase the likelihood of generating profanity by making the model less deterministic, not reduce it.

592
MCQmedium

A marketing team uses a Gemini model to generate ad copy. They notice the outputs are repetitive and lack variety across multiple runs for the same prompt. They want more diverse creative options without sacrificing relevance. Which parameter adjustment should they make?

A.Set top-k to 1 so the model always picks the single most likely token.
B.Decrease the temperature to reduce repetition.
C.Reduce the max output tokens to force the model to vary its wording.
D.Increase the temperature to allow more varied token sampling.
AnswerD

Higher temperature flattens the probability distribution, allowing less likely but still relevant tokens to be selected. This produces more diverse creative outputs across runs. The team should balance temperature to avoid incoherence while gaining variety. It directly addresses the repetition issue without changing the prompt or model.

Why this answer

Increasing temperature broadens token sampling, yielding more varied creative outputs while still following the prompt. Lower temperature, top-k of 1, and reduced max output tokens all make outputs more deterministic or shorter, not more diverse. Temperature is the primary control for balancing creativity and coherence in ad copy generation.

Exam trap

The trap here is confusing length controls or greedy decoding with creativity controls, when temperature is the parameter that governs output diversity.

593
MCQeasy

A product manager wants to quickly build a conversational agent that can answer FAQs from the company's help center articles. They have limited coding experience. Which Google Cloud service is BEST suited for this task?

A.Vertex AI Pipelines
B.Vertex AI Agent Builder
C.Vertex AI Studio
D.Model Garden
AnswerB

Vertex AI Agent Builder provides a low-code environment for assembling conversational agents grounded in enterprise documents, so help centre articles can be indexed and queried without bespoke development. This satisfies the stem's constraint of limited coding experience combined with rapid FAQ agent delivery.

Why this answer

Vertex AI Agent Builder is the best choice because it provides a no-code/low-code interface specifically designed for building conversational agents and search experiences. It allows the product manager to connect help center articles as a data source and automatically generate a FAQ-answering agent without writing code, making it ideal for someone with limited coding experience.

Exam trap

Google often tests the distinction between tools for building agents (Agent Builder) versus tools for model experimentation (Studio) or model selection (Model Garden), and candidates mistakenly choose Studio or Model Garden because they think any generative AI tool can build a chatbot, ignoring the specific no-code agent-building capability required.

How to eliminate wrong answers

Option A is wrong because Vertex AI Pipelines is a tool for orchestrating and automating ML workflows (e.g., training, deployment) and requires pipeline definition via code or SDK, not for quickly building a conversational agent. Option C is wrong because Vertex AI Studio is a platform for experimenting with and tuning generative AI models (like prompts and foundation models), but it does not provide a built-in agent builder for connecting to help center articles and generating FAQ responses without custom development. Option D is wrong because Model Garden is a repository of pre-trained foundation models and does not include the agent-building or data-connecting capabilities needed to create a conversational FAQ agent.

594
Multi-Selecteasy

A developer is using the Gemini API to generate text summaries. They want to control the creativity and diversity of the output. Which THREE parameters can they adjust?

Select 3 answers
A.Context window
B.Top-p
C.Embedding dimension
D.Top-k
E.Temperature
AnswersB, D, E

Top-p (nucleus sampling) restricts token selection to the smallest cumulative-probability set, directly controlling output diversity. Adjusting it satisfies the requirement to tune creativity, since lower values yield focused summaries and higher values broaden word choice.

Why this answer

Temperature (E) is correct because it scales the logits before sampling, so lower values make the output more deterministic and higher values increase randomness and creativity. Top-p (B), or nucleus sampling, is correct because it restricts sampling to the smallest set of tokens whose cumulative probability exceeds p, directly controlling diversity. Top-k (D) is correct because it limits sampling to the k most likely tokens, which also tunes creativity and variety.

Context window (A) only defines the maximum number of tokens the model can consider, and embedding dimension (C) is a fixed property of the embedding model, so neither controls generation creativity.

Exam trap

Often, the distinction between parameters that affect input processing (context window, embedding dimension) versus those that control output generation (temperature, top-k, top-p) is tested, leading candidates to mistakenly select context window as a creativity parameter.

595
MCQhard

A research lab is fine-tuning a large language model on a small dataset of medical records. They observe that the model overfits, memorizing specific patient details and producing outputs that violate privacy regulations. Which technique should they apply to improve generalization and reduce memorization?

A.Increase the batch size to 64
B.Increase the number of training epochs
C.Use early stopping based on validation loss
D.Apply differential privacy (DP-SGD) during fine-tuning
AnswerD

DP-SGD injects calibrated noise into per-sample gradients during fine-tuning, bounding any single record's influence on learned parameters. This directly curbs the memorisation of individual patient details that breaches privacy rules, while the noise also regularises training so the model generalises better to unseen records.

Why this answer

Differential privacy (DP-SGD) is the correct technique because it directly addresses memorization of sensitive patient data by adding calibrated noise to the gradient updates during fine-tuning. This bounds the model's ability to encode any single individual's information, improving generalization and ensuring compliance with privacy regulations like HIPAA.

Exam trap

Google Cloud often tests the misconception that early stopping or batch size adjustments can prevent memorization, when in fact only techniques like differential privacy directly bound the influence of individual training examples.

How to eliminate wrong answers

Option A is wrong because increasing batch size to 64 reduces gradient variance but does not prevent memorization of specific patient details; it may even accelerate overfitting on a small dataset. Option B is wrong because increasing the number of training epochs exacerbates overfitting, causing the model to memorize more training examples and worsen privacy violations. Option C is wrong because early stopping based on validation loss only halts training when validation performance degrades, but it does not impose any privacy guarantee or fundamentally limit memorization of unique patient records.

596
MCQmedium

A team wants to improve the factual accuracy of their chatbot responses regarding internal company policies. What is the most effective approach?

A.Use few-shot prompting with example Q&A pairs
B.Increase the model's maximum tokens
C.Fine-tune the model on policy documents
D.Use RAG with Vertex AI Search indexing the policies
AnswerD

RAG retrieves relevant passages from the indexed policy corpus at query time and supplies them as grounding context, so answers cite actual internal documents rather than relying on parametric memory. Vertex AI Search handles indexing and retrieval, satisfying the requirement for factual accuracy on company-specific policies.

Why this answer

RAG with Vertex AI Search is the most effective approach because it retrieves relevant, up-to-date policy documents from a curated index and injects them into the prompt context at inference time, grounding the chatbot's responses in authoritative sources without modifying the underlying model. This ensures factual accuracy for dynamic or evolving policies, as the model can reference the exact text rather than relying on static training data.

Exam trap

A common misconception in the Google Gen AI Leader exam is that fine-tuning (Option C) is the best approach to improve factual accuracy for dynamic knowledge. In reality, RAG with Vertex AI Search is superior because it retrieves up-to-date policies from a curated index without retraining the model, and provides verifiable source citations.

How to eliminate wrong answers

Option A is wrong because few-shot prompting provides example Q&A pairs but does not guarantee the model will recall or cite the correct policy details, especially for nuanced or updated policies; it relies on the model's parametric memory, which can be incomplete or outdated. Option B is wrong because increasing the maximum tokens only expands the output length, not the factual grounding; it does nothing to improve the accuracy of the content generated. Option C is wrong because fine-tuning on policy documents embeds static knowledge into the model weights, making it difficult to update when policies change and risking catastrophic forgetting of other capabilities; it also does not provide a mechanism to cite specific sources or handle real-time retrieval.

597
MCQeasy

A company notices that their AI chatbot occasionally generates incorrect information. Which technique can best reduce hallucinations without retraining?

A.Use a longer system prompt without examples
B.Use system instructions to constrain the model to only answer from provided context
C.Set top_p to 0.1
D.Increase temperature to 0.9
AnswerB

System instructions steer generation by restricting the model to answer solely from supplied context, so unsupported claims are refused rather than invented. This constrains decoding behaviour at inference time, reducing hallucinations without any retraining, fine-tuning or modification of model weights.

Why this answer

Constraining the model to answer only from provided context directly addresses the root cause of hallucinations—the model generating information not grounded in verified sources. This technique, often implemented via system instructions or retrieval-augmented generation (RAG) pipelines, forces the model to rely on a trusted knowledge base rather than its parametric memory, effectively eliminating unsupported fabrications without requiring retraining.

Exam trap

Google often tests the misconception that adjusting sampling parameters (like top_p or temperature) can fix hallucinations, when in reality these parameters control randomness, not factual grounding, and the correct solution is to constrain the model's output to a trusted context.

How to eliminate wrong answers

Option A is wrong because using a longer system prompt without examples does not prevent hallucinations; it may actually increase the risk by introducing more ambiguous or conflicting instructions, and without explicit grounding constraints, the model can still generate unverified content. Option C is wrong because setting top_p to 0.1 reduces the diversity of token sampling but does not enforce factual accuracy—it merely makes outputs more deterministic, which can still produce confident hallucinations if the model's internal knowledge is flawed. Option D is wrong because increasing temperature to 0.9 increases randomness and creativity in outputs, which exacerbates hallucination risk by making the model more likely to generate improbable or fabricated information.

598
Multi-Selecthard

Which THREE benefits does Vertex AI Agent Builder provide over building a custom conversational agent from scratch?

Select 3 answers
A.Automatic scaling and load balancing
B.Pre-built integration for grounding on enterprise data sources
C.Full control over the underlying ML model architecture
D.Built-in safety filters and guardrails
E.Guaranteed lower inference latency
AnswersA, B, D

Vertex AI Agent Builder manages the underlying serving infrastructure, automatically scaling instances and distributing traffic as demand fluctuates. This removes the capacity planning and load-balancing work a custom-built agent would otherwise require the team to implement and maintain.

Why this answer

Option A is correct because Vertex AI Agent Builder is a managed service that automatically handles scaling and load balancing of the agent infrastructure, removing the need to provision or tune servers yourself. Option B is correct because it provides pre-built connectors and integrations for grounding responses on enterprise data sources such as Vertex AI Search, BigQuery, and other Google Cloud data stores, which would otherwise require custom retrieval pipelines. Option D is correct because it includes built-in safety filters and guardrails (e.g., responsible AI controls, content moderation, and policy enforcement) that a from-scratch agent would need to implement manually.

Option C is not correct because Agent Builder abstracts the underlying model and does not give full control over the ML model architecture, which is actually a limitation rather than a benefit. Option E is not correct because Agent Builder does not guarantee lower inference latency; latency depends on the selected model, region, and workload, and no such guarantee is offered.

Exam trap

The trap here is that candidates may confuse 'full control' (Option C) with the flexibility of Vertex AI Agent Builder, which actually limits architectural control in favor of managed simplicity, and may assume managed services always provide lower latency (Option E) without considering that custom optimizations can outperform generic managed solutions.

599
MCQhard

A global software company uses a generative AI model to produce localized release notes. Outputs are accurate but inconsistently formatted: some use tables, some bullet lists, and some paragraphs, which breaks the publishing pipeline. The team wants stable, machine-parseable formatting across all runs. Which technique should they prioritize?

A.Set the model's temperature to a moderate value and rely on repeated runs to average out formatting differences.
B.Instruct the model to think step by step about the release content before formatting the output.
C.Fine-tune the model on a dataset of previously published release notes to teach consistent formatting.
D.Define a strict output schema, such as a JSON or Markdown template, and require the model to conform to it.
AnswerD

A strict output schema removes formatting ambiguity by defining exactly which structure every response must follow. The model fills the template rather than choosing a format, so the publishing pipeline receives consistent, parseable output. This is prompt-level and enforceable with validation, requiring no retraining or sampling changes.

Why this answer

Inconsistent formatting is a specification gap. Declaring a strict output schema, such as a fixed JSON structure or Markdown template, forces every generation into the same shape and makes the output machine-parseable. Temperature averaging, reasoning prompts, and fine-tuning influence variability, reasoning, or style bias but do not guarantee structural conformity.

Exam trap

The trap here is thinking that fine-tuning or repeated sampling stabilizes formatting, when only an explicit output schema reliably constrains the structure of every response.

600
Multi-Selectmedium

A developer is using the Gemini API to generate marketing copy. They want the output to be diverse and creative but still relevant to the topic. Which THREE parameter adjustments would help achieve this? (Choose 3)

Select 3 answers
A.Increase temperature to 0.9
B.Increase top-k to 50
C.Decrease temperature to 0.1
D.Decrease top-k to 10
E.Increase top-p to 0.95
AnswersA, B, E

Raising temperature to 0.9 flattens the softmax probability distribution, so lower-probability tokens are sampled more often. This directly satisfies the stem's demand for diverse, creative marketing copy while the prompt's topic framing keeps output relevant.

Why this answer

Increasing temperature to 0.9 (option A) is correct because higher temperature flattens the probability distribution, making the model more likely to pick less probable tokens and thus produce more diverse, creative text while still grounded in the prompt. Increasing top-k to 50 (option B) is correct because a larger top-k widens the candidate pool from which the next token is sampled, allowing more varied word choices and greater output diversity. Increasing top-p to 0.95 (option E) is correct because a higher nucleus sampling threshold includes more tokens cumulatively comprising 95% of the probability mass, which increases randomness and creativity while excluding only the least likely tokens.

Decreasing temperature to 0.1 (option C) is not correct because low temperature makes the model nearly deterministic and conservative, reducing diversity. Decreasing top-k to 10 (option D) is not correct because a smaller top-k restricts sampling to fewer high-probability tokens, which narrows creativity rather than enhancing it.

Page 7

Page 8 of 14

Page 9