Courseiva

CCNA Business Strategies for Generative AI Solutions Questions

56 of 131 questions · Page 2/2 · Business Strategies for Generative AI Solutions · Answers revealed

76
MCQhard

A financial services firm wants to deploy a generative AI assistant that summarizes earnings call transcripts for its analysts. The firm's risk committee requires that the assistant never produce investment recommendations and that all outputs be traceable to source transcript passages. Which design choice most directly enforces both constraints?

A.Store all prompts and responses in BigQuery and run a nightly batch job to flag outputs that look like recommendations.
B.Deploy the assistant with a content filter that blocks any output containing financial terminology.
C.Use retrieval-augmented generation over the transcripts and apply a system instruction that prohibits recommendations, returning cited source chunks.
D.Fine-tune a Gemini model on historical earnings transcripts and deploy it with a low temperature setting.
AnswerC

RAG grounds every response in retrieved transcript passages and can return citations to those passages, satisfying traceability. A system instruction that explicitly forbids investment recommendations constrains the model's behavior at generation time. Together they directly address both the no-recommendation policy and the requirement that outputs map back to source text.

Why this answer

Grounding responses in retrieved transcript passages with RAG gives analysts verifiable citations, while a system instruction establishes a hard behavioral boundary against investment recommendations. This combination enforces both the provenance and the policy constraint at generation time, rather than relying on post-processing or broad content filters that cannot distinguish summary from advice.

Exam trap

The trap here is treating traceability and policy enforcement as model-tuning problems, when they are better solved through retrieval grounding and explicit generative instructions.

77
Multi-Selecteasy

Which TWO Google Cloud services can be used together to implement a RAG (retrieval-augmented generation) pipeline? (Select 2)

Select 2 answers
A.Cloud SQL
B.Vertex AI Vector Search
C.Bigtable
D.Vertex AI PaLM API
E.Cloud Functions
AnswersB, D

Provides vector similarity search for retrieval.

Why this answer

Vertex AI Vector Search (option B) is correct because it provides a managed vector database for storing and querying embeddings, which is essential for the retrieval step in a RAG pipeline. It enables semantic similarity search over large datasets, allowing the system to fetch relevant context documents based on a user query.

Exam trap

Google Cloud often tests the misconception that any database (like Cloud SQL or Bigtable) can serve as a vector store for RAG, but they lack native vector indexing and similarity search, making them unsuitable for efficient retrieval at scale.

78
Multi-Selectmedium

A company is considering whether to use Vertex AI's Generative AI Studio. Which TWO are benefits?

Select 2 answers
A.It is always cheaper than using third-party APIs
B.It integrates seamlessly with Vertex AI Pipelines for MLOps
C.It generates outputs that are always more accurate than custom models
D.It provides built-in tools for prompt engineering and iterative testing
E.It requires no coding or machine learning expertise to use
AnswersB, D

Integration allows automating deployment, monitoring, and retraining.

Why this answer

Vertex AI Generative AI Studio is designed to work natively with Vertex AI Pipelines, enabling users to incorporate generative models into end-to-end MLOps workflows for automation, monitoring, and retraining. This integration allows seamless orchestration of prompt tuning, model evaluation, and deployment within the same managed environment, reducing operational overhead.

Exam trap

Google Cloud often tests the misconception that 'no-code' tools eliminate the need for any ML expertise, but the trap here is that Generative AI Studio still requires understanding of prompt engineering, model evaluation, and cost trade-offs to avoid poor outputs or unexpected expenses.

79
MCQhard

A financial services firm wants to deploy a generative AI assistant that summarizes analyst research for internal advisors. The security team requires that the assistant never expose confidential client data in its responses, and the business wants to move quickly without building a custom model from scratch. Which approach best balances speed, control, and data protection on Google Cloud?

A.Fine-tune a foundation model on all historical analyst reports, including confidential client sections.
B.Deploy an open-source model on a single Compute Engine VM and rely on network firewalls for data protection.
C.Store all analyst reports in a public Cloud Storage bucket so the model can retrieve them without authentication.
D.Use a managed foundation model with grounding on an access-controlled document store and Sensitive Data Protection filters on inputs and outputs.
AnswerD

Grounding the assistant on an access-controlled corpus ensures it summarizes only documents the advisor is permitted to see, while Sensitive Data Protection inspection on prompts and responses adds a layer that can detect and redact confidential client identifiers. This uses managed models, so the firm avoids building from scratch, and it keeps sensitive data out of model weights, which supports both speed and the security team's non-exposure requirement.

Why this answer

Combining a managed foundation model with grounded retrieval over an access-controlled store and Sensitive Data Protection filtering keeps confidential client data out of model weights and out of responses. Advisors receive summaries only from documents they are authorized to view, and inspection adds detection and redaction for sensitive identifiers. This satisfies the security requirement while avoiding the cost and delay of training a custom model.

Exam trap

The trap here is assuming that network perimeter controls or fine-tuning alone protect confidential data, when the real risk is sensitive content appearing in model context or generated responses.

80
MCQeasy

A retail chain's leadership wants to understand how generative AI could improve its customer service operations but is unsure where to start. They ask a cloud consultant to identify candidate use cases, estimate potential value, and flag which ones are realistically achievable with current technology. Which activity should the consultant perform first?

A.Fine-tune a foundation model on the retailer's historical customer service transcripts and measure its response quality.
B.Run a generative AI use case discovery workshop with business stakeholders to inventory candidate scenarios and prioritize them by value and feasibility.
C.Deploy a general-purpose chatbot on the public website and monitor engagement metrics for one quarter.
D.Migrate the retailer's contact center platform to a new SaaS vendor that advertises built-in generative AI features.
AnswerB

Discovery workshops bring business and technical stakeholders together to enumerate candidate scenarios, then rank them by expected business value and technical feasibility. This directly answers leadership's request for candidate use cases, value estimates, and realism checks, and it produces a prioritized shortlist that later pilots can draw from, which is the standard first move in a generative AI business strategy engagement.

Why this answer

The consultant should begin with structured use case discovery, because the organization needs an inventory of candidate scenarios ranked by business value and technical feasibility before committing to any build or procurement. This produces the prioritized shortlist that later pilots, model choices, and investment decisions can be based on, aligning generative AI work with measurable business outcomes.

Exam trap

The trap here is equating progress with building something, so the first instinct becomes fine-tuning a model or deploying a chatbot instead of first identifying and prioritizing which use cases deserve investment.

81
MCQmedium

A regional insurance company wants to let claims adjusters ask natural-language questions about 12 years of policy documents and claim histories stored in Cloud Storage. Leadership wants a working prototype in two weeks, minimal model-tuning effort, and answers that cite the exact source document. Which Google Cloud approach should the team implement?

A.Build a retrieval-augmented generation flow using Vertex AI Search to index the corpus and ground Gemini responses in retrieved passages.
B.Increase the Gemini model's context window to the maximum and paste the entire document corpus into every prompt.
C.Fine-tune a Gemini model on the historical claims corpus, then deploy the tuned endpoint behind an internal web app.
D.Train a custom text embedding model from scratch on the claims corpus, then serve similarity search from a self-managed cluster.
AnswerA

Vertex AI Search indexes the Cloud Storage corpus and returns relevant passages at query time, which Gemini then uses as grounding context. This delivers citable, source-linked answers without retraining, and the managed indexing pipeline keeps the prototype within a two-week window while honoring the citation requirement.

Why this answer

Retrieval-augmented generation pairs a managed search index with a foundation model, so answers are grounded in retrieved passages and can point back to the originating document. Vertex AI Search handles ingestion and indexing of the Cloud Storage corpus, while Gemini synthesizes the retrieved context into an answer. This avoids costly fine-tuning and oversized prompts while meeting the citation and speed requirements.

Exam trap

The trap here is assuming that fine-tuning or a larger context window can substitute for retrieval when the real requirement is traceable, up-to-date grounding in a large document corpus.

82
Multi-Selectmedium

Which TWO strategies are effective for reducing latency in a generative AI chat application deployed on Vertex AI? (Select 2)

Select 2 answers
A.Deploy on TPU instead of GPU
B.Use streaming responses
C.Increase the max output tokens
D.Enable model quantization
E.Use larger batch sizes
AnswersB, D

Reduces perceived latency.

Why this answer

Streaming responses reduce perceived latency by sending tokens to the client as they are generated, rather than waiting for the full response. This leverages server-sent events (SSE) or chunked transfer encoding to deliver partial results immediately, improving user experience in chat applications.

Exam trap

Google Cloud often tests the distinction between reducing actual latency (e.g., model optimization) versus reducing perceived latency (e.g., streaming), and candidates mistakenly choose options that increase throughput (like larger batch sizes) without realizing they harm per-request latency.

83
MCQmedium

A startup is building a generative AI tool that helps users write code. They want to launch quickly but need to ensure the generated code is secure and does not introduce vulnerabilities. They have a small team of developers with some ML experience. The tool should be cloud-hosted. Which approach balances speed, security, and cost?

A.Deploy the tool without any security checks and rely on manual review
B.Train a custom code generation model from scratch on a large dataset
C.Use a pre-trained code model (e.g., Codey) and add a security filtering layer
D.Use a smaller model and restrict outputs to only simple code patterns
AnswerC

A pre-trained code model removes the cost and delay of training from scratch, letting a small ML team launch quickly, while the added security filtering layer scans generated code for vulnerabilities. This satisfies the need to balance speed, security and cost.

Why this answer

Using a pre-trained code model such as Codey provides immediate capability without the cost and time of training from scratch, while adding a security filtering layer (SAST-style scanning, output sanitization, policy checks) addresses the security requirement. This balances speed, security, and cost for a small team.

Exam trap

The trap is assuming that a pre-trained model alone is secure, or that training from scratch yields better security — in reality, guardrails around a pre-trained model deliver the best speed/security/cost balance.

How to eliminate wrong answers

Option A is wrong because relying solely on manual review does not scale and defeats the purpose of automated generation — security vulnerabilities would slip through. Option B is wrong because training a custom code model from scratch requires massive datasets, GPU compute, and ML expertise the small team lacks, blowing the time and cost budget. Option D is wrong because restricting outputs to simple patterns severely limits usefulness and does not guarantee security — simple code can still contain injection flaws.

84
MCQhard

A healthcare startup is exploring GenAI for clinical note summarization. They have concerns about patient data privacy. Which Google Cloud approach best addresses privacy while still using powerful models?

A.Deploy open-source models on-premises
B.Use a third-party API with anonymization of patient data
C.Use Vertex AI with model customization (fine-tuning)
D.Use Vertex AI with data residency controls and no external data sharing
AnswerD

Vertex AI keeps prompts and patient data within the selected region, and Google contractually excludes them from model training, satisfying the residency and confidentiality constraints. Data residency controls pin storage and processing to approved locations, while no external sharing prevents third-party exposure — meeting healthcare privacy requirements without sacrificing access to powerful models.

Why this answer

Vertex AI with data residency controls and no external data sharing ensures that patient data remains within specified geographic boundaries and is not used for model training or improvement, directly addressing healthcare privacy regulations like HIPAA. This approach leverages Google Cloud's powerful models while maintaining strict data governance, unlike options that risk data exposure or lack enterprise-grade controls.

Exam trap

The trap here is that candidates often assume fine-tuning (Option C) inherently provides privacy, but without explicit data residency and no-sharing policies, it fails to meet strict healthcare compliance requirements.

How to eliminate wrong answers

Option A is wrong because deploying open-source models on-premises, while offering data control, often lacks the advanced summarization capabilities and scalability of Vertex AI's foundation models, and still requires significant effort to ensure HIPAA compliance without Google's built-in privacy safeguards. Option B is wrong because using a third-party API, even with anonymization, introduces risks of data leakage or re-identification, and typically does not provide contractual guarantees against model training on patient data, violating many healthcare privacy policies. Option C is wrong because fine-tuning a model on Vertex AI without explicit data residency controls and no external data sharing may still allow Google to process data outside desired regions or use it for service improvements, failing to meet strict data privacy requirements.

85
Multi-Selecthard

A financial institution is deploying a generative AI solution that generates investment advice. They must ensure fairness, avoid toxic outputs, and comply with regulations like GDPR. Which TWO strategies should they implement? (Choose two.)

Select 2 answers
A.Use Vertex AI Safety Attributes to filter harmful content in both input and output.
B.Set the model temperature to 0 to eliminate creativity and reduce bias.
C.Implement a human review process for any advice above a certain risk threshold.
D.Fine-tune the model exclusively on compliant financial documents.
E.Disable request logging to avoid storing sensitive data.
AnswersA, C

Vertex AI Safety Attributes provides built-in safety filters that detect and block harmful content (e.g., hate speech, toxicity, financial misinformation) in both user prompts and model outputs, directly addressing the need to avoid toxic outputs and comply with regulations like GDPR.

Why this answer

Vertex AI Safety Attributes provides built-in safety filters that can detect and block harmful content (e.g., hate speech, toxicity, financial misinformation) in both user prompts and model outputs. This directly addresses the need to avoid toxic outputs and comply with regulations like GDPR, which require protecting users from harmful or biased advice.

Exam trap

A common misconception is that reducing model temperature or fine-tuning on compliant data alone can ensure safety and regulatory compliance, when in fact these measures do not address dynamic, context-dependent toxic outputs or logging requirements.

86
MCQhard

A company deployed a large language model on Vertex AI using the configuration shown in the exhibit. During peak usage, users report high latency. Which change is most likely to improve latency?

A.Remove the accelerator to simplify deployment.
B.Increase minReplicaCount to 3.
C.Switch to a GPU with more memory, such as NVIDIA_TESLA_A100.
D.Change machineType to n1-standard-4 to reduce cost.
AnswerB

Vertex AI scales replicas between minReplicaCount and maxReplicaCount; raising the minimum to 3 keeps additional instances warm, so peak traffic is spread across more replicas and per-request latency falls. Cold-start delays from scaling up are avoided.

Why this answer

Increasing minReplicaCount to 3 ensures that at least three instances of the model are always running and ready to serve requests. This reduces cold-start latency and distributes the load across multiple replicas, directly addressing high latency during peak usage by providing more concurrent serving capacity.

Exam trap

The Generative AI Leader exam often tests the misconception that upgrading hardware (GPU memory or type) is the primary fix for latency, when in fact scaling out replicas is the more direct solution for handling concurrent request load.

How to eliminate wrong answers

Option A is wrong because removing the accelerator (GPU/TPU) would force the model to run on CPU, drastically increasing inference latency, especially for large language models. Option C is wrong because switching to a GPU with more memory (NVIDIA_TESLA_A100) does not directly improve latency; it addresses memory capacity issues, not the throughput bottleneck caused by insufficient replicas. Option D is wrong because changing machineType to n1-standard-4 reduces CPU and memory resources, which would likely increase latency rather than improve it, and cost reduction is not the goal here.

87
MCQeasy

A retail company with a large FAQ database wants to build a generative AI customer service chatbot that can answer questions accurately with up-to-date information. Which business strategy should they prioritize?

A.Use retrieval-augmented generation (RAG) with vector search on the FAQ database.
B.Train a new model from scratch using the FAQ data.
C.Fine-tune a foundational model on the entire FAQ dataset.
D.Use a general-purpose language model without any customization.
AnswerA

RAG with vector search retrieves relevant FAQ passages at query time and grounds the model's responses in them, keeping answers accurate and current without retraining. This directly satisfies the up-to-date information constraint while leveraging the existing FAQ database.

Why this answer

Retrieval-augmented generation (RAG) with vector search allows the chatbot to dynamically retrieve the most relevant, up-to-date FAQ entries from a large database at inference time, grounding the generative model's responses in verified content without requiring retraining. This approach combines the flexibility of a pre-trained language model with the accuracy of real-time information retrieval, ensuring answers reflect the latest FAQ updates.

Exam trap

Google Cloud often tests the misconception that fine-tuning is the best way to inject domain knowledge, but the trap here is that fine-tuning cannot efficiently handle frequently changing data, whereas RAG provides a modular, update-friendly architecture that avoids retraining costs.

How to eliminate wrong answers

Option B is wrong because training a new model from scratch on FAQ data is computationally prohibitive, requires massive datasets and resources, and still cannot guarantee up-to-date answers without frequent retraining. Option C is wrong because fine-tuning a foundational model on the entire FAQ dataset risks catastrophic forgetting of general language capabilities and does not inherently handle dynamic updates; any FAQ change would require re-fine-tuning. Option D is wrong because a general-purpose language model without customization lacks domain-specific knowledge and cannot access the company's proprietary FAQ database, leading to hallucinated or outdated answers.

88
MCQeasy

A startup with limited budget wants to quickly test a generative AI use case for personalized email marketing. Which approach minimizes time-to-market and cost?

A.Hire a team of AI researchers to build a solution.
B.Develop a custom model from scratch.
C.Fine-tune a large open-source model on internal data.
D.Use a managed API like the PaLM API with prompt engineering.
AnswerD

A managed API such as the PaLM API removes infrastructure, training, and hosting overhead, letting the startup validate personalised email marketing through prompt engineering alone. This directly minimises both time-to-market and cost, matching the limited-budget constraint.

Why this answer

Using a managed API like the PaLM API with prompt engineering eliminates the need for infrastructure setup, model training, and data preparation. This approach leverages a pre-trained model via a simple REST API call, allowing the startup to iterate on prompts and achieve personalized email content in hours rather than weeks, minimizing both time-to-market and cost.

Exam trap

Google Cloud often tests the misconception that fine-tuning (Option C) is always the fastest and cheapest path for customization, but the trap here is that fine-tuning still requires significant compute and data preparation, whereas prompt engineering on a managed API is truly zero-infrastructure and pay-per-use, making it the optimal choice for a quick, low-cost test.

How to eliminate wrong answers

Option A is wrong because hiring a team of AI researchers is expensive and time-consuming, requiring salaries, compute resources, and months of development, which contradicts the limited budget and quick testing goal. Option B is wrong because developing a custom model from scratch demands vast amounts of labeled data, significant GPU/TPU compute, and deep expertise, making it cost-prohibitive and slow for a rapid proof-of-concept. Option C is wrong because fine-tuning a large open-source model still requires substantial compute for training (e.g., GPU hours for LoRA or full fine-tuning), data curation, and deployment overhead, which exceeds the minimal cost and speed constraints of a quick test.

89
MCQhard

A global bank wants to deploy a generative AI assistant for employees across multiple European countries, each with strict data residency laws. Which deployment strategy is most compliant?

A.Deploy separate model instances in each country's cloud region.
B.Use a federated learning approach where data stays on-premises.
C.Deploy a single model in a US region and use data masking.
D.Use a third-party API that processes data outside Europe.
AnswerA

Ensures data never leaves the country, meeting local compliance requirements.

Why this answer

Deploying separate model instances in each country's cloud region ensures that data never crosses national borders, directly complying with strict data residency laws like the GDPR's data localization requirements. This strategy uses regional cloud infrastructure (e.g., AWS eu-central-1, Azure westeurope) to keep both training and inference data within the specific jurisdiction, avoiding any cross-border data transfer.

Exam trap

Google Cloud often tests the misconception that data masking or anonymization alone satisfies data residency laws, but the trap here is that data residency requires the data to physically remain within the jurisdiction, not just be obfuscated.

How to eliminate wrong answers

Option B is wrong because federated learning only keeps training data on-premises, but the model parameters or gradients must still be exchanged with a central server, which can violate data residency if that server is outside the country. Option C is wrong because deploying a single model in a US region and using data masking does not prevent the underlying data from being processed or stored in the US, which violates EU data residency laws like GDPR. Option D is wrong because using a third-party API that processes data outside Europe directly violates data residency requirements, as the data physically leaves the European Economic Area (EEA) without adequate safeguards.

90
MCQmedium

A regional insurance company wants to launch a generative AI assistant that answers policyholder questions. Legal requires that every customer-facing answer be traceable to an approved policy document, and that the system never invent coverage terms. The company has about 40,000 internal policy PDFs that change quarterly. Which approach should the GenAI Leader recommend?

A.Deploy the assistant on Gemini with a system instruction telling it to be accurate and to avoid making up policy details.
B.Fine-tune a foundation model on the full set of policy PDFs and deploy the tuned model to a Vertex AI endpoint for the assistant to call.
C.Use a large context window model and paste the entire policy library into every prompt so the model always sees all approved documents.
D.Ground the assistant with retrieval-augmented generation using Vertex AI Search over the policy document corpus, and require citations in responses.
AnswerD

Retrieval-augmented generation retrieves the most relevant approved policy passages at query time and instructs the model to answer only from that retrieved context, so every answer can cite a source document. Because Vertex AI Search indexes the corpus and can be re-indexed each quarter, the assistant stays aligned with current policy language instead of relying on parametric memory that cannot be audited.

Why this answer

Grounding the assistant in the approved corpus with retrieval-augmented generation satisfies traceability because each response is generated from retrieved passages that can be cited, and it keeps answers current because the index is refreshed when policies change. Fine-tuning, whole-corpus prompting, and instruction-only approaches all leave the model without a verifiable source for coverage terms.

Exam trap

The trap here is assuming that fine-tuning on internal documents produces citations, when tuning only changes model weights and leaves no retrievable source to reference.

91
MCQeasy

A business wants to build a generative AI application but has limited data science resources. What is the recommended path?

A.Use Vertex AI's AutoML and pre-built APIs to accelerate development
B.Hire a team of ML engineers to develop an in-house solution
C.Purchase a third-party generative AI SaaS product off-the-shelf
D.Build a custom model from scratch using TensorFlow
AnswerA

Vertex AI's AutoML and pre-built APIs let teams with limited data science expertise train custom models and integrate generative capabilities without building pipelines from scratch. This accelerates development, satisfying the constraint of scarce specialist resources.

Why this answer

Vertex AI's AutoML and pre-built APIs are the recommended path because they allow the business to leverage Google's managed infrastructure and pre-trained models, significantly reducing the need for in-house data science expertise. AutoML automates model training, tuning, and deployment, while pre-built APIs (e.g., for vision, language) provide immediate access to generative capabilities without custom development. This approach accelerates time-to-market and lowers the barrier to entry for organizations with limited ML resources.

Exam trap

A common mistake is to assume that limited data science resources require outsourcing all AI work (Option C) or building from scratch (Option D), when the correct answer uses Google's managed services to reduce the need for in-house expertise while still allowing customization.

How to eliminate wrong answers

Option B is wrong because hiring a full team of ML engineers is resource-intensive and contradicts the premise of limited data science resources; it also introduces significant overhead in recruitment, management, and infrastructure. Option C is wrong because purchasing a third-party SaaS product off-the-shelf may not offer the customization, data privacy controls, or integration flexibility needed for a generative AI application, and it can lock the business into a vendor's roadmap. Option D is wrong because building a custom model from scratch using TensorFlow requires deep ML expertise, extensive training data, and computational resources, which is impractical for a team with limited data science capabilities and would delay deployment.

92
MCQhard

A utility company's board approves a generative AI program and asks the program lead to present a plan showing how investment will be governed and how value will be tracked from pilot through production. The lead wants a framework that ties each initiative to a business owner, defines stage gates for continued funding, and specifies which metrics justify scaling. Which approach best meets the board's expectation?

A.Adopt an AI acceptable use policy that prohibits employees from entering confidential data into public generative AI tools.
B.Standardize on a single foundation model and require every initiative to use it to simplify procurement and support.
C.Create a value realization framework that links each generative AI initiative to a business owner, defines stage-gate criteria for continued investment, and specifies scale-up metrics.
D.Build a centralized generative AI center of excellence that provides prompt libraries and engineering support to all business units.
AnswerC

The board asked for governance of investment and tracking of value across the lifecycle. A value realization framework assigns accountability, sets explicit stage gates that determine whether funding continues after each phase, and names the metrics that must be met before scaling, which is precisely the structure needed to govern a multi-initiative generative AI program from pilot to production.

Why this answer

The board is asking for investment governance and value tracking across the program lifecycle. A value realization framework supplies exactly that: a named business owner per initiative, stage gates that decide whether funding continues after each phase, and pre-agreed metrics that must be met before scaling. This makes continued investment evidence-based rather than momentum-driven.

Exam trap

The trap here is answering a funding governance question with a technical or policy control such as model standardization or an acceptable use policy, neither of which defines ownership, stage gates, or scale-up metrics.

93
MCQmedium

Refer to the exhibit. A sudden surge of traffic reaches 15,000 requests per second, but the endpoint can only handle 1,000 req/s per replica. What will happen to new requests?

A.They will be processed, and replicas will exceed maxReplicaCount.
B.They will be redirected to a different model.
C.They will receive HTTP 429 (Too Many Requests) errors.
D.They will be queued until capacity becomes available.
AnswerC

When incoming traffic exceeds the endpoint's per-replica capacity, the front end throttles rather than queues indefinitely, returning HTTP 429 responses to excess callers. The 15,000 req/s surge against 1,000 req/s per replica therefore produces Too Many Requests errors for new requests.

Why this answer

When a surge of 15,000 requests per second hits an endpoint configured with a maxReplicaCount (e.g., 10 replicas at 1,000 req/s each = 10,000 req/s capacity), any excess requests beyond that capacity are rejected with an HTTP 429 (Too Many Requests) status code. This is standard behavior in autoscaling systems: once the replica count reaches its maximum limit, the service cannot scale further, and new requests are throttled to prevent overload.

Exam trap

The trap here is that candidates assume autoscaling can handle any traffic surge indefinitely, ignoring the hard limit of maxReplicaCount, and thus incorrectly choose Option A or D, failing to recognize that HTTP 429 is the standard throttling mechanism when capacity is exhausted.

How to eliminate wrong answers

Option A is wrong because the maxReplicaCount is a hard upper limit; replicas cannot exceed this configured value, so new requests are not processed beyond that capacity. Option B is wrong because traffic redirection to a different model is not a standard behavior for capacity overflow; it would require explicit routing rules or a load balancer configured for failover, which is not implied in the scenario. Option D is wrong because queuing is not the default behavior for HTTP-based endpoints in this context; while some systems support request queuing (e.g., with message brokers), the exhibit describes a direct endpoint handling, and HTTP 429 is the standard response for rate limiting per RFC 6585.

94
MCQmedium

A national retail chain wants to deploy a generative AI assistant that recommends products to shoppers in six countries. The company's legal team requires that customer conversations never leave the country of origin, while the engineering team wants one consistent deployment pattern across all regions. Which Google Cloud approach best satisfies both requirements?

A.Configure VPC Service Controls per country and continue serving all shoppers from one central region.
B.Use a multi-region bucket for conversation logs and set a Cloud Storage lifecycle rule that deletes objects after 30 days.
C.Deploy the assistant separately in a Google Cloud region within each country and keep conversation storage in that same region.
D.Deploy the assistant in a single global region and rely on Google's global load balancing to route users to the nearest edge point of presence.
AnswerC

Running the assistant in a region inside each country keeps both inference and stored conversation data within national borders, satisfying the residency requirement. Reusing the same architecture, prompts, and deployment tooling in every region preserves the single consistent pattern the engineering team wants, so neither team's constraint is compromised.

Why this answer

Data residency for generative AI workloads requires that both inference and any retained conversation data physically stay inside the required jurisdiction. Only a per-country regional deployment achieves that while letting the team reuse one design, one prompt library, and one deployment mechanism everywhere. Perimeter controls, global routing, and multi-region storage change access or durability, not the physical location of processing.

Exam trap

The trap here is assuming that global load balancing or VPC Service Controls change where data is processed, when they only affect traffic routing and access boundaries.

95
Multi-Selecthard

A financial services firm is preparing a board presentation on its generative AI program. Directors want assurance that spending is disciplined and that failures are contained. Which TWO practices should the program adopt? (Choose two.)

Select 2 answers
A.Require every generative AI use case to run on the same foundation model to simplify vendor management.
B.Defer all evaluation of model output quality until after the production launch to save time.
C.Fund generative AI initiatives through staged gates, releasing additional budget only when agreed outcome metrics are met.
D.Centralize all generative AI engineering in one team that owns every use case end to end.
E.Define explicit stop conditions and decommission criteria for each initiative before development begins.
AnswersC, E

Staged funding converts a large upfront commitment into a series of smaller, evidence-gated investments. Each gate requires demonstrated progress against the agreed outcome metric, so weak initiatives stop early and capital flows to the strongest candidates. This directly answers the board's demand for disciplined spending while preserving room to scale winners quickly.

Why this answer

Board-level confidence in a generative AI program comes from disciplined capital allocation and bounded downside. Staged funding tied to outcome metrics ensures money follows evidence, while predefined stop and decommission criteria guarantee that underperforming initiatives end quickly and cheaply. Together they make the portfolio's risk explicit and controllable, which is what directors are asking to see.

Exam trap

The trap here is assuming discipline comes from standardization or centralization, when it actually comes from how funding is gated and how failures are bounded.

96
MCQeasy

A retail marketing team wants to generate product descriptions in five languages. The team has no machine learning engineers and wants to avoid managing any infrastructure. Which Google Cloud option should the GenAI Leader recommend first?

A.Provision a Vertex AI training job with custom containers and train a multilingual model from scratch on product data.
B.Call the Gemini API on Vertex AI with prompts requesting the product description in each target language.
C.Create a Google Kubernetes Engine cluster with GPU node pools and self-host an open model behind an inference server.
D.Build a BigQuery ML remote model that calls a public translation endpoint and post-process the output with SQL transformations.
AnswerB

The Gemini API on Vertex AI is a fully managed service, so the team sends prompts and receives multilingual text without provisioning servers, GPUs, or training pipelines. Gemini handles multiple languages natively, letting a small marketing team generate descriptions in five languages through simple API calls or the Google Cloud console, which matches their skill set and no-infrastructure constraint.

Why this answer

A managed foundation model API is the fastest route to multilingual generation for a team without ML engineers, because Google operates the model, scaling, and serving. The other choices require training pipelines, cluster operations, or an analytics-only pattern that does not generate marketing copy from product attributes.

Exam trap

The trap here is equating 'no infrastructure' with 'no cloud service,' when a managed API is precisely the option that removes infrastructure responsibility.

97
MCQhard

A company is using generative AI for code generation and wants to evaluate the quality of generated code for security vulnerabilities. Which metric is most appropriate?

A.BLEU score
B.Automatic static analysis
C.Human evaluation
D.Perplexity
AnswerB

Automatic static analysis scans generated code for insecure patterns such as injection flaws, hardcoded credentials and unsafe API calls, giving an objective, repeatable security measure. Functional or similarity metrics cannot detect vulnerabilities, so static analysis directly satisfies the requirement to evaluate security quality.

Why this answer

(Automatic static analysis) is correct because it directly scans code for security vulnerabilities, making it the most appropriate metric for this purpose. Option A (BLEU score) measures text similarity, not security. Option C (Human evaluation) is subjective and less scalable.

Option D (Perplexity) measures language model confidence, not code security.

98
Multi-Selecteasy

A company is adopting generative AI for customer support. Which TWO strategies should they implement to manage risks related to brand reputation?

Select 2 answers
A.Establish a human-in-the-loop escalation process for sensitive interactions.
B.Publish a disclaimer that the AI may make mistakes.
C.Implement automated monitoring for toxic or off-brand language.
D.Deploy the model without any content filters to maximize helpfulness.
E.Disable customer support AI entirely to avoid any risk.
AnswersA, C

Sensitive or ambiguous queries can produce harmful or off-brand replies, so routing them to a human before the response reaches the customer prevents reputational damage. This satisfies the brand-reputation risk constraint by keeping a person accountable for high-stakes interactions.

Why this answer

Option A is correct because a human-in-the-loop escalation process ensures that sensitive or high-stakes customer interactions are reviewed by a person before or during resolution, which directly protects brand reputation by preventing an AI from delivering inappropriate, harmful, or incorrect responses in delicate situations. Option C is correct because automated monitoring for toxic or off-brand language allows the company to detect and remediate reputation-damaging outputs in real time or near real time, maintaining consistent brand voice and catching harmful content that could erode customer trust. Option B is not a sufficient risk-management strategy because a disclaimer only shifts liability and does not prevent reputational harm from bad AI outputs.

Option D is incorrect because deploying a model without content filters increases the risk of toxic, biased, or off-brand responses, which is the opposite of reputation risk management. Option E is incorrect because disabling the AI entirely avoids risk only by eliminating the business benefit, which is not a viable strategy for a company adopting generative AI for customer support.

Exam trap

Google Cloud often tests the distinction between passive risk communication (like disclaimers) and active risk mitigation (like human-in-the-loop or automated monitoring), trapping candidates who think a disclaimer is sufficient to manage brand reputation risk.

99
MCQmedium

A global logistics firm wants to add a generative AI feature that drafts replies to customer shipment inquiries. The team must prove business value to executives within one quarter, keep engineering effort low, and later swap in a different model without rewriting the application. Which design decision best supports those goals?

A.Embed model-specific SDK calls directly in the application so each provider's latest features are immediately available.
B.Deploy an open-weights model on a single virtual machine and expose it through a custom REST endpoint.
C.Call Gemini through the Vertex AI API and isolate prompt construction and model selection behind an internal abstraction layer.
D.Train a bespoke language model on historical shipment correspondence and serve it from a self-managed GPU cluster.
AnswerC

Vertex AI provides a managed, enterprise-ready endpoint for Gemini, and placing prompt building and model selection behind an internal abstraction lets the team change models or parameters with minimal application changes. This keeps initial engineering effort low while preserving future flexibility, supporting a one-quarter value demonstration.

Why this answer

Using a managed Vertex AI endpoint for Gemini accelerates delivery and removes infrastructure work, while an internal abstraction layer for prompts and model selection decouples the application from any single provider. That combination lets the firm demonstrate value quickly and later change models by adjusting one integration point rather than rewriting business logic.

Exam trap

The trap here is equating rapid delivery with tight coupling to one model's SDK, when the stated requirement to swap models later makes that coupling a liability.

100
Multi-Selecteasy

A company is using Vertex AI generative models for a high-volume text summarization service. Which two strategies can reduce operational costs?

Select 2 answers
A.Increase the model's max output tokens to 2048.
B.Implement retry logic with exponential backoff.
C.Lower the temperature parameter to 0.
D.Use batch prediction instead of online prediction.
E.Reduce the size of the model (e.g., switch from text-bison@002 to text-bison-light).
AnswersD, E

Batch prediction has lower per-request cost for large jobs compared to online prediction.

Why this answer

Batch prediction reduces costs by processing multiple requests in a single batch job, which avoids the per-request overhead and idle compute time associated with online prediction. This is especially cost-effective for high-volume, non-real-time workloads like text summarization, as you pay only for the compute time used during the batch job rather than for each individual inference.

Exam trap

Google Cloud often tests the misconception that adjusting inference parameters like temperature or output length can reduce costs, when in reality only reducing model size or switching to batch processing directly lowers operational expenses.

101
MCQeasy

A marketing team wants to use generative AI to draft product descriptions. Legal requires that no customer data or proprietary campaign plans appear in any prompt sent to the model. Which practice should the team adopt first?

A.Ask the model provider to sign a data processing addendum covering all prompts submitted by the team.
B.Configure the model with a low temperature to reduce the chance that sensitive text is reproduced.
C.Enable a content filter on the model output to block any sensitive information that appears in drafts.
D.Establish prompt input guidelines and a review step that excludes customer data and proprietary campaign details.
AnswerD

The requirement is about what enters the prompt, so the first control is defining what may and may not be included and verifying it before submission. Clear guidelines plus a review step directly prevent sensitive content from being sent, and they create an auditable practice the legal team can inspect.

Why this answer

Because the restriction concerns what is placed into prompts, the effective first step is to define allowable prompt content and review submissions against it. That prevents sensitive data from leaving the organization and creates evidence of compliance. Temperature settings, output filtering, and contractual addenda do not control what data is included in the prompt, so they cannot satisfy the stated legal constraint.

Exam trap

The trap here is confusing output-side controls with input-side data governance when the restriction is specifically about what is sent to the model.

102
MCQmedium

A regional insurance company wants to launch a generative AI assistant that drafts policyholder responses for its claims team. Before any code is written, the CIO asks the team to produce a document that articulates the intended business outcome, the target user group, the success metrics, and the boundaries of what the assistant may and may not do. Which artifact best matches this request?

A.A Vertex AI model card documenting the training data, evaluation results, and known limitations of the underlying foundation model.
B.A generative AI use case definition that states the business objective, target users, measurable success criteria, and explicit scope boundaries.
C.A service-level objective document specifying latency percentiles and uptime targets for the assistant's API endpoint.
D.A cloud architecture diagram showing the assistant's front end, API layer, retrieval store, and model endpoint.
AnswerB

This artifact captures exactly what the CIO asked for: the business outcome, the intended user group, quantifiable success metrics, and the guardrails defining what the assistant should and should not handle. Framing these elements before implementation keeps the project anchored to measurable value and prevents scope drift once engineering begins, which is the foundational step of a generative AI business strategy on Google Cloud.

Why this answer

A generative AI initiative should begin with a use case definition that ties the technology to a business outcome, names the users it serves, defines how success will be measured, and sets explicit boundaries. This keeps the claims-drafting assistant aligned to measurable value and gives later technical and governance work a stable reference point before any build activity begins.

Exam trap

The trap here is assuming that a governance or architecture artifact such as a model card or diagram satisfies an upfront business framing request, when it describes the solution rather than the business problem.

103
MCQhard

A university wants to deploy a generative AI study assistant across six faculties. Each faculty has its own curriculum documents and privacy rules, and the central IT team must prevent one faculty's content from appearing in another faculty's answers. Which architecture decision best satisfies this governance requirement?

A.Fine-tune the model separately on each faculty's curriculum and expose a single endpoint that selects the adapter at random.
B.Deploy the assistant in a single project with one service account and apply IAM conditions based on the user's email domain.
C.Use one shared data store for all curricula and rely on prompt instructions telling the model to answer only from the requesting faculty's material.
D.Create a separate grounded data store and retrieval configuration per faculty, and route each request only to its own faculty's store.
AnswerD

Isolating each faculty's curriculum in its own grounded data store means retrieval can only return documents from the requesting faculty, so cross-faculty leakage is prevented by architecture rather than by instruction. Each faculty can also apply its own privacy rules to its store, and central IT retains one platform to operate across all six.

Why this answer

Separating grounded data stores per faculty enforces isolation at the retrieval layer, so answers can only draw on the requesting faculty's curriculum regardless of prompt wording. Each faculty keeps control of its own privacy rules, while central IT operates a single assistant platform. This structural separation is the only option that prevents leakage by design rather than by model compliance.

Exam trap

The trap here is trusting prompt instructions to enforce data boundaries, when retrieval isolation must be enforced in the architecture before the model ever sees the context.

104
Multi-Selecthard

Which THREE factors should be considered when choosing between a fine-tuned model and a prompted foundation model for a generative AI solution? (Select 3)

Select 3 answers
A.Need for domain-specific vocabulary
B.Inference latency requirements
C.Size of training data available
D.Whether the model is open-source
E.Token cost per request
AnswersA, C, E

Fine-tuning can incorporate domain language.

Why this answer

Fine-tuning allows the model to learn domain-specific vocabulary and terminology that may not be well-represented in the foundation model's pre-training data. This is critical for specialized fields like legal, medical, or technical domains where precise language is required for accurate outputs.

Exam trap

Google Cloud often tests the misconception that inference latency is a deciding factor between fine-tuning and prompting, when in reality both can be optimized for speed, and the key differentiators are data availability, domain specificity, and cost per token.

105
MCQeasy

A large e-commerce company is experiencing high costs for their generative AI product recommendation system. The system generates personalized product descriptions for millions of users daily. The team wants to reduce cost while maintaining quality. They are using a fine-tuned version of a large foundation model hosted on Vertex AI. The current cost is driven by the number of tokens processed. Which approach should they take?

A.Optimize prompts to generate shorter, more concise descriptions
B.Switch to a larger, more capable foundation model
C.Retrain the model with more product data to improve efficiency
D.Increase the batch size of inference requests
AnswerA

Shorter outputs use fewer tokens, reducing cost.

Why this answer

Prompt engineering to reduce output length decreases token usage per request, directly lowering cost without model changes. Option B (switching to a larger model) increases cost. Option C (increasing batch size) may not reduce per-request cost.

Option D (retraining with more data) does not affect inference cost.

106
Multi-Selecteasy

A company is choosing a generative AI model for code generation. Which TWO considerations are most important?

Select 2 answers
A.The total number of model parameters
B.Whether the model's training data includes the target programming languages
C.The open-source license of the model
D.The maximum context length supported by the model
E.The latency of the model's inference endpoint
AnswersB, D

Code generation depends on the model having seen the target programming languages during training; without that coverage, syntax, idioms and library usage are unreliable, so training-data language coverage directly determines output quality for the required languages.

Why this answer

Option B is correct because a code-generation model must have been trained on the target programming languages (e.g., Python, Java, C++) to produce syntactically valid and idiomatic code; a model lacking exposure to a language will generate unreliable or non-compiling output. Option D is correct because maximum context length determines how much source code, documentation, and surrounding files the model can ingest at once, which is critical for tasks like multi-file refactoring, repository-level completion, and understanding large codebases. Option A is not decisive because parameter count alone does not guarantee code quality; a smaller model fine-tuned on code can outperform a larger general-purpose model.

Option C is not a primary technical consideration for code-generation capability, though licensing matters for legal/commercial deployment rather than model effectiveness. Option E affects user experience and cost but is a deployment concern, not a core capability consideration for choosing a code-generation model.

Exam trap

The trap here is that candidates often assume more parameters (A) or lower latency (E) are always better, but Google tests the understanding that domain-specific training data relevance (B) and context length (D) are critical for code generation accuracy and handling long code sequences.

107
MCQeasy

A mid-sized insurance company wants to adopt generative AI to help claims adjusters draft customer correspondence. Executives are unsure how to prioritize the first use case and want a framework that balances business value with implementation feasibility. Which first step best aligns with a value-driven generative AI adoption strategy?

A.Wait until competitors publicly report their generative AI results before committing any budget.
B.Begin with the most technically complex use case to prove the organization's engineering capability.
C.Deploy generative AI across all claims functions simultaneously to maximize transformation speed.
D.Select the use case with the highest volume of manual effort and a clear measurable baseline for time saved.
AnswerD

Prioritizing a high-volume, effort-intensive task with an existing measurable baseline lets the organization demonstrate tangible return quickly and build momentum. Drafting correspondence is repetitive, so time saved per adjuster is easy to quantify. This aligns with a value-driven approach that pairs business impact with feasibility rather than chasing the most technically ambitious option first.

Why this answer

A value-driven generative AI strategy starts where effort is high, the task is repetitive, and a baseline already exists to measure improvement. High-volume correspondence drafting offers fast, quantifiable time savings that build credibility for later, more complex initiatives. This sequencing balances impact with feasibility instead of optimizing for technical novelty or waiting on the sidelines.

Exam trap

The trap here is equating ambition with strategy, assuming the most complex or broadest rollout is the best first move when measurable quick wins build the case for scaling.

108
MCQhard

A large insurance company is using generative AI to automate claims processing. They have deployed a custom fine-tuned model on Vertex AI that reads claim documents and extracts key information. Recently, they noticed that the model’s performance degrades over time for certain claim types, leading to incorrect payouts. The team needs to detect and address model drift with minimal manual intervention. They have a data pipeline that captures incoming claims and user feedback on predictions. Which approach should they take?

A.Implement a human review process for all claims the model processes
B.Set up continuous evaluation with automated retraining pipelines based on performance metrics
C.Switch to a simpler rule-based system to avoid drift
D.Manually retrain the model monthly using a snapshot of recent claims
AnswerB

Continuous evaluation monitors live performance metrics against baselines, and automated retraining pipelines refresh the model when drift is detected, correcting degradation with minimal manual intervention. This directly satisfies the requirement to detect and address drift automatically.

Why this answer

It establishes a closed-loop MLOps pipeline where continuous evaluation of performance metrics (e.g., precision, recall, or F1-score on streaming data) triggers automated retraining when drift is detected. This minimizes manual intervention while ensuring the model adapts to distribution shifts in claim types, which is critical for maintaining accurate payouts in production.

Exam trap

Google Cloud often tests the misconception that periodic manual retraining (Option D) is sufficient, but the trap here is that it ignores the need for real-time drift detection and automated response, which is essential for production systems handling high-stakes financial decisions.

How to eliminate wrong answers

Option A is wrong because implementing human review for all claims defeats the purpose of automation and introduces significant operational cost and latency, failing the requirement for minimal manual intervention. Option C is wrong because switching to a simpler rule-based system cannot handle the complexity and variability of claim documents, and it will still suffer from drift as claim patterns evolve over time. Option D is wrong because manually retraining monthly on a snapshot ignores real-time drift detection and may miss sudden shifts between retraining cycles, leading to prolonged periods of degraded performance.

109
MCQmedium

A media company uses generative AI to produce personalized news summaries for subscribers. They notice that the summaries sometimes contain factual inaccuracies, leading to customer complaints. The team needs to improve accuracy without slowing down the generation speed. They are using a pre-trained model via Vertex AI. What strategy should they implement?

A.Switch to a larger, more accurate foundation model
B.Fine-tune the model on a dataset of verified news articles
C.Implement retrieval-augmented generation (RAG) with a trusted knowledge base
D.Add a human-in-the-loop review for every summary
AnswerC

RAG grounds generation in a trusted knowledge base by retrieving relevant verified content and injecting it into the prompt, so summaries reflect accurate sources. This improves factual accuracy without retraining or slowing inference, unlike fine-tuning or larger models.

Why this answer

Retrieval-augmented generation (RAG) grounds the model's output in a trusted, external knowledge base, allowing it to retrieve verified facts in real time without retraining. This directly addresses factual inaccuracies while maintaining generation speed, as the pre-trained model remains unchanged and only the retrieval step is added. RAG avoids the latency of human review and the computational cost of fine-tuning or switching models.

Exam trap

Google Cloud often tests the misconception that fine-tuning is the default solution for accuracy issues, but the trap here is that RAG provides a faster, more scalable way to ground outputs in verified data without retraining, which is critical when speed and accuracy must both be maintained.

How to eliminate wrong answers

Option A is wrong because switching to a larger foundation model would increase inference latency and computational cost, contradicting the requirement to not slow down generation speed, and it does not guarantee improved factual accuracy without additional grounding. Option B is wrong because fine-tuning on a dataset of verified news articles requires significant time, data, and compute resources, and it may not prevent hallucinations on unseen topics, while also risking catastrophic forgetting of the model's general capabilities. Option D is wrong because adding a human-in-the-loop review for every summary introduces unacceptable latency and operational overhead, making it impractical for real-time personalized news generation at scale.

110
MCQhard

A large enterprise wants to deploy multiple generative AI models across different business units while ensuring cost governance and usage tracking. Which Google Cloud solution is best suited?

A.Use Vertex AI Endpoint with monitoring
B.Deploy each model in a separate project with IAM policies
C.Implement a custom cost allocation using labels
D.Use Cloud Billing budgets and alerts per model
AnswerB

Projects isolate resources and costs per business unit.

Why this answer

Deploying each model in a separate Google Cloud project with IAM policies provides the strongest isolation for cost governance and usage tracking. This approach ensures that each business unit's model usage is billed to its own project, enabling granular cost allocation and independent monitoring without cross-project interference. It also allows per-project budget alerts and usage quotas, directly addressing the enterprise's need for decentralized cost control.

Exam trap

The trap here is that candidates often confuse cost allocation mechanisms (like labels or budgets) with true cost isolation, assuming that tagging or alerting alone can enforce per-business-unit governance without the structural separation that separate projects provide.

How to eliminate wrong answers

Option A is wrong because Vertex AI Endpoint with monitoring tracks model performance and latency but does not inherently isolate costs per business unit; it aggregates usage under a single project, making per-unit cost governance difficult. Option C is wrong because custom cost allocation using labels requires manual tagging and can be inconsistent or incomplete, leading to inaccurate cost tracking; labels are metadata, not a billing boundary. Option D is wrong because Cloud Billing budgets and alerts per model are not natively supported; budgets apply at the project or billing account level, not per model, and cannot enforce cost isolation across multiple business units.

111
MCQeasy

A retail company plans to use Vertex AI's generative AI to create product descriptions. They need to ensure descriptions are factually accurate and do not misrepresent products. Which strategy should they prioritize?

A.Implement human-in-the-loop review
B.Use prompt engineering
C.Use a larger model
D.Increase temperature parameter
AnswerA

Human-in-the-loop review places a person between model output and publication, catching factual errors and misrepresentations before customers see them. For retail product descriptions where accuracy is a hard requirement, this verification step directly satisfies the constraint that generated text must not misstate product attributes.

Why this answer

Human-in-the-loop (HITL) review is the correct strategy because it directly addresses the need for factual accuracy and prevention of misrepresentation. While generative AI can produce fluent text, it lacks a reliable grounding mechanism for product-specific facts, making human oversight essential to catch hallucinations, verify claims, and ensure compliance with advertising standards. This approach aligns with responsible AI practices and is a core recommendation for high-stakes content generation.

Exam trap

Google Cloud often tests the misconception that prompt engineering or model size alone can solve factual accuracy issues, when in reality, generative AI's inherent lack of ground truth makes human validation indispensable for high-stakes content.

How to eliminate wrong answers

Option B is wrong because prompt engineering, while useful for guiding output style and structure, does not guarantee factual accuracy; it cannot prevent the model from generating plausible-sounding but incorrect product details. Option C is wrong because using a larger model may improve fluency and reduce some errors, but it does not eliminate hallucinations or misrepresentations, and can even introduce more subtle inaccuracies. Option D is wrong because increasing the temperature parameter makes the model's output more random and creative, which increases the risk of generating factually incorrect or misleading descriptions, the opposite of what is needed.

112
MCQeasy

A retail company wants to use gen AI for customer service chatbots. They have a large volume of customer interactions. What is the primary business consideration for deploying a gen AI solution?

A.Minimizing latency at any cost
B.Using open-source models only
C.Choosing the most complex model
D.Ensuring data privacy and compliance
AnswerD

Handling large volumes of customer interactions means personal data flows through the chatbot, so data privacy and compliance is the primary business consideration; it governs lawful processing, retention and regulatory risk under GDPR-style rules, directly satisfying the scenario's scale-driven exposure.

Why this answer

Ensuring data privacy and compliance is the primary business consideration when handling customer data, especially with regulatory requirements like GDPR or CCPA. Option A is wrong because while latency matters, minimizing it at any cost is not the primary consideration and can lead to excessive expenses. Option B is wrong because using only open-source models may limit flexibility and can raise data privacy concerns if not properly vetted.

Option C is wrong because choosing the most complex model often increases cost and latency without necessarily improving business outcomes.

113
MCQmedium

A mid-size retail company wants to launch a generative AI assistant that drafts promotional product descriptions for its marketing team. The team expects a rapid pilot, but the CIO insists that any generated text must never expose the company's unreleased product roadmap or pricing data that may exist in internal documents. The company has no dedicated AI engineering staff and prefers a managed approach on Google Cloud. Which strategy best balances rapid pilot delivery with this data-exposure requirement?

A.Use a managed generative AI API with grounding restricted to an approved, curated product catalog and clear system instructions limiting the assistant to that content.
B.Build a custom large language model from scratch using the company's historical marketing copy as the sole training corpus.
C.Deploy an open-weight model on a Compute Engine VM and allow the assistant to search the entire shared drive so the marketing team gets the most complete answers.
D.Fine-tune a foundational model on all internal product documents so the assistant learns the company's writing style and terminology.
AnswerA

A managed API removes infrastructure and model-ops work, enabling a fast pilot without AI engineering staff. Restricting grounding to an approved catalog and using system instructions keeps generation anchored to vetted content, so unreleased roadmap and pricing documents are never retrieved or exposed. This directly satisfies both the speed goal and the CIO's data-boundary requirement.

Why this answer

The scenario couples a fast, low-skill pilot with a hard boundary around sensitive internal content. A fully managed generative AI API eliminates infrastructure and model-tuning overhead, while grounding limited to an approved catalog plus explicit system instructions confines generation to vetted material, preventing unreleased roadmap and pricing data from appearing in drafts. This combination delivers speed without weakening the data-exposure control the CIO demanded.

Exam trap

The trap here is assuming that fine-tuning or self-hosting is required for brand consistency, when grounding and instructions actually control content boundaries without embedding sensitive data in model weights.

114
MCQmedium

A company wants to deploy a generative AI chatbot for customer service but is concerned about cost unpredictability due to variable usage. Which pricing model should they choose to best manage costs?

A.Committed use discounts
B.Free tier
C.Pay-as-you-go
D.Provisioned throughput
AnswerD

Provides fixed capacity with predictable monthly cost.

Why this answer

Provisioned throughput provides fixed capacity with predictable monthly cost, ideal for managing cost uncertainty. Option A (committed use discounts) requires a commitment but costs can still vary if usage exceeds the committed amount. Option B (free tier) is too limited for a full-scale chatbot.

Option C (pay-as-you-go) is variable and leads to cost unpredictability.

115
MCQeasy

A retail company wants to integrate generative AI into its customer service chatbot to handle routine inquiries. They have a limited budget and want to launch quickly. Which strategy is most appropriate?

A.Partner with a generative AI vendor for a custom solution
B.Use pre-trained models via Google Cloud's Generative AI Studio API
C.Fine-tune an open-source model on their customer service logs
D.Build a custom LLM from scratch using the company's own data
AnswerB

Pre-trained models accessed through an API remove the cost and time of training or hosting custom models, letting the retailer launch quickly within budget. The API handles inference, so only integration work remains, satisfying both the limited-budget and fast-launch constraints.

Why this answer

Using pre-trained models via Google Cloud's Generative AI Studio API allows the company to leverage existing, powerful models without the high cost and time investment of custom development or fine-tuning. This approach enables rapid deployment on a limited budget by simply integrating the API into their chatbot, handling routine inquiries effectively without requiring extensive machine learning expertise or infrastructure.

Exam trap

Google Cloud often tests the misconception that fine-tuning or custom models are always better for domain-specific tasks, but the trap here is that for routine inquiries with limited budget and time, pre-trained APIs offer the fastest and most cost-effective solution without sacrificing quality.

How to eliminate wrong answers

Option A is wrong because partnering with a generative AI vendor for a custom solution typically involves significant upfront costs, long development cycles, and vendor lock-in, which contradicts the company's limited budget and need for quick launch. Option C is wrong because fine-tuning an open-source model on customer service logs requires substantial computational resources, data preparation, and machine learning expertise, making it slower and more expensive than using a pre-trained API. Option D is wrong because building a custom LLM from scratch is extremely resource-intensive, requiring massive datasets, specialized hardware, and months of training, which is impractical for a company with limited budget and a need for speed.

116
MCQeasy

An e-commerce company is using a generative AI model to recommend products. They notice that the recommendations are often irrelevant. What is the most likely cause?

A.Using an outdated model version
B.Incorrect regional endpoint configuration
C.Inadequate prompt engineering
D.Overfitting on training data
AnswerC

Irrelevant recommendations typically stem from vague or poorly structured prompts that fail to supply the model with sufficient context about customer preferences, product attributes, or business goals. Refining prompt engineering—adding explicit constraints, examples, and desired output format—directly addresses this, making it the most likely cause given the scenario's lack of retrieval or training defects.

Why this answer

Inadequate prompt engineering is the most likely cause because generative AI models rely heavily on the quality and specificity of the input prompt to produce relevant outputs. If the prompts used to generate product recommendations are vague, poorly structured, or lack context (e.g., not including user preferences or historical behavior), the model will return generic or irrelevant suggestions. This is a common failure point in recommendation systems where the prompt acts as the primary interface for steering model behavior.

Exam trap

Google Cloud often tests the misconception that model performance issues are always due to training data or model version problems, when in fact prompt engineering is the most immediate and common cause of output irrelevance in generative AI systems.

How to eliminate wrong answers

Option A is wrong because using an outdated model version may affect performance or feature availability, but it does not directly cause irrelevant recommendations; the model would still generate outputs consistent with its training, and relevance is more tied to prompt quality. Option B is wrong because incorrect regional endpoint configuration would cause connectivity or latency issues (e.g., API timeouts or routing errors), not irrelevant content generation; the model's output relevance is independent of the endpoint's geographic location. Option D is wrong because overfitting on training data would cause the model to memorize specific patterns and perform poorly on new or diverse inputs, but in a recommendation context, overfitting typically leads to overly narrow or repetitive suggestions, not broadly irrelevant ones; the primary issue with irrelevant outputs is prompt misalignment, not training data memorization.

117
MCQeasy

A company is choosing between Google's Gemini API and an open-source model. Which factor is most important for a business with limited ML expertise?

A.Ease of integration and availability of support
B.Model parameter count
C.Cost per token
D.Community size
AnswerA

Managed APIs like Gemini ship with SDKs, documentation and vendor support, so teams lacking ML engineers avoid model hosting, tuning and MLOps overhead. That directly satisfies the limited-expertise constraint, whereas self-hosting an open-source model demands in-house skills for deployment, scaling and maintenance.

Why this answer

For a business with limited ML expertise, ease of integration and availability of support are paramount because they reduce the need for in-house machine learning engineering talent. Google's Gemini API offers managed infrastructure, pre-built SDKs, and enterprise-grade support (e.g., SLA-backed uptime, dedicated account management), which directly lowers the barrier to entry and operational risk. In contrast, open-source models require significant expertise for deployment, scaling, and troubleshooting, making them unsuitable for teams without deep ML skills.

Exam trap

The Generative AI Leader exam often tests the misconception that technical metrics like parameter count or cost per token are the primary decision factors, when in reality, for a non-expert team, operational simplicity and vendor support are the critical success factors that determine whether a GenAI project can be delivered at all.

How to eliminate wrong answers

Option B is wrong because model parameter count (e.g., 7B vs 175B) is a technical metric that does not directly address the business's lack of ML expertise; a larger parameter count can actually increase complexity and resource requirements, making it harder to integrate without expert knowledge. Option C is wrong because cost per token, while important for budgeting, is secondary to the ability to actually use the model; without easy integration and support, even a low-cost model can become expensive due to hidden engineering costs and downtime. Option D is wrong because community size, though helpful for troubleshooting, does not provide the structured, guaranteed support and SLAs that a business with limited ML expertise needs; community forums lack accountability and may not offer timely or accurate solutions for production-critical issues.

118
Multi-Selecthard

An energy utility is preparing a generative AI assistant that drafts responses to regulator inquiries. Before launch, the GenAI Leader must define how the program will be evaluated and governed on Google Cloud. Which TWO practices should be included? (Choose two.)

Select 2 answers
A.Grant the assistant's service account broad project-level Owner permissions so it can retrieve any internal document it may need.
B.Disable all logging of prompts and responses to avoid storing potentially sensitive regulatory content in Cloud Logging.
C.Rely solely on the foundation model's built-in safety filters as the complete control set for regulatory accuracy.
D.Track model and prompt versions alongside evaluation results so changes to the assistant can be compared and rolled back.
E.Establish a human review and approval step for every drafted regulatory response before it is sent.
AnswersD, E

Versioning the model, prompt template, and evaluation scores lets the team detect regressions when any component changes and restore a known-good configuration quickly. Without this traceability, a prompt tweak that reduces factual accuracy could go unnoticed in production. It is a foundational control for operating generative AI in a regulated environment where behavior must be explainable after the fact.

Why this answer

A defensible program pairs human approval of each external response with versioned tracking of models, prompts, and evaluation results so behavior is reproducible and reversible. Together these provide accountability and change control. Broad permissions, disabled logging, and reliance on safety filters alone all weaken oversight without addressing factual accuracy.

Exam trap

The trap here is treating built-in model safety filters as sufficient governance for factual and legal accuracy in a regulated workflow.

119
Multi-Selecthard

A financial services firm must comply with regulations when using gen AI. Which two measures are critical?

Select 2 answers
A.Implement audit trails
B.Deploy without risk assessment
C.Use a closed-source model
D.Use explainable AI
E.Use only synthetic data
AnswersA, D

Audit trails provide accountability and support regulatory reviews.

Why this answer

Audit trails are critical for compliance because they provide a tamper-evident, chronological record of all AI model inputs, outputs, and decisions. This enables firms to demonstrate regulatory adherence (e.g., under GDPR or SOX) by reconstructing the exact sequence of events that led to a specific AI-generated output, which is essential for accountability and forensic review.

Exam trap

Google Cloud often tests the misconception that 'closed-source models are inherently more compliant' or that 'synthetic data eliminates privacy risks,' when in reality, compliance hinges on transparency, auditability, and risk assessment rather than the model's source or data origin.

120
MCQmedium

A healthcare organization wants to use generative AI for medical report summaries. What is the primary concern?

A.Ensuring HIPAA compliance and data security when using cloud AI services
B.The model's ability to generate fluent and coherent summaries
C.Minimizing the cost of each API call to stay within budget
D.Latency of responses for real-time use cases
AnswerA

Medical report summaries contain protected health information, so sending that data to a cloud generative AI service raises HIPAA compliance and data security obligations. The organisation must ensure the provider signs a business associate agreement and safeguards the data.

Why this answer

The primary concern for a healthcare organization using generative AI for medical report summaries is ensuring HIPAA compliance and data security when using cloud AI services. Medical data is protected health information (PHI), and any cloud-based AI service must have a Business Associate Agreement (BAA) in place and enforce encryption at rest and in transit to avoid regulatory penalties and data breaches.

Exam trap

Google Cloud often tests the misconception that technical performance (fluency, cost, latency) is the top priority, when in regulated industries like healthcare, compliance and data security are the non-negotiable primary concerns.

How to eliminate wrong answers

Option B is wrong because while fluency and coherence are important for summary quality, they are secondary to the legal and security obligations of handling PHI; a fluent summary that leaks data is non-compliant. Option C is wrong because cost minimization is an operational concern, not the primary risk; HIPAA violations carry fines up to $50,000 per violation, far outweighing API call costs. Option D is wrong because latency is a performance metric relevant for real-time use, but medical report summarization is typically asynchronous or batch-processed, and compliance takes precedence over speed.

121
Multi-Selecteasy

Which THREE are essential components of a responsible AI strategy for GenAI? (Select three.)

Select 3 answers
A.Use of only open-source models
B.Maximum model size
C.Human oversight for critical decisions
D.Model transparency and explainability
E.Bias detection and mitigation
AnswersC, D, E

Human oversight prevents harmful automated decisions and ensures ethical use.

Why this answer

Human oversight for critical decisions (C) is essential because GenAI models can produce plausible but incorrect or harmful outputs. A responsible AI strategy mandates that a human-in-the-loop reviews high-stakes outputs, such as medical diagnoses or financial approvals, to prevent automated errors from causing real-world harm. This aligns with the principle of human accountability in AI governance frameworks like the NIST AI Risk Management Framework.

Exam trap

Google Cloud often tests the misconception that technical attributes like model size or open-source licensing are core to responsible AI, when in fact the focus is on governance practices like transparency, bias mitigation, and human oversight.

122
MCQhard

Refer to the exhibit. This JSON describes a Vertex AI endpoint with a deployed model. Which statement about scaling is true?

A.The endpoint uses only dedicated resources, no automatic scaling
B.The endpoint will automatically scale based on GPU utilization
C.The endpoint will scale from 1 to 3 replicas based on load using automatic scaling
D.The endpoint can scale to zero when not in use
AnswerA

DedicatedResources with min/max replicas means manual scaling.

Why this answer

The JSON shows that the endpoint is configured with `dedicatedResources` and no `autoscalingMetricSpecs` or `minReplicaCount`/`maxReplicaCount` fields. In Vertex AI, when you specify only `machineSpec` and a fixed `minReplicaCount` (here implicitly 1) without a `maxReplicaCount` or autoscaling metrics, the endpoint uses dedicated resources with no automatic scaling — the model will always run on exactly the number of replicas you define, regardless of load.

Exam trap

Google Cloud often tests the misconception that any endpoint with a `minReplicaCount` and `maxReplicaCount` automatically enables scaling, but the trap here is that without `autoscalingMetricSpecs`, the endpoint uses dedicated resources and does not scale dynamically — the `maxReplicaCount` is ignored if autoscaling metrics are absent.

How to eliminate wrong answers

Option B is wrong because Vertex AI automatic scaling is based on CPU utilization or custom metrics, not GPU utilization; GPU utilization is not a supported metric for autoscaling in Vertex AI endpoints. Option C is wrong because the JSON does not include `autoscalingMetricSpecs` or a `maxReplicaCount` field, which are required to enable automatic scaling from a minimum to a maximum number of replicas; without these, the endpoint uses a fixed replica count. Option D is wrong because Vertex AI endpoints with dedicated resources cannot scale to zero; scaling to zero is only possible with private endpoints using manual scaling or when using Vertex AI Prediction with a custom container that supports scale-to-zero, but dedicated resources always maintain at least one replica.

123
MCQmedium

A regional insurance provider wants to launch a generative AI claims assistant. The executive sponsor insists the solution be built on a foundation model the company can host inside its own Google Cloud project, with no dependence on a vendor's externally managed endpoint. Which decision does the sponsor need to make first?

A.Whether claim documents should be stored in Cloud Storage or BigQuery.
B.Which prompt engineering technique will produce the most accurate claim summaries.
C.Whether to use a fully managed model API or a model whose weights the company can deploy and control itself.
D.How many claims adjusters will be granted access to the assistant at launch.
AnswerC

The sponsor's constraint is about control over hosting and endpoints, which is fundamentally a build-versus-buy model-hosting decision. Resolving whether the company runs open-weight models on its own infrastructure or consumes a managed API determines architecture, cost model, staffing, and compliance posture, so it must be settled before any other design choice.

Why this answer

The sponsor's requirement is about where the model executes and who controls the endpoint, which is a hosting and control decision. Choosing between a managed model API and self-deployed model weights shapes architecture, operating cost, and compliance, and every other choice depends on it. Prompt design, user scope, and storage selection are important but subordinate to that first commitment.

Exam trap

The trap here is jumping to prompt or data decisions before resolving the hosting and control model that the stated constraint actually targets.

124
MCQmedium

A company's generative AI model is producing biased outputs. What is the most effective mitigation strategy?

A.Use a larger model with more parameters to improve overall accuracy
B.Fine-tune the model using a balanced, representative dataset and implement output filtering
C.Use prompt engineering to instruct the model to avoid biased language
D.Increase the diversity of input samples by random sampling
AnswerB

Fine-tuning on a balanced, representative dataset reduces the skewed correlations producing biased outputs, while output filtering catches residual harmful generations at inference. Together they address both the model's learned behaviour and runtime leakage, which prompt tweaks alone cannot reliably fix.

Why this answer

Fine-tuning on a balanced, representative dataset directly addresses the root cause of biased outputs by correcting the model's learned associations, while output filtering provides a safety net to catch residual bias. This combination is more effective than superficial fixes because it modifies the model's internal weights rather than just masking outputs.

Exam trap

Google Cloud often tests the misconception that prompt engineering or model scaling alone can fix bias, when in fact only retraining or fine-tuning with balanced data addresses the underlying weight distribution.

How to eliminate wrong answers

Option A is wrong because increasing model size does not inherently reduce bias; larger models can amplify biases present in training data due to higher capacity to memorize spurious correlations. Option C is wrong because prompt engineering only provides a surface-level instruction that the model may ignore or fail to generalize, especially if the bias is deeply embedded in its parameters. Option D is wrong because random sampling of inputs does not address the model's biased internal representations; it only diversifies the prompts, not the training data that caused the bias.

125
MCQmedium

A logistics company is selecting a generative AI use case to fund first. Leadership wants a project that demonstrates value quickly, has accessible data, and carries limited regulatory exposure. Which use case best fits these selection criteria?

A.Automating final customs classification decisions for international shipments without human review.
B.Generating draft responses to routine internal IT helpdesk tickets using an approved knowledge base.
C.Generating personalized medical advice for drivers based on wearable health data.
D.Replacing all human dispatchers with an autonomous agent that negotiates carrier contracts.
AnswerB

Drafting responses to routine internal helpdesk tickets uses an existing approved knowledge base, serves an internal audience, and has limited regulatory exposure compared with customer-facing or health-related data. Value can appear quickly through reduced handling time and faster resolution, and the scope is narrow enough to evaluate. This combination of accessible data, low compliance risk, and measurable productivity gain matches the leadership criteria for a first funded project.

Why this answer

An internal helpdesk drafting assistant draws on an approved knowledge base, affects only employees, and avoids the heavy regulatory exposure of customs, medical, or contract-negotiation scenarios. It can show measurable reductions in handling time within a short pilot, giving leadership evidence of value before funding riskier, customer-facing or regulated use cases. Accessible data and a narrow scope further support fast, defensible results.

Exam trap

The trap here is equating high business impact with suitability for a first project, when regulatory exposure, data accessibility, and time to demonstrable value should drive the initial selection.

126
MCQeasy

A logistics firm wants every generative AI proposal to be judged on whether it reduces cost per shipment. Leadership asks the AI team to define the metric before any project starts. Which practice does this illustrate?

A.Selecting a foundation model based on its published benchmark scores.
B.Adopting a responsible AI review board to approve all generative AI use cases.
C.Choosing a deployment region that minimizes network egress charges.
D.Establishing a measurable business outcome tied to an existing operational key performance indicator.
AnswerD

Tying the generative AI effort to cost per shipment connects the initiative to an operational metric the business already tracks and trusts. That makes value demonstrable, comparable across proposals, and defensible to finance, which is exactly what defining the metric before work begins is meant to achieve.

Why this answer

Value-driven generative AI programs start by naming the business outcome and linking it to a metric the organization already measures, such as cost per shipment. That anchor lets leadership compare proposals, size expected returns, and confirm impact after launch. Model benchmarks, egress savings, and governance approvals are supporting concerns rather than definitions of business value.

Exam trap

The trap here is mistaking a technical or governance activity for a business value definition, when value must be expressed in an operational outcome the business already tracks.

127
MCQeasy

A logistics firm's leadership wants a plain-language summary of how generative AI could reduce costs in dispatch operations before approving any budget. The team has no data scientists available. Which first step best aligns with a generative AI business strategy?

A.Ask each regional dispatch manager to independently experiment with public consumer chatbots and report anecdotes.
B.Commission a multi-year research program to train a custom dispatch model on historical route data.
C.Purchase GPU hardware for an on-premises cluster so the firm controls its own model hosting from day one.
D.Run a short, scoped pilot using a managed generative AI model on a sample of dispatch tasks and report measured outcomes.
AnswerD

A scoped pilot with a managed model produces concrete evidence, such as time saved per dispatch decision, without requiring data science hires or large capital. Because managed services remove infrastructure work, a small operations team can execute it. The measured results then give leadership the plain-language, cost-focused summary they requested before any budget commitment.

Why this answer

A short pilot on a managed generative AI model converts an abstract idea into measured dispatch outcomes, such as handling time or routing decision quality, using the team already in place. That evidence lets leadership evaluate cost impact in plain language before committing budget, which is the essence of a staged generative AI business strategy.

Exam trap

The trap here is equating credible generative AI adoption with acquiring hardware or launching research, when a small measured pilot is what actually informs an approval decision.

128
MCQmedium

A global news agency is using a generative AI model to summarize breaking news articles in real-time. The model is deployed on Vertex AI across multiple regions (us-central1, europe-west4, asia-southeast1) for low latency worldwide. The agency has a Service Level Objective (SLO) of 99.9% availability and p99 latency under 2 seconds. Recently, during a major event, traffic spiked 10x, and the europe-west4 region experienced latency spikes over 5 seconds and some 503 errors. The team suspects the regional endpoint is under-provisioned. Which combination of actions should they take to meet the SLO consistently?

A.Enable the global endpoint feature in Vertex AI with automatic traffic splitting, and increase the minimum replicas for each regional endpoint
B.Increase the maximum replicas for the europe-west4 endpoint and reduce the min replicas in other regions
C.Implement Cloud CDN caching for common summaries and reduce the number of regions to two
D.Configure a global load balancer with a single Vertex AI endpoint and increase max replicas globally
AnswerA

Global endpoint distributes traffic and increases capacity; higher min replicas prevent cold starts during spikes.

Why this answer

It enables the global endpoint feature with automatic traffic splitting, allowing traffic to be routed to healthy regions and providing failover. Additionally, increasing minimum replicas per region ensures each regional endpoint has baseline capacity to handle spikes, preventing under-provisioning. Option B only increases max replicas in europe-west4, which does not address traffic shifts, and reducing min replicas elsewhere risks capacity issues.

Option C suggests Cloud CDN, which is for static content, not model inference. Option D configures a global load balancer with a single endpoint, which does not optimally use Vertex AI's regional endpoints and may not meet latency SLO.

129
MCQhard

A large enterprise runs a generative AI solution serving millions of daily inference requests. To reduce costs, they propose using serverless endpoints (Vertex AI Prediction) with a custom container, but they notice high latency during cold starts. Which strategy best addresses this problem while minimizing cost?

A.Set a minimum number of replicas to maintain a baseline of always-on instances.
B.Upgrade to GPU-accelerated machines for all replicas.
C.Implement client-side request batching to reduce the number of inference calls.
D.Use prewarmed containers by setting an idle timeout to keep instances alive.
AnswerA

Correct. Setting a minimum number of replicas ensures a baseline of always-on instances, which eliminates cold starts for the majority of requests. This directly addresses the latency spike caused by container initialization and model loading, and the cost is limited to the minimum replicas rather than scaling all instances.

Why this answer

Setting a minimum number of replicas ensures that a baseline of always-on instances is maintained, eliminating cold starts for the majority of requests. This directly addresses the latency spike caused by container initialization and model loading in serverless endpoints, while the cost impact is limited to the minimum replicas rather than scaling all instances.

Exam trap

Google Cloud often tests the misconception that prewarming via idle timeout is a configurable parameter in serverless ML services, but in Vertex AI Prediction, the idle timeout is fixed and not user-adjustable, making minimum replicas the correct approach.

How to eliminate wrong answers

Option B is wrong because upgrading to GPU-accelerated machines increases cost significantly without solving cold start latency; GPUs primarily improve per-request throughput, not initialization time. Option C is wrong because client-side request batching reduces the number of inference calls but does not affect cold start latency; it may even increase perceived latency for individual requests. Option D is wrong because setting an idle timeout to keep instances alive is not a supported mechanism in Vertex AI Prediction; the service uses an internal keep-alive policy, and user-configurable idle timeouts are not available, making this option technically infeasible.

130
Multi-Selectmedium

A financial institution is implementing a generative AI chatbot to handle customer inquiries. The institution must comply with regulatory requirements (e.g., GDPR, SOX) and ensure data privacy. Which TWO actions should the institution take?

Select 2 answers
A.Establish a Center of Excellence (CoE) for AI governance to oversee model deployment and monitoring.
B.Use Vertex AI without additional data governance controls to simplify deployment.
C.Use a pre-trained model without customization to reduce development time.
D.Implement model validation and testing to ensure outputs meet regulatory standards.
E.Deploy the model on-premises only to keep data within local infrastructure.
AnswersA, D

A CoE centralises AI governance, giving the financial institution the oversight structure needed to enforce GDPR and SOX controls across model deployment and monitoring. It assigns accountability, standardises review gates and ensures privacy requirements are applied consistently.

Why this answer

Option A is correct because establishing a Center of Excellence (CoE) for AI governance provides the oversight, policy enforcement, and monitoring needed to keep a generative AI chatbot compliant with GDPR and SOX, covering model deployment, risk management, and accountability. Option D is correct because model validation and testing are essential to verify that chatbot outputs meet regulatory standards, detect bias or data leakage, and ensure ongoing compliance before and after deployment. Option B is incorrect because using Vertex AI without additional data governance controls fails to address GDPR and SOX privacy and audit requirements.

Option C is incorrect because a pre-trained model without customization does not by itself satisfy regulatory compliance or data privacy obligations. Option E is incorrect because on-premises deployment alone does not guarantee compliance and may not be feasible or sufficient for the institution's regulatory and operational needs.

Exam trap

The trap here is choosing 'on-premises only' as a silver bullet for data privacy — candidates conflate physical data location with regulatory compliance, but GDPR and SOX require governance, validation, and auditability regardless of where the model runs.

131
MCQhard

A company has been using an on-premises ML infrastructure for generative AI and wants to migrate to Google Cloud. They have a pipeline that fine-tunes a large language model weekly using a proprietary dataset. The migration must minimize downtime and data transfer costs. Which approach best addresses these requirements?

A.Use Vertex AI Pipelines to orchestrate the fine-tuning process, and use Vertex AI Managed Datasets to incrementally sync new data with BigQuery as the source.
B.Use AutoML to train a new model directly from the dataset without fine-tuning.
C.Deploy the existing pipeline on a Google Kubernetes Engine cluster and use Google Cloud Filestore for shared storage.
D.Use Cloud Storage Transfer Service to move all data to Cloud Storage, then set up a Vertex AI custom training job to run the fine-tuning.
AnswerA

Vertex AI Pipelines offers a managed orchestration service that can schedule the weekly fine-tuning workflow with minimal operational overhead. Combined with Vertex AI Managed Datasets, which incrementally sync new data from BigQuery, this approach reduces data transfer costs by avoiding full dataset copies and minimizes downtime by enabling automated, scheduled execution without manual intervention.

Why this answer

Vertex AI Pipelines provides a managed, serverless orchestration service that can run the weekly fine-tuning workflow with minimal operational overhead, while Vertex AI Managed Datasets can incrementally sync new data from BigQuery, reducing data transfer costs by avoiding full dataset copies. This combination minimizes downtime because the pipeline can be triggered on a schedule without manual intervention, and incremental syncs avoid re-transferring the entire proprietary dataset each week.

Exam trap

The trap here is that candidates often assume full data migration (e.g., Cloud Storage Transfer Service) is necessary, overlooking incremental sync capabilities of Vertex AI Managed Datasets with BigQuery, which directly addresses cost and downtime minimization.

How to eliminate wrong answers

Option B is wrong because AutoML is designed for training models from scratch on labeled data, not for fine-tuning an existing large language model, and it would require a full dataset transfer, increasing costs and downtime. Option C is wrong because deploying the existing pipeline on GKE with Filestore for shared storage does not address data transfer costs (Filestore still requires initial data migration) and introduces additional operational complexity for managing Kubernetes clusters, which does not minimize downtime compared to a managed service. Option D is wrong because using Cloud Storage Transfer Service to move all data to Cloud Storage incurs high initial data transfer costs and does not leverage incremental sync capabilities, and the custom training job setup lacks the orchestration and scheduling benefits of Vertex AI Pipelines, leading to more downtime during migration.

← PreviousPage 2 of 2 · 131 questions total

Ready to test yourself?

Try a timed practice session using only Business Strategies for Generative AI Solutions questions.