Generative AI Leader Google Cloud's Generative AI Offerings Practice Question
Which TWO actions can reduce the cost of using Vertex AI Gemini API? (Choose two.)
⚠ Common exam trap
Candidates often mistakenly believe that increasing max output tokens or using a larger model improves quality without cost impact, but both directly increase token consumption and per-token pricing.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Use batch prediction instead of online
Option A is correct because batch prediction processes many requests asynchronously in a single job and is priced at a discount (typically 50%) compared to online prediction, directly lowering per-request cost for workloads that tolerate latency. Option E is correct because context caching lets you store frequently reused input tokens (e.g., long system prompts or documents) and pay a reduced rate for cached tokens plus a small storage fee, cutting costs when the same context is sent repeatedly. Option B is wrong because increasing max output tokens generates more billable output tokens, raising cost. Option C is wrong because grounding with Google Search adds grounding charges and does not reduce API cost. Option D is wrong because larger models have higher per-token prices, increasing rather than reducing cost.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✓
Use batch prediction instead of online
Why this is correct
Batch prediction processes many prompts asynchronously in a single job, avoiding the per-request overhead and premium pricing of synchronous online calls. This directly satisfies the stem's cost-reduction constraint, since Vertex AI charges less per token for batch workloads than for real-time online inference.
- ✗
Increase the max output tokens
Why it's wrong here
Raising max output tokens permits longer responses, and billing is per output token, so cost rises. It is tempting because a higher ceiling prevents truncation on long-form generation, but the correct lever is capping output length, not extending it.
- ✗
Use grounding with Google Search
Why it's wrong here
Grounding with Google Search adds retrieval calls and billed search queries, increasing rather than reducing cost. It is tempting because grounding improves factual accuracy and freshness for knowledge-intensive prompts, but that quality gain is paid for through extra tokens and per-query charges.
- ✗
Use a larger model
Why it's wrong here
A larger model raises cost per token, directly increasing spend rather than reducing it. Larger models are chosen when accuracy or reasoning depth on complex tasks matters more than budget, so this suits quality-critical workloads, not cost reduction.
- ✓
Use context caching
Why this is correct
Context caching stores frequently reused prompt prefixes so subsequent requests reuse the cached tokens rather than reprocessing them. Vertex AI charges reduced rates for cached input tokens, directly lowering per-call cost when prompts share large, repeated context such as system instructions or documents.
Go deeper
Related to this question
About these practice questions
Courseiva writes every Generative AI Leader question from scratch — 1,008 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →
JA
Written by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
This Generative AI Leader practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the Generative AI Leader exam.