Courseiva
Business Strategies for Generative AI SolutionseasyMultiple SelectObjective-mapped

Cost Optimization Strategies for Generative AI

A company is using Vertex AI generative models for a high-volume text summarization service. Which two strategies can reduce operational costs?

Quick Answer

The answer is to reduce the model size and use batch prediction. Switching from a larger model like text-bison@002 to a lighter variant such as text-bison-light directly lowers the per-token compute cost, while batch prediction processes multiple requests in a single job, eliminating the per-request overhead and idle compute time of online inference. This combination is ideal for high-volume, non-real-time workloads like text summarization, where you pay only for the batch job’s compute duration rather than individual predictions. On the Google Cloud Generative AI Leader exam, this question tests your understanding of cost optimization for generative AI by distinguishing between online and batch inference patterns—a common trap is assuming that only model downsizing matters. Remember the mnemonic “Lighter and Together”: use a smaller model and batch your requests to cut costs.

⚠ Common exam trap

Google Cloud often tests the misconception that adjusting inference parameters like temperature or output length can reduce costs, when in reality only reducing model size or switching to batch processing directly lowers operational expenses.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

Use batch prediction instead of online prediction.

Batch prediction reduces costs by processing multiple requests in a single batch job, which avoids the per-request overhead and idle compute time associated with online prediction. This is especially cost-effective for high-volume, non-real-time workloads like text summarization, as you pay only for the compute time used during the batch job rather than for each individual inference.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • Increase the model's max output tokens to 2048.

    Why it's wrong here

    Increasing output tokens increases cost per request.

  • Implement retry logic with exponential backoff.

    Why it's wrong here

    Retries increase the number of requests, raising costs.

  • Lower the temperature parameter to 0.

    Why it's wrong here

    Temperature does not affect cost; it only changes output randomness.

  • Use batch prediction instead of online prediction.

    Why this is correct

    Batch prediction has lower per-request cost for large jobs compared to online prediction.

  • Reduce the size of the model (e.g., switch from text-bison@002 to text-bison-light).

    Why this is correct

    Smaller models have lower cost per token, reducing operational costs.

About these practice questions

Courseiva writes every Generative AI Leader question from scratch — 683 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

Same concept, more angles

2 more ways this is tested on Generative AI Leader

These questions test the same concept from different angles. Work through them to make sure you can recognise it however the exam phrases it.

Variation 1. A company is using a generative AI model for internal report generation. They notice costs are high because each request processes large amounts of text. Which business strategy would most effectively reduce costs while maintaining quality?

medium
  • A.Fine-tune a smaller model on a specialized dataset.
  • B.Use a more powerful model to reduce retries.
  • C.Implement caching for repeated requests.
  • D.Increase the batch size for online predictions.

Why A: Fine-tuning a smaller model on a specialized dataset reduces computational cost per inference because smaller models have fewer parameters and require less memory and processing power. By tailoring the model to the company's specific domain (e.g., internal reports), it can maintain output quality comparable to a larger general-purpose model, directly addressing the cost-per-request issue without sacrificing accuracy.

Variation 2. Which of the following is a key consideration when selecting a GenAI model for a cost-sensitive application?

easy
  • A.Model size in parameters
  • B.Latency and throughput requirements
  • C.Number of training epochs
  • D.The model's training data source

Why B: For cost-sensitive applications, latency and throughput requirements directly impact infrastructure costs, as lower latency often requires more expensive compute resources (e.g., higher GPU memory, faster inference hardware) and higher throughput may necessitate scaling out instances. Model size in parameters is a secondary factor that influences latency and throughput, but the primary cost driver is the operational performance needed to meet service-level agreements (SLAs).

JA

Written by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

This Generative AI Leader practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the Generative AI Leader exam.