Courseiva
Business Strategies for Generative AI SolutionshardMultiple SelectObjective-mapped

Generative AI Leader Practice Question: Business Strategies for Generative AI Solutions

Which TWO strategies can effectively reduce the operational costs of a generative AI model in production without significantly degrading user experience?

⚠ Common exam trap

Google Cloud often tests the misconception that increasing batch sizes or retraining frequency inherently reduces costs, when in fact these actions typically increase resource usage or introduce operational overhead without guaranteeing cost savings.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

Cache frequent prompt completions

Caching frequent prompt completions reduces operational costs by eliminating redundant inference calls for identical or similar user requests. This directly lowers compute usage and latency without degrading user experience, as cached responses are served instantly. It is a common optimization in production LLM deployments, especially for high-traffic applications with repetitive queries.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • Use larger batch sizes for inference

    Why it's wrong here

    Batching is not always applicable for real-time responses and may increase latency.

  • Increase the frequency of model retraining to improve efficiency

    Why it's wrong here

    Retraining costs money and may not reduce inference cost.

  • Cache frequent prompt completions

    Why this is correct

    Caching reduces duplicate inference calls, lowering cost.

  • Adopt a pay-per-use pricing model instead of a flat rate

    Why this is correct

    Pay-per-use ensures you only pay for actual usage.

  • Deploy multiple models and route requests by complexity

    Why it's wrong here

    Managing multiple models increases operational overhead.

About these practice questions

This Generative AI Leader question is part of Courseiva's 683-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

This Generative AI Leader practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the Generative AI Leader exam.