Courseiva

Generative AI Leader Practice Question: Business Strategies for Generative AI Solutions

A company wants to use Generative AI for customer support chatbots. They are concerned about cost and latency. Which deployment option best balances these concerns?

⚠ Common exam trap

Google Cloud often tests the misconception that 'larger model = better accuracy always' or that 'on-premise is always cheaper,' ignoring the total cost of ownership, scaling overhead, and the efficiency gains from fine-tuning and caching for specific use cases.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Use a fine-tuned version of a smaller model on Vertex AI with response caching

Using a fine-tuned smaller model on Vertex AI with response caching reduces both cost and latency. Smaller models require fewer computational resources, and caching avoids redundant inference calls, directly addressing the company's concerns without sacrificing accuracy for the specific task.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    Deploy an open-source model on-premise to avoid cloud costs

    Why it's wrong here

    On-premise open-source deployment removes cloud API fees but requires GPU hardware, operations and scaling capacity, so cost shifts rather than balances, and latency depends on local infrastructure. It is tempting because it avoids per-call charges, and would suit strict data-residency or offline requirements with existing hardware.

  • ✗

    Rely on a third-party chatbot API that abstracts the model

    Why it's wrong here

    A third-party API abstracting the model leaves cost and latency governed by the vendor's routing and pricing, giving no control over model size or hosting to balance the two concerns. It is tempting because it removes operational overhead, and would suit teams lacking ML engineering capacity who accept the provider's trade-offs.

  • ✗

    Use the largest available foundation model via API for highest accuracy

    Why it's wrong here

    The largest foundation model via API maximises per-token cost and inference latency, directly contradicting the stated cost and latency concerns. It is tempting because such models give the highest accuracy on complex reasoning, and would be the right pick when answer quality outweighs budget and response-time constraints.

  • ✓

    Use a fine-tuned version of a smaller model on Vertex AI with response caching

    Why this is correct

    A fine-tuned smaller model cuts inference cost and latency versus a large general model, while response caching avoids repeated generation for common queries. Running on Vertex AI provides managed scaling, together balancing the stated cost and latency concerns for the chatbot.

About these practice questions

One of 1,008 original Generative AI Leader practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

This Generative AI Leader practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the Generative AI Leader exam.