Courseiva

Generative AI Leader Practice Question: Business Strategies for Generative AI Solutions

A media company wants to build a multi-modal generative app that accepts text, image, and video inputs and produces summaries. The app must handle variable-length videos up to 10 minutes. Which architecture is most scalable and cost-effective?

⚠ Common exam trap

A common misconception is that a single multi-modal endpoint is inherently scalable for any input size. In practice, directly ingesting raw, variable-length video without preprocessing (such as splitting into clips and extracting key frames) can cause high token costs and latency. A pipeline approach with context caching is more practical and cost-effective for variable-length multi-modal inputs.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Use a pipeline to split videos into short clips, extract key frames, and process with Gemini 1.5 Pro (with context caching) to generate summaries.

Option A is the most scalable and cost-effective because it preprocesses variable-length videos by splitting them into short clips and extracting key frames, reducing token usage and computational load. Gemini 1.5 Pro's context caching further improves efficiency by reusing processed context across requests. Option B loses visual context and adds an extra API, reducing accuracy and scalability. Option C discards multi-modal information entirely. Option D is less practical for variable-length videos because directly ingesting raw video into a single endpoint without preprocessing can lead to high token costs and latency, especially for videos up to 10 minutes.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✓

    Use a pipeline to split videos into short clips, extract key frames, and process with Gemini 1.5 Pro (with context caching) to generate summaries.

    Why this is correct

    Correct. Splitting videos into short clips and extracting key frames reduces computational load and token usage, while Gemini 1.5 Pro's context caching efficiently handles variable-length videos by reusing processed context across requests. This balance of scalability and cost-effectiveness is ideal for multi-modal summarization.

  • ✗

    Use Video Intelligence API to generate video captions, then feed captions to a text model.

    Why it's wrong here

    Captioning strips visual and temporal detail before summarisation, and Video Intelligence API captioning is billed per minute, so ten-minute videos inflate cost while losing frames. It would suit searchable transcripts, not multi-modal summarisation requiring native video understanding.

  • ✗

    Convert all inputs to text descriptions and use a text-only model.

    Why it's wrong here

    Converting images and video to text descriptions discards visual and temporal information before the model sees it, so summaries cannot reflect on-screen content. It would suit text-only pipelines, not a multi-modal app that must accept image and video inputs natively.

  • ✗

    Deploy a single Vertex AI endpoint with a model that can ingest multi-modal data directly.

    Why it's wrong here

    Incorrect. Deploying a single Vertex AI endpoint that ingests raw video directly would lead to prohibitive token costs and latency, especially for variable-length videos up to 10 minutes, making it impractical compared to a pipeline with preprocessing and caching.

About these practice questions

One of 1,008 original Generative AI Leader practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

This Generative AI Leader practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the Generative AI Leader exam.