Courseiva

Generative AI Leader Google Cloud's Generative AI Offerings Practice Question

A logistics company needs a generative AI model that can accept both text and images as input, and produce text output for describing shipping damage. They want to use a Google Cloud model through the Vertex AI API. Which Gemini model capability should they select?

⚠ Common exam trap

The trap here is assuming any Gemini model automatically handles images, when the input modality must be explicitly supported and configured for the chosen model.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Gemini 1.5 Pro with multimodal input support

A multimodal Gemini model such as Gemini 1.5 Pro can take both text and images and return text, directly satisfying the need to describe shipping damage from photos and written notes. The other services either restrict input to text, produce vectors instead of language, or focus on vision classification rather than generative text output.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    Gemini 1.5 Flash with text-only input

    Why it's wrong here

    Gemini 1.5 Flash is optimized for high-volume, lower-cost workloads, but configuring it for text-only input would prevent the model from analyzing damage photos. The scenario explicitly requires image input alongside text, so a text-only configuration cannot satisfy the core requirement even though the model family supports multimodal input in other configurations.

  • ✗

    Vertex AI Vision product for image classification

    Why it's wrong here

    Vertex AI Vision is designed for computer vision pipelines such as classification, detection, and tracking, not for open-ended text generation. It could label damage categories but would not produce the flexible written descriptions the logistics team wants, and it does not provide the conversational generative behavior expected from a Gemini model.

  • ✓

    Gemini 1.5 Pro with multimodal input support

    Why this is correct

    Gemini 1.5 Pro accepts interleaved text, image, and video input and returns text, which matches the requirement to describe shipping damage from photos plus written context. Calling it through the Vertex AI API gives the company managed access to this multimodal capability without hosting the model itself.

  • ✗

    Vertex AI Embeddings API for text and image vectors

    Why it's wrong here

    The Embeddings API converts content into numeric vectors for search, clustering, and similarity tasks. It does not generate natural-language descriptions of shipping damage. Using embeddings would require an additional model to turn vectors into text, adding complexity and not directly meeting the request for a generative model that produces text output.

About these practice questions

Courseiva writes every Generative AI Leader question from scratch — 1,008 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official Google Cloud exam blueprint

This Generative AI Leader practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the Generative AI Leader exam.