Generative AI Leader Google Cloud's Generative AI Offerings Practice Question
A logistics company needs a generative AI model that can accept both text and images as input, and produce text output for describing shipping damage. They want to use a Google Cloud model through the Vertex AI API. Which Gemini model capability should they select?
⚠ Common exam trap
The trap here is assuming any Gemini model automatically handles images, when the input modality must be explicitly supported and configured for the chosen model.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Gemini 1.5 Pro with multimodal input support
A multimodal Gemini model such as Gemini 1.5 Pro can take both text and images and return text, directly satisfying the need to describe shipping damage from photos and written notes. The other services either restrict input to text, produce vectors instead of language, or focus on vision classification rather than generative text output.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Gemini 1.5 Flash with text-only input
Why it's wrong here
Gemini 1.5 Flash is optimized for high-volume, lower-cost workloads, but configuring it for text-only input would prevent the model from analyzing damage photos. The scenario explicitly requires image input alongside text, so a text-only configuration cannot satisfy the core requirement even though the model family supports multimodal input in other configurations.
- ✗
Vertex AI Vision product for image classification
Why it's wrong here
Vertex AI Vision is designed for computer vision pipelines such as classification, detection, and tracking, not for open-ended text generation. It could label damage categories but would not produce the flexible written descriptions the logistics team wants, and it does not provide the conversational generative behavior expected from a Gemini model.
- ✓
Gemini 1.5 Pro with multimodal input support
Why this is correct
Gemini 1.5 Pro accepts interleaved text, image, and video input and returns text, which matches the requirement to describe shipping damage from photos plus written context. Calling it through the Vertex AI API gives the company managed access to this multimodal capability without hosting the model itself.
- ✗
Vertex AI Embeddings API for text and image vectors
Why it's wrong here
The Embeddings API converts content into numeric vectors for search, clustering, and similarity tasks. It does not generate natural-language descriptions of shipping damage. Using embeddings would require an additional model to turn vectors into text, adding complexity and not directly meeting the request for a generative model that produces text output.
Go deeper
Related to this question
About these practice questions
Courseiva writes every Generative AI Leader question from scratch — 1,008 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official Google Cloud exam blueprint
This Generative AI Leader practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the Generative AI Leader exam.