mediumMultiple Choice
Generative AI Leader Practice Question: A media company wants to automatically generate…
A media company wants to automatically generate captions for video content in multiple languages. The captions should be synced with the audio timeline. Which combination of Google Cloud services is most appropriate?
⚠ Common exam trap
Candidates often mistakenly choose Cloud Video Intelligence API for captioning tasks because it can analyze video content, but caption generation requires audio transcription via Speech-to-Text API, not just visual analysis.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Speech-to-Text and Translation API
The workflow requires first transcribing the audio track into text using Speech-to-Text API, which provides timestamps for each word or phrase to ensure synchronization with the video timeline. The resulting transcript is then passed to the Translation API to generate captions in the target languages, preserving the original timing data for alignment.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Cloud Video Intelligence API and Translation API
Why it's wrong here
Cloud Video Intelligence API performs label, shot and explicit-content detection on video, not speech transcription, so no caption text or timeline alignment is produced. It would be correct for content classification, moderation or scene analysis rather than generating subtitles.
- ✗
Document AI and Translation API
Why it's wrong here
Document AI extracts structured data from scanned or digital documents, not audio streams, so it cannot transcribe speech or align captions to the video timeline. It would be correct for parsing invoices, forms or PDFs into structured fields.
- ✓
Speech-to-Text and Translation API
Why this is correct
Speech-to-Text transcribes the audio with timestamps, and the Translation API renders that text into other languages, satisfying the multilingual captioning requirement. The timestamps preserve audio-timeline sync, which a generic translation service alone could not provide.
- ✗
Text-to-Speech and Translation API
Why it's wrong here
Text-to-Speech synthesises spoken audio from text; it does not transcribe existing video audio, so no timeline-synced captions can be produced. It would be the correct choice for generating narration or accessibility voiceovers from a script, not for captioning recorded speech.
Visual reference
Go deeper
Related to this question
About these practice questions
One of 1,008 original Generative AI Leader practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →
JA
Written by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
This Generative AI Leader practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the Generative AI Leader exam.