Courseiva
mediumMultiple Choice

Generative AI Leader Practice Question: A media company wants to automatically generate…

A media company wants to automatically generate captions for video content in multiple languages. The captions should be synced with the audio timeline. Which combination of Google Cloud services is most appropriate?

⚠ Common exam trap

Candidates often mistakenly choose Cloud Video Intelligence API for captioning tasks because it can analyze video content, but caption generation requires audio transcription via Speech-to-Text API, not just visual analysis.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Speech-to-Text and Translation API

The workflow requires first transcribing the audio track into text using Speech-to-Text API, which provides timestamps for each word or phrase to ensure synchronization with the video timeline. The resulting transcript is then passed to the Translation API to generate captions in the target languages, preserving the original timing data for alignment.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    Cloud Video Intelligence API and Translation API

    Why it's wrong here

    Cloud Video Intelligence API performs label, shot and explicit-content detection on video, not speech transcription, so no caption text or timeline alignment is produced. It would be correct for content classification, moderation or scene analysis rather than generating subtitles.

  • ✗

    Document AI and Translation API

    Why it's wrong here

    Document AI extracts structured data from scanned or digital documents, not audio streams, so it cannot transcribe speech or align captions to the video timeline. It would be correct for parsing invoices, forms or PDFs into structured fields.

  • ✓

    Speech-to-Text and Translation API

    Why this is correct

    Speech-to-Text transcribes the audio with timestamps, and the Translation API renders that text into other languages, satisfying the multilingual captioning requirement. The timestamps preserve audio-timeline sync, which a generic translation service alone could not provide.

  • ✗

    Text-to-Speech and Translation API

    Why it's wrong here

    Text-to-Speech synthesises spoken audio from text; it does not transcribe existing video audio, so no timeline-synced captions can be produced. It would be the correct choice for generating narration or accessibility voiceovers from a script, not for captioning recorded speech.

Visual reference

Client Server SYN (seq=100) SYN-ACK (seq=200, ack=101) ACK (ack=201) Connection established — data transfer begins

About these practice questions

One of 1,008 original Generative AI Leader practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

This Generative AI Leader practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the Generative AI Leader exam.