easyMultiple Choice
Generative AI Leader Practice Question: Which Google Cloud AI service would you use to…
Which Google Cloud AI service would you use to transcribe customer service call recordings into text for subsequent analysis?
⚠ Common exam trap
Watch out — candidates often confuse Speech-to-Text with Text-to-Speech or assuming that Translation API can handle audio input, when in fact it only works on text, leading candidates to pick a service that does not perform audio transcription.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Speech-to-Text
Speech-to-Text (STT) is the correct service because it is specifically designed to convert audio speech into written text using automatic speech recognition (ASR) models. For customer service call recordings, STT can handle domain-specific vocabulary, multiple speakers, and various audio formats, enabling downstream analysis like sentiment analysis or keyword extraction.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✓
Speech-to-Text
Why this is correct
Google Cloud Speech-to-Text converts audio into written text using automatic speech recognition, directly satisfying the requirement to transcribe call recordings. Unlike Text-to-Speech, which synthesises audio from text, it ingests audio and outputs transcripts, enabling the subsequent analysis step described in the scenario.
- ✗
Text-to-Speech
Why it's wrong here
Text-to-Speech synthesises spoken audio from written input, the reverse direction of this task, so it cannot produce a transcript from recordings. It is the right service when you already hold text and need natural-sounding speech output, such as voice responses in an IVR system.
- ✗
Translation API
Why it's wrong here
Translation API converts text between languages; it neither accepts audio input nor emits transcripts, so call recordings cannot be processed. It is correct when you already have text in one language and need it rendered in another, for example localising support articles.
- ✗
Document AI
Why it's wrong here
Document AI extracts structured fields from documents such as invoices and forms; it does not process audio streams, so recordings cannot be transcribed. It is the right choice when parsing scanned PDFs or images into structured data, not for speech-to-text conversion.
Go deeper
Related to this question
About these practice questions
One of 1,008 original Generative AI Leader practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →
JA
Written by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
This Generative AI Leader practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the Generative AI Leader exam.