mediumMultiple Choice
Generative AI Leader Practice Question: A developer wants to add real-time speech…
A developer wants to add real-time speech transcription to a customer call center application. They need low latency and high accuracy for multiple languages. Which Google AI API is most appropriate?
⚠ Common exam trap
The Generative AI Leader exam often tests the distinction between APIs that process text (Natural Language, Translation) versus those that process audio (Speech-to-Text, Text-to-Speech), and the trap here is confusing the direction of conversion (speech-to-text vs. text-to-speech) or assuming a translation API can handle raw audio input.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Speech-to-Text API
The Speech-to-Text API is the correct choice because it is specifically designed to convert audio into text in real time, supporting over 125 languages and variants with low-latency streaming. It offers features like automatic punctuation, speaker diarization, and domain-specific models (e.g., phone call) that directly meet the requirements of a customer call center application needing high accuracy across multiple languages.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✓
Speech-to-Text API
Why this is correct
Google Cloud Speech-to-Text API delivers streaming recognition with low latency, supporting real-time transcription of live audio. It covers over 125 languages and variants with high accuracy, satisfying the call centre's multilingual requirement. Its synchronous streaming mode returns interim results as speech occurs, which batch alternatives cannot match.
- ✗
Natural Language API
Why it's wrong here
Natural Language API analyses text for entities, sentiment and syntax; it accepts no audio input, so it cannot transcribe speech. Streaming multilingual transcription needs Speech-to-Text. Natural Language API would be correct when the requirement is extracting meaning or sentiment from text that has already been transcribed.
- ✗
Text-to-Speech API
Why it's wrong here
Text-to-Speech synthesises spoken audio from written text, the opposite direction to transcription, so it cannot convert caller speech into text at all. It is tempting because it also handles multiple languages and low-latency streaming, and would be the right choice when generating spoken prompts or IVR responses for callers.
- ✗
Translation API
Why it's wrong here
Translation API converts text between languages; it does not process audio or produce transcripts. Real-time multilingual speech transcription requires Speech-to-Text, which accepts streaming audio and returns text. Translation API would be correct when the input is already text and the goal is localisation into another language.
Go deeper
Related to this question
About these practice questions
Courseiva writes every Generative AI Leader question from scratch — 1,008 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →
JA
Written by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
This Generative AI Leader practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the Generative AI Leader exam.