Courseiva
mediumMultiple Choice

Generative AI Leader Practice Question: A developer wants to add real-time speech…

A developer wants to add real-time speech transcription to a customer call center application. They need low latency and high accuracy for multiple languages. Which Google AI API is most appropriate?

⚠ Common exam trap

The Generative AI Leader exam often tests the distinction between APIs that process text (Natural Language, Translation) versus those that process audio (Speech-to-Text, Text-to-Speech), and the trap here is confusing the direction of conversion (speech-to-text vs. text-to-speech) or assuming a translation API can handle raw audio input.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Speech-to-Text API

The Speech-to-Text API is the correct choice because it is specifically designed to convert audio into text in real time, supporting over 125 languages and variants with low-latency streaming. It offers features like automatic punctuation, speaker diarization, and domain-specific models (e.g., phone call) that directly meet the requirements of a customer call center application needing high accuracy across multiple languages.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✓

    Speech-to-Text API

    Why this is correct

    Google Cloud Speech-to-Text API delivers streaming recognition with low latency, supporting real-time transcription of live audio. It covers over 125 languages and variants with high accuracy, satisfying the call centre's multilingual requirement. Its synchronous streaming mode returns interim results as speech occurs, which batch alternatives cannot match.

  • ✗

    Natural Language API

    Why it's wrong here

    Natural Language API analyses text for entities, sentiment and syntax; it accepts no audio input, so it cannot transcribe speech. Streaming multilingual transcription needs Speech-to-Text. Natural Language API would be correct when the requirement is extracting meaning or sentiment from text that has already been transcribed.

  • ✗

    Text-to-Speech API

    Why it's wrong here

    Text-to-Speech synthesises spoken audio from written text, the opposite direction to transcription, so it cannot convert caller speech into text at all. It is tempting because it also handles multiple languages and low-latency streaming, and would be the right choice when generating spoken prompts or IVR responses for callers.

  • ✗

    Translation API

    Why it's wrong here

    Translation API converts text between languages; it does not process audio or produce transcripts. Real-time multilingual speech transcription requires Speech-to-Text, which accepts streaming audio and returns text. Translation API would be correct when the input is already text and the goal is localisation into another language.

About these practice questions

Courseiva writes every Generative AI Leader question from scratch — 1,008 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

This Generative AI Leader practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the Generative AI Leader exam.