mediumMultiple Choice
Generative AI Leader Practice Question: An enterprise needs to generate natural-sounding…
An enterprise needs to generate natural-sounding speech from text for a voice assistant. They require low latency and support for custom voice models. Which service should they use?
⚠ Common exam trap
Many candidates confuse the Text-to-Speech API with the Speech-to-Text API, as candidates often mix up the direction of conversion (text-to-audio vs. audio-to-text) under time pressure.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Text-to-Speech API
The Text-to-Speech API (A) is correct because it is specifically designed to convert text into natural-sounding speech with low latency, and it supports custom voice models through features like Custom Voice and WaveNet voices. This directly meets the enterprise's requirements for a voice assistant that needs real-time, high-quality speech synthesis.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✓
Text-to-Speech API
Why this is correct
The Text-to-Speech API generates natural-sounding speech with low latency and supports custom voice models, matching the voice assistant's requirements exactly. Competing services lack either the custom voice training or the real-time synthesis latency the scenario demands.
- ✗
Cloud Translation API
Why it's wrong here
Cloud Translation API converts text between languages, returning text rather than synthesised audio, and cannot train custom voice models. It is tempting because it is a low-latency managed language service often paired with speech components; it would be correct for localising a voice assistant's responses, not for generating the speech itself.
- ✗
Vertex AI Text Generation
Why it's wrong here
Vertex AI Text Generation outputs written language tokens, producing no audio waveform and offering no voice-model customisation. It is tempting because Vertex AI hosts many generative models and supports tuning; it would be correct for chatbots, summarisation or content drafting, not for a voice assistant needing natural-sounding speech.
- ✗
Speech-to-Text API
Why it's wrong here
Speech-to-Text performs the inverse conversion, transcribing audio into text, so it cannot synthesise speech from a text prompt or host custom voice models. It is tempting because it is the sibling speech service and shares low-latency infrastructure; it would be correct for building transcription, captioning or voice-command input pipelines.
Go deeper
Related to this question
About these practice questions
One of 1,008 original Generative AI Leader practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →
JA
Written by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
This Generative AI Leader practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the Generative AI Leader exam.