An enterprise needs to generate natural-sounding speech from text for a voice assistant. They require low latency and support for custom voice models. Which service should they use?
The Text-to-Speech API generates natural-sounding speech with low latency and supports custom voice models, matching the voice assistant's requirements exactly. Competing services lack either the custom voice training or the real-time synthesis latency the scenario demands.
Why this answer
The Text-to-Speech API (A) is correct because it is specifically designed to convert text into natural-sounding speech with low latency, and it supports custom voice models through features like Custom Voice and WaveNet voices. This directly meets the enterprise's requirements for a voice assistant that needs real-time, high-quality speech synthesis.
Exam trap
The trap here is confusing the Text-to-Speech API with the Speech-to-Text API, as candidates often mix up the direction of conversion (text-to-audio vs. audio-to-text) under time pressure.
How to eliminate wrong answers
Option B (Cloud Translation API) is wrong because it translates text between languages, not text to speech, and does not generate audio output. Option C (Vertex AI Text Generation) is wrong because it generates text content (e.g., chat responses, summaries) rather than synthesizing speech from text. Option D (Speech-to-Text API) is wrong because it performs the inverse operation—converting audio speech into text—and does not produce speech output.