Courseiva
mediumMultiple Choice

Generative AI Leader Practice Question: An enterprise needs to generate natural-sounding…

An enterprise needs to generate natural-sounding speech from text for a voice assistant. They require low latency and support for custom voice models. Which service should they use?

⚠ Common exam trap

Many candidates confuse the Text-to-Speech API with the Speech-to-Text API, as candidates often mix up the direction of conversion (text-to-audio vs. audio-to-text) under time pressure.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Text-to-Speech API

The Text-to-Speech API (A) is correct because it is specifically designed to convert text into natural-sounding speech with low latency, and it supports custom voice models through features like Custom Voice and WaveNet voices. This directly meets the enterprise's requirements for a voice assistant that needs real-time, high-quality speech synthesis.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✓

    Text-to-Speech API

    Why this is correct

    The Text-to-Speech API generates natural-sounding speech with low latency and supports custom voice models, matching the voice assistant's requirements exactly. Competing services lack either the custom voice training or the real-time synthesis latency the scenario demands.

  • ✗

    Cloud Translation API

    Why it's wrong here

    Cloud Translation API converts text between languages, returning text rather than synthesised audio, and cannot train custom voice models. It is tempting because it is a low-latency managed language service often paired with speech components; it would be correct for localising a voice assistant's responses, not for generating the speech itself.

  • ✗

    Vertex AI Text Generation

    Why it's wrong here

    Vertex AI Text Generation outputs written language tokens, producing no audio waveform and offering no voice-model customisation. It is tempting because Vertex AI hosts many generative models and supports tuning; it would be correct for chatbots, summarisation or content drafting, not for a voice assistant needing natural-sounding speech.

  • ✗

    Speech-to-Text API

    Why it's wrong here

    Speech-to-Text performs the inverse conversion, transcribing audio into text, so it cannot synthesise speech from a text prompt or host custom voice models. It is tempting because it is the sibling speech service and shares low-latency infrastructure; it would be correct for building transcription, captioning or voice-command input pipelines.

About these practice questions

One of 1,008 original Generative AI Leader practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

This Generative AI Leader practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the Generative AI Leader exam.