easyMultiple Choice
Generative AI Leader Practice Question: Which Google Cloud AI API is used to convert…
Which Google Cloud AI API is used to convert spoken language into text?
⚠ Common exam trap
Test-takers frequently confuse the direction of conversion — candidates often mix up Speech-to-Text (audio → text) with Text-to-Speech (text → audio), so reading the question's direction carefully is essential.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Speech-to-Text API
The Speech-to-Text API is Google Cloud's dedicated service for automatic speech recognition (ASR), converting audio input into written text using models like Chirp and the Universal speech model. It accepts audio from files or streaming sources and returns transcriptions with optional timestamps, speaker diarization, and word-level confidence. This directly matches the requirement of turning spoken language into text.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✓
Speech-to-Text API
Why this is correct
The Speech-to-Text API performs automatic speech recognition, transcribing audio input into written text. It directly satisfies the stem's requirement to convert spoken language into text, unlike Vision, Translation or Natural Language APIs, which handle images, language translation and text analysis respectively.
- ✗
Text-to-Speech API
Why it's wrong here
The Text-to-Speech API performs the inverse conversion, synthesising audio from written text, so it cannot transcribe speech. It is tempting because it sits in the same speech domain, and would be correct if the requirement were generating spoken audio output from text, such as reading responses aloud.
- ✗
Natural Language API
Why it's wrong here
The Natural Language API performs entity, sentiment and syntax analysis on existing text; it cannot accept audio input at all. It is tempting because it handles text processing, and would be correct if the requirement were extracting entities or sentiment from written records rather than transcribing speech.
- ✗
Translation API
Why it's wrong here
The Translation API converts text between languages, taking text input and returning text output, so it never processes audio. It is tempting because it also handles language data, and would be the right choice if the task were translating already-transcribed text into another language.
Go deeper
Related to this question
About these practice questions
Courseiva writes every Generative AI Leader question from scratch — 1,008 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official Google Cloud exam blueprint
This Generative AI Leader practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the Generative AI Leader exam.