PMLE Architecting Low-Code ML Solutions Practice Question
A company wants to transcribe customer service calls in real-time to detect sentiment and identify urgent issues. They need a solution with low latency. Which combination of pre-built APIs should they use?
⚠ Common exam trap
PMLE often tests whether candidates confuse the direction of the audio APIs, so picking Text-to-Speech (which synthesizes speech) instead of Speech-to-Text (which transcribes) is the classic wrong-answer trap for transcription scenarios.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Speech-to-Text and Natural Language API
Real-time call transcription requires Speech-to-Text to convert audio to text with low latency, and Natural Language API to analyze that text for sentiment and entity/urgency detection. Together they form the standard streaming audio-analysis pipeline on Google Cloud. Speech-to-Text supports streaming recognition, and Natural Language provides sentiment and entity analysis on the resulting transcript.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Text-to-Speech and Natural Language API
Why it's wrong here
Text-to-Speech converts text into audio, the opposite direction from transcribing calls, and the Natural Language API analyses text sentiment but cannot process speech. It is tempting because both APIs handle language tasks, but they suit generating spoken output from written text, not real-time call transcription.
- ✗
Speech-to-Text and Translation API
Why it's wrong here
Speech-to-Text does transcribe calls, but the Translation API converts between languages, which the scenario never requires; sentiment detection needs the Natural Language API instead. It is tempting because both are speech and language APIs, yet translation suits multilingual content, not sentiment analysis of calls.
- ✗
Video Intelligence API
Why it's wrong here
Video Intelligence API analyses stored or streamed video for content, labels and shot detection; it does not transcribe audio at all, so real-time call sentiment is impossible. It would suit indexing video archives or moderating visual content, not speech-to-text on live calls.
- ✓
Speech-to-Text and Natural Language API
Why this is correct
Speech-to-Text streams audio into text with low latency, and the Natural Language API then performs sentiment analysis and entity detection on that text. Together they deliver real-time transcription plus sentiment and urgency identification without training custom models.
Go deeper
Related to this question
About these practice questions
Courseiva writes every PMLE question from scratch — 775 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official Google Cloud exam blueprint
This PMLE practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the PMLE exam.