Courseiva

PMLE Architecting Low-Code ML Solutions Practice Question

A company wants to transcribe customer service calls in real-time to detect sentiment and identify urgent issues. They need a solution with low latency. Which combination of pre-built APIs should they use?

⚠ Common exam trap

PMLE often tests whether candidates confuse the direction of the audio APIs, so picking Text-to-Speech (which synthesizes speech) instead of Speech-to-Text (which transcribes) is the classic wrong-answer trap for transcription scenarios.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Speech-to-Text and Natural Language API

Real-time call transcription requires Speech-to-Text to convert audio to text with low latency, and Natural Language API to analyze that text for sentiment and entity/urgency detection. Together they form the standard streaming audio-analysis pipeline on Google Cloud. Speech-to-Text supports streaming recognition, and Natural Language provides sentiment and entity analysis on the resulting transcript.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    Text-to-Speech and Natural Language API

    Why it's wrong here

    Text-to-Speech converts text into audio, the opposite direction from transcribing calls, and the Natural Language API analyses text sentiment but cannot process speech. It is tempting because both APIs handle language tasks, but they suit generating spoken output from written text, not real-time call transcription.

  • ✗

    Speech-to-Text and Translation API

    Why it's wrong here

    Speech-to-Text does transcribe calls, but the Translation API converts between languages, which the scenario never requires; sentiment detection needs the Natural Language API instead. It is tempting because both are speech and language APIs, yet translation suits multilingual content, not sentiment analysis of calls.

  • ✗

    Video Intelligence API

    Why it's wrong here

    Video Intelligence API analyses stored or streamed video for content, labels and shot detection; it does not transcribe audio at all, so real-time call sentiment is impossible. It would suit indexing video archives or moderating visual content, not speech-to-text on live calls.

  • ✓

    Speech-to-Text and Natural Language API

    Why this is correct

    Speech-to-Text streams audio into text with low latency, and the Natural Language API then performs sentiment analysis and entity detection on that text. Together they deliver real-time transcription plus sentiment and urgency identification without training custom models.

About these practice questions

Courseiva writes every PMLE question from scratch — 775 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official Google Cloud exam blueprint

This PMLE practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the PMLE exam.