PMLE Architecting Low-Code ML Solutions Practice Question
A developer needs to transcribe phone calls with high accuracy for a call center analytics application. The audio is in English and has background noise. Which Speech-to-Text model should they choose?
⚠ Common exam trap
PMLE often tests the misconception that any Speech-to-Text model can be used interchangeably, but the trap is failing to recognize that telephony audio requires a specialized model due to its unique acoustic properties.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
telephony
The 'telephony' model is specifically optimized for transcribing audio from phone calls, which often includes background noise, low fidelity, and narrowband audio. It is designed to handle the unique characteristics of telephony audio, such as 8kHz sampling rate and compression artifacts, providing higher accuracy for call center analytics. Other models like 'latest_short' and 'latest_long' are general-purpose and may not perform as well on noisy phone calls.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✓
telephony
Why this is correct
The telephony model is trained on 8 kHz narrowband audio, matching the bandwidth of phone calls, and is optimised to suppress background noise. This directly satisfies the stem's call centre constraint: English telephone speech with noise, where the default model's wideband training would degrade accuracy.
- ✗
latest_short
Why it's wrong here
latest_short targets short utterances like commands or voice search, not extended two-party conversations. Its design assumes brief, clean speech, so continuous noisy phone calls fall outside its intended use; the telephony model covers that scenario.
- ✗
latest_long
Why it's wrong here
latest_long suits long-form audio such as voicemail or dictated documents, but it is not tuned for telephony bandwidth or noisy call-centre audio. The telephony model exists precisely for phone-call transcription, so latest_long misses the required acoustic domain.
- ✗
Any model; they are all equivalent
Why it's wrong here
Claiming all models perform identically ignores that Azure's Speech-to-Text offerings differ by training data and noise handling; the base model degrades on noisy call-centre audio, whereas the custom or enhanced models target exactly this. It is tempting when treating transcription as a commodity, but equivalence would only hold for clean, single-speaker, standard-vocabulary recordings.
Go deeper
Related to this question
About these practice questions
Courseiva writes every PMLE question from scratch — 775 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →
Same concept, more angles
5 more ways this is tested on PMLE
These questions test the same concept from different angles. Work through them to make sure you can recognise it however the exam phrases it.
Variation 1. A media company wants to transcribe audio files from customer support calls into text for analysis. The audio is in English with clear speech and no background noise. They want a quick solution with no ML model training. Which Google Cloud service should they use?
easy- A.Translation API to translate the audio
- B.AutoML NLP to train a transcription model
- C.Vertex AI Workbench to train a custom speech recognition model
- ✓ D.Speech-to-Text API with the latest_long model
Why D: The Speech-to-Text API is Google Cloud's fully managed, pre-trained automatic speech recognition (ASR) service that requires no model training. The latest_long model is specifically optimized for transcribing long-form audio content such as customer support calls, providing high accuracy for clear English speech. Since the audio is clear with no background noise and the company wants a quick, training-free solution, Speech-to-Text API with latest_long is the ideal fit.
Variation 2. A company wants to analyze customer reviews for sentiment (positive, negative, neutral) using a pre-trained model with no training. They have text data stored in BigQuery. Which Google Cloud service should they use?
medium- A.Translation API
- B.Speech-to-Text API
- ✓ C.Natural Language API
- D.AutoML NLP
Why C: The Natural Language API is a pre-trained service that provides sentiment analysis out of the box, requiring no training. It can directly process text data from BigQuery (via integration or export) to classify sentiment as positive, negative, or neutral. This matches the requirement of using a pre-trained model with no training.
Variation 3. A company wants to transcribe audio from customer service calls and then analyze the sentiment of the transcribed text. Which TWO Google Cloud services should they use?
easy- ✓ A.Natural Language API
- B.Document AI
- ✓ C.Speech-to-Text
- D.Translation API
- E.Vision API
Why A: Speech-to-Text (option C) is correct because it converts audio from the customer service calls into written text, which is the required transcription step. Natural Language API (option A) is correct because it performs sentiment analysis on text, allowing the company to analyze the sentiment of the transcribed call content. Document AI (option B) is not appropriate here because it processes documents and forms rather than audio or general sentiment analysis. Translation API (option D) only translates text between languages and does not transcribe audio or analyze sentiment. Vision API (option E) analyzes images, so it cannot handle audio transcription or text sentiment analysis.
Variation 4. A company wants to transcribe customer service calls in real-time to detect sentiment and identify urgent issues. They need a solution with low latency. Which combination of pre-built APIs should they use?
easy- A.Text-to-Speech and Natural Language API
- B.Speech-to-Text and Translation API
- C.Video Intelligence API
- ✓ D.Speech-to-Text and Natural Language API
Why D: Real-time call transcription requires Speech-to-Text to convert audio to text with low latency, and Natural Language API to analyze that text for sentiment and entity/urgency detection. Together they form the standard streaming audio-analysis pipeline on Google Cloud. Speech-to-Text supports streaming recognition, and Natural Language provides sentiment and entity analysis on the resulting transcript.
Variation 5. A developer wants to add text translation to a mobile app. They need to translate user-generated content into multiple languages, and latency is critical. Which pre-built API should they use?
easy- ✓ A.Translation API
- B.Vision API
- C.Text-to-Speech API
- D.Natural Language API
Why A: Translation API provides fast, real-time translation for text. Natural Language API is for analysis. Text-to-Speech is for audio. Vision API is for images.
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official Google Cloud exam blueprint
This PMLE practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the PMLE exam.