PMLE Architecting Low-Code ML Solutions Practice Question
A media company wants a low-code pipeline that ingests uploaded video files, detects scenes and on-screen text, and stores structured metadata for search. They prefer managed services and minimal custom code. Which TWO Google Cloud capabilities should they combine? (Choose two.)
⚠ Common exam trap
The trap here is substituting frame-by-frame Cloud Vision API calls for the video-native Video Intelligence API, which ignores the extra code required to sample frames and align timestamps.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Cloud Storage and Pub/Sub to trigger processing when new videos arrive
The Video Intelligence API supplies managed shot change and text detection on video, producing the scene and on-screen text metadata needed for search. Pairing it with Cloud Storage object notifications to Pub/Sub creates an event-driven trigger so each upload is analyzed automatically. Together they deliver a managed, low-code pipeline without custom frame extraction or model training.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Cloud Vision API for detecting labels and text in sampled video frames
Why it's wrong here
Cloud Vision API works on individual still images; using it on video requires you to extract frames and stitch timestamps yourself, adding custom code. That frame-sampling plumbing contradicts the low-code, managed pipeline goal and duplicates what the video-native API already handles.
- ✗
Vertex AI Vision for building a custom object-tracking pipeline with trained detectors
Why it's wrong here
Vertex AI Vision is aimed at streaming analytics and custom detector pipelines, which demands more configuration and model work than the low-code requirement allows. It also targets live streams and object tracking rather than the simple scene and text metadata extraction described.
- ✗
Cloud Data Fusion for building a visual ETL pipeline for video ingestion
Why it's wrong here
Cloud Data Fusion is a visual ETL tool for batch and streaming data records, not for analyzing video content such as scenes or on-screen text. It could move files but would not produce the semantic metadata, so it does not satisfy the detection requirement.
- ✓
Cloud Storage and Pub/Sub to trigger processing when new videos arrive
Why this is correct
A Cloud Storage bucket receiving uploads can publish object notifications to Pub/Sub, which then triggers the video analysis step. This serverless eventing is the standard low-code way to automate ingestion so each new video is processed and its metadata stored without manual intervention.
- ✓
Video Intelligence API for shot change detection and text detection
Why this is correct
The Video Intelligence API offers managed shot change detection and text detection on video, returning timestamps and recognized text without custom model code. It directly supplies the scene segmentation and on-screen text metadata the media company wants for search indexing.
Go deeper
Related to this question
About these practice questions
Courseiva writes every PMLE question from scratch — 775 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official Google Cloud exam blueprint
This PMLE practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the PMLE exam.