hardMultiple Select
PMLE Practice Question: An ML engineer is building a monitoring dashboard…
An ML engineer is building a monitoring dashboard for a Vertex AI pipeline that includes training, evaluation, and batch prediction. Which THREE components should be included to provide comprehensive observability? (Select THREE.)
⚠ Common exam trap
It's easy for candidates to confuse infrastructure monitoring (CPU/memory logs) or serving-layer metrics (online prediction latency) with pipeline-specific observability, leading them to select options that are relevant to different stages of the ML lifecycle rather than the pipeline itself.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Pipeline execution status, duration, and failure rates for each component.
Option A is correct because tracking pipeline execution status, duration, and failure rates per component gives the operational observability needed to detect where a training, evaluation, or batch prediction step fails or slows down. Option C is correct because model evaluation metrics such as accuracy and AUC after training and validation are essential to monitor model quality and detect degradation before deployment. Option D is correct because data validation reports with anomaly counts and feature statistics surface input data drift, skew, and schema issues that directly affect pipeline and model reliability. Option B is not appropriate because raw Compute Engine CPU and memory logs for each step are low-level infrastructure telemetry, not the pipeline-level observability components the scenario requires. Option E is not appropriate because online prediction latency and request counts apply to a deployed endpoint, whereas this pipeline uses batch prediction and has no online serving endpoint.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✓
Pipeline execution status, duration, and failure rates for each component.
Why this is correct
Pipeline execution status, duration, and failure rates expose orchestration-level telemetry for each training, evaluation, and batch prediction step, satisfying the requirement for comprehensive observability across the whole Vertex AI pipeline. These metrics reveal where runs stall or fail, which resource-level metrics alone cannot attribute to a specific component.
- ✗
Compute engine CPU and memory logs for each pipeline step.
Why it's wrong here
Pipeline-step CPU and memory metrics describe the underlying Compute Engine VMs, not the ML workflow itself. Vertex AI Pipelines exposes task-level execution metadata, artefacts, and lineage; the dashboard needs those pipeline-run signals. Host metrics would be the right choice when diagnosing infrastructure saturation on self-managed training VMs.
- ✓
Model evaluation metrics (e.g., accuracy, AUC) after training and validation.
Why this is correct
Model evaluation metrics capture predictive quality after training and validation, satisfying the pipeline's evaluation stage. Vertex AI writes these to Vertex ML Metadata, letting the dashboard track accuracy and AUC drift across runs. Without them, observability covers infrastructure and data but not model behaviour, leaving a monitoring gap the stem explicitly requires.
- ✓
Data validation reports showing anomaly counts and feature statistics.
Why this is correct
Data validation reports expose input drift and schema violations before they corrupt downstream training or predictions, satisfying the pipeline-wide observability requirement. Vertex AI's built-in validation surfaces anomaly counts and feature statistics per split, letting the dashboard flag skew between training and serving data — a gap that resource metrics alone cannot detect.
- ✗
Online prediction latency and request count from the deployed model endpoint.
Why it's wrong here
The pipeline covers training, evaluation, and batch prediction, so endpoint request metrics sit outside its scope — no deployed online endpoint is involved. Online latency and request counts belong on a dashboard for a real-time prediction service, where traffic volume and tail latency drive autoscaling and SLO alerting.
Go deeper
Related to this question
About these practice questions
This PMLE question is part of Courseiva's 775-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →
JA
Written by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
This PMLE practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the PMLE exam.