easyMultiple Choice
PMLE Practice Question: A company deploys an online prediction model…
A company deploys an online prediction model serving 100 requests per second. They are optimizing for both latency and throughput. Which monitoring strategy should they use?
⚠ Common exam trap
The trap here is that candidates often focus on a single metric (e.g., error rate or p99 latency) and overlook the need for multi-metric correlation, especially the latency-throughput trade-off, which is a core concept in monitoring ML systems under production load.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Monitor both the p50 and p99 latency, and the request count. Create a dashboard showing latency vs. throughput at different load levels.
Monitoring both p50 and p99 latency alongside request count provides a comprehensive view of system performance under load. Latency percentiles reveal tail behavior (p99) and typical user experience (p50), while request count tracks throughput. A dashboard correlating latency vs. throughput at different load levels is essential for identifying performance cliffs or degradation before failures occur, aligning with best practices for production ML inference systems.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Monitor only the request count and set an alert if it drops below a threshold.
Why it's wrong here
Request count alone reveals traffic volume, saying nothing about how long each prediction takes or how many complete per second, so latency regressions stay invisible. It tempts because volume monitoring is the standard first signal for capacity planning and detecting traffic drops or outages.
- ✗
Set a single alert on the 99th percentile latency and ignore throughput since it's already high.
Why it's wrong here
Alerting on p99 latency while dismissing throughput ignores saturation: as load rises, throughput plateaus and tail latency climbs, so both must be tracked together. It tempts because tail latency is the standard user-experience metric, and it suffices when throughput headroom is genuinely guaranteed.
- ✗
Monitor the error rate and set an alert if it exceeds 1%.
Why it's wrong here
Error rate captures failed or malformed responses only; a model answering every request correctly but slowly leaves latency and throughput unmeasured. It tempts because error-rate alerting is the conventional reliability baseline, and it is the right primary signal when correctness, not performance, is the service objective.
- ✓
Monitor both the p50 and p99 latency, and the request count. Create a dashboard showing latency vs. throughput at different load levels.
Why this is correct
Monitoring p50 and p99 latency alongside request count directly satisfies the latency-and-throughput constraint at 100 requests per second. Percentiles expose tail latency that averages hide, while correlating both metrics against load levels reveals saturation points, letting the team tune the model for throughput without breaching latency targets.
Go deeper
Related to this question
About these practice questions
Courseiva writes every PMLE question from scratch — 775 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →
JA
Written by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
This PMLE practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the PMLE exam.