hardMultiple Choice
PMLE Practice Question: A financial institution needs to deploy a fraud…
A financial institution needs to deploy a fraud detection model with strict latency <100ms per prediction and high throughput (1000 predictions/sec). The model is a deep neural network. Which architecture on Google Cloud meets these requirements?
⚠ Common exam trap
Google Cloud often tests the distinction between batch and online prediction services, where candidates mistakenly choose batch prediction for real-time requirements because they focus on throughput without considering latency constraints.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Use Vertex AI Prediction with autoscaling enabled and GPU machine types
Vertex AI Prediction with autoscaling and GPU machine types is correct because it provides low-latency online serving with autoscaling to handle high throughput (1000 predictions/sec) while keeping latency under 100ms. GPUs accelerate deep neural network inference, and autoscaling ensures resources match demand without over-provisioning.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Deploy the model on AI Platform Training with a single large VM
Why it's wrong here
AI Platform Training provisions resources for training jobs, not for serving online predictions, and a single VM cannot sustain 1000 predictions/sec with sub-100ms latency for a deep network. It is tempting because it runs on Google Cloud compute and can host custom code, but it lacks an inference endpoint.
- ✗
Deploy the model as a Cloud Function triggered by Cloud Pub/Sub
Why it's wrong here
Cloud Functions with Pub/Sub introduces cold starts and per-invocation overhead, and a deep neural network exceeds practical function size and timeout limits, so sub-100ms latency at 1000 predictions/sec is unattainable. It is tempting for event-driven, low-volume inference where asynchronous triggering suits sporadic workloads.
- ✗
Use Vertex AI Batch Prediction with a fixed number of machines
Why it's wrong here
Batch Prediction processes jobs asynchronously over stored input, returning results after completion, so it cannot serve per-prediction requests under 100ms. It is tempting for large offline scoring runs where throughput matters but latency does not, such as nightly fraud model evaluation over historical data.
- ✓
Use Vertex AI Prediction with autoscaling enabled and GPU machine types
Why this is correct
Vertex AI Prediction with autoscaling and GPU machine types provides the parallel compute needed for sub-100ms deep neural network inference at 1000 predictions/sec. GPUs accelerate matrix operations, while autoscaling adds nodes under load, satisfying both the latency and throughput constraints.
Go deeper
Related to this question
About these practice questions
Courseiva writes every PMLE question from scratch — 775 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →
JA
Written by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
This PMLE practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the PMLE exam.