easyMultiple Choice
PDE Practice Question: A startup is deploying a machine learning model…
A startup is deploying a machine learning model for real-time fraud detection. They need low latency and automatic scaling during peak hours. Which Google Cloud service should they use?
⚠ Common exam trap
It's easy for candidates to confuse Cloud Functions or Batch Prediction for real-time serving, overlooking that Vertex AI Endpoints are the only option purpose-built for low-latency, autoscaling online predictions in the modern Vertex AI ecosystem.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Vertex AI Endpoints
Vertex AI Endpoints provide managed, autoscaling infrastructure designed for low-latency online predictions, making them ideal for real-time fraud detection. They automatically scale the number of compute nodes based on incoming traffic, ensuring peak-hour demand is met without manual intervention.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Cloud Functions
Why it's wrong here
Cloud Functions runs event-driven code with cold starts and execution time limits, so it cannot host a model for sustained low-latency inference under peak load. It is tempting because it scales automatically, but it is designed for lightweight glue logic, not model serving.
- ✗
Batch Prediction on Vertex AI
Why it's wrong here
Batch Prediction processes jobs asynchronously against stored data, so it cannot return per-transaction scores within the milliseconds fraud detection demands. It is tempting because it is the right tool for scheduled bulk scoring of large datasets, where latency is irrelevant and throughput matters.
- ✗
Cloud AI Platform Prediction with custom containers
Why it's wrong here
AI Platform Prediction is the legacy precursor to Vertex AI; it lacks the current autoscaling and low-latency online serving features the scenario requires. It is tempting because custom containers genuinely suit models with bespoke dependencies, but that flexibility does not deliver the required real-time scaling here.
- ✓
Vertex AI Endpoints
Why this is correct
Vertex AI Endpoints serves models behind a managed, autoscaling HTTP interface, so replicas scale with traffic to hold latency low during peaks. This directly satisfies the real-time fraud detection requirement for low latency plus automatic scaling, unlike batch prediction or custom serving.
Go deeper
Related to this question
About these practice questions
One of 747 original PDE practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →
JA
Written by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
This PDE practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the PDE exam.