Courseiva
hardMultiple Choice

PMLE Practice Question: A financial institution needs to deploy a fraud…

A financial institution needs to deploy a fraud detection model with strict latency <100ms per prediction and high throughput (1000 predictions/sec). The model is a deep neural network. Which architecture on Google Cloud meets these requirements?

⚠ Common exam trap

Google Cloud often tests the distinction between batch and online prediction services, where candidates mistakenly choose batch prediction for real-time requirements because they focus on throughput without considering latency constraints.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Use Vertex AI Prediction with autoscaling enabled and GPU machine types

Vertex AI Prediction with autoscaling and GPU machine types is correct because it provides low-latency online serving with autoscaling to handle high throughput (1000 predictions/sec) while keeping latency under 100ms. GPUs accelerate deep neural network inference, and autoscaling ensures resources match demand without over-provisioning.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    Deploy the model on AI Platform Training with a single large VM

    Why it's wrong here

    AI Platform Training provisions resources for training jobs, not for serving online predictions, and a single VM cannot sustain 1000 predictions/sec with sub-100ms latency for a deep network. It is tempting because it runs on Google Cloud compute and can host custom code, but it lacks an inference endpoint.

  • ✗

    Deploy the model as a Cloud Function triggered by Cloud Pub/Sub

    Why it's wrong here

    Cloud Functions with Pub/Sub introduces cold starts and per-invocation overhead, and a deep neural network exceeds practical function size and timeout limits, so sub-100ms latency at 1000 predictions/sec is unattainable. It is tempting for event-driven, low-volume inference where asynchronous triggering suits sporadic workloads.

  • ✗

    Use Vertex AI Batch Prediction with a fixed number of machines

    Why it's wrong here

    Batch Prediction processes jobs asynchronously over stored input, returning results after completion, so it cannot serve per-prediction requests under 100ms. It is tempting for large offline scoring runs where throughput matters but latency does not, such as nightly fraud model evaluation over historical data.

  • ✓

    Use Vertex AI Prediction with autoscaling enabled and GPU machine types

    Why this is correct

    Vertex AI Prediction with autoscaling and GPU machine types provides the parallel compute needed for sub-100ms deep neural network inference at 1000 predictions/sec. GPUs accelerate matrix operations, while autoscaling adds nodes under load, satisfying both the latency and throughput constraints.

About these practice questions

Courseiva writes every PMLE question from scratch — 775 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

This PMLE practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the PMLE exam.