Courseiva
mediumMultiple Choice

PDE Practice Question: A retail company needs to generate product…

A retail company needs to generate product recommendations for millions of users every few hours. The model is a small scikit-learn model. Which prediction method should be used to minimize infrastructure cost while meeting the latency requirements?

⚠ Common exam trap

Google Cloud often tests the distinction between online (real-time) and batch (asynchronous) prediction patterns, and the trap here is that candidates assume 'predictions' always require a live endpoint, overlooking that batch jobs are the correct choice when latency requirements are in hours and the workload is massive and periodic.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Use a Vertex AI batch prediction job that reads from BigQuery and writes results back to BigQuery or Cloud Storage.

Batch prediction is the most cost-effective approach for generating recommendations for millions of users every few hours. Vertex AI batch prediction jobs process large datasets in parallel without maintaining always-on infrastructure, and they can read from BigQuery and write results directly to BigQuery or Cloud Storage, minimizing compute costs while meeting the latency requirement of 'every few hours' (not real-time).

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    Use Cloud Run to host the model and invoke it for each user request.

    Why it's wrong here

    Cloud Run bills per request and cold-starts each invocation, so millions of individual user calls incur continuous per-request charges and latency spikes. It suits low-volume, interactive online prediction where per-request scaling matters, not bulk batch scoring of millions of users on a schedule.

  • ✗

    Export the model as a container and run on Google Kubernetes Engine with cluster autoscaling.

    Why it's wrong here

    A GKE cluster with autoscaling keeps node pools provisioned between the every-few-hours runs, so the team pays for idle capacity. GKE fits sustained, always-on serving with steady request volume, not periodic batch jobs where resources should exist only during the run.

  • ✗

    Deploy the model to a Vertex AI endpoint with a single replica for online predictions.

    Why it's wrong here

    A single-replica endpoint stays provisioned continuously, billing for idle capacity between the few-hourly batch runs. Batch prediction submits a job that provisions resources only for the run, then releases them, which matches the periodic schedule and minimises cost.

  • ✓

    Use a Vertex AI batch prediction job that reads from BigQuery and writes results back to BigQuery or Cloud Storage.

    Why this is correct

    Batch prediction processes the entire input dataset in a single managed job, so Vertex AI provisions resources only for the job's duration rather than continuously. This satisfies the stem's cost-minimisation constraint while handling millions of users every few hours.

About these practice questions

One of 747 original PDE practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

This PDE practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the PDE exam.