PMLE Serving and Scaling Models Practice Question
You are designing a batch prediction pipeline using Vertex AI. The input data is 100 TB of images stored in Cloud Storage. The model is a custom TensorFlow model that expects TFRecord format. The pipeline must be cost-effective and run within a time window of 2 hours. Which THREE steps should you include?
⚠ Common exam trap
Google often tests the misconception that Cloud Functions can handle large-scale data processing tasks, but the trap here is that Cloud Functions have strict timeout and memory limits, making them unsuitable for converting 100 TB of images to TFRecord format.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Create a Vertex AI batch prediction job with input from GCS (TFRecord files).
Option B is correct because Vertex AI batch prediction natively accepts TFRecord input files from Cloud Storage, which matches the custom TensorFlow model's expected format and lets the managed service handle the large-scale inference job within the 2-hour window. Option C is correct because Dataflow is a scalable, cost-effective managed service for reading 100 TB of images from GCS and transforming them into TFRecord files, a task Cloud Functions cannot handle at this volume. Option D is correct because Vertex AI batch prediction writes its output to Cloud Storage, which is the appropriate and cost-effective destination for large result sets from 100 TB of input. Option A is not appropriate because BigQuery is not the default or cost-effective sink for massive batch prediction outputs and would add unnecessary loading overhead. Option E is not appropriate because Cloud Functions has execution time, memory, and concurrency limits that make it unsuitable for converting 100 TB of images to TFRecord.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Store batch prediction results in BigQuery.
Why it's wrong here
Writing predictions to BigQuery adds load and cost without serving the pipeline's requirement to produce predictions for 100 TB of images within two hours. It is tempting as a destination for structured analytics, and would be correct when downstream querying or reporting on prediction output is required.
- ✓
Create a Vertex AI batch prediction job with input from GCS (TFRecord files).
Why this is correct
A batch prediction job reading TFRecord files directly from Cloud Storage matches the model's expected input format, avoiding costly conversion of 100 TB. Vertex AI distributes the job across managed workers, meeting the two-hour window cost-effectively.
- ✓
Use Dataflow to read images and write TFRecord files to GCS.
Why this is correct
Dataflow is the managed, autoscaling Apache Beam service that parallelises reading 100 TB of images from Cloud Storage and serialising them into TFRecord shards. This satisfies the two-hour window and cost-effectiveness constraint, since TFRecord conversion cannot be done natively by Vertex AI batch prediction.
- ✓
Store batch prediction results in GCS.
Why this is correct
Vertex AI batch prediction writes output to a Cloud Storage destination, so specifying GCS as the sink is required to persist the 100 TB of predictions. It satisfies the pipeline's storage requirement, keeping results durable and accessible for downstream processing without extra services.
- ✗
Use Cloud Functions to convert images to TFRecord.
Why it's wrong here
Cloud Functions caps execution at minutes and memory, so converting 100 TB of images cannot complete within the two-hour window. It is tempting for lightweight event-driven transformation, and would suit small per-file conversions triggered by uploads rather than bulk preprocessing at this scale.
Quick reference
Cloud Service Model Comparison
| Model | You Manage | Provider Manages | Examples |
|---|---|---|---|
| IaaS | OS, runtime, apps, data | Hardware, hypervisor, networking | EC2, Azure VMs, GCP Compute Engine |
| PaaS | Apps and data | OS, runtime, middleware, hardware | Elastic Beanstalk, Azure App Service |
| SaaS | Data and settings only | Everything else | Microsoft 365, Salesforce, Workday |
| FaaS / Serverless | Function code only | Infra, scaling, runtime | Lambda, Azure Functions, Cloud Run |
| CaaS | Containers and apps | Kubernetes, OS, hardware | EKS, AKS, GKE |
Go deeper
Related to this question
About these practice questions
One of 775 original PMLE practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →
JA
Written by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
This PMLE practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the PMLE exam.