PMLE Scaling Prototypes into ML Models Practice Question
A machine learning team is deploying a PyTorch model on Vertex AI Prediction for real-time inference. The model was trained with preprocessing that includes tokenization and normalization. They want to embed the preprocessing logic in the model to reduce prediction latency and avoid additional service calls. Which approach should they take?
⚠ Common exam trap
PMLE often tests the misconception that preprocessing must be a separate service — candidates pick Cloud Functions or Flask microservices because they are familiar patterns, but the question explicitly asks to embed preprocessing in the model, and TorchScript is the PyTorch-native way to do that.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Use TorchScript to trace the preprocessing steps and export the entire pipeline as a single scripted model
TorchScript allows you to trace or script a PyTorch model, including preprocessing operations like tokenization and normalization, into a single serialized artifact. By embedding preprocessing in the TorchScript model, the entire pipeline runs in one forward pass on the Vertex AI endpoint, eliminating extra service calls and reducing latency. This is the standard approach for consolidating preprocessing with a PyTorch model for real-time inference.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Deploy the preprocessing logic as a Cloud Function and invoke it before calling the prediction endpoint
Why it's wrong here
A Cloud Function adds a separate network hop and service call before prediction, which is exactly the latency and extra-call overhead the team wants to eliminate. Cloud Functions suit decoupled, event-driven preprocessing outside the model. Embedding the tokenisation and normalisation layers inside the PyTorch module keeps inference self-contained.
- ✗
Wrap the preprocessing logic in a Flask application and deploy it as a separate microservice in front of the prediction endpoint
Why it's wrong here
Deploying preprocessing as a separate Flask microservice introduces network round-trip latency between the microservice and the prediction endpoint, directly contradicting the requirement to reduce prediction latency. This approach is tempting because it cleanly separates concerns and would be correct if the preprocessing logic needed independent scaling or was shared across multiple services, but here the goal is to embed logic within the model artifact to avoid any inter-service calls.
- ✓
Use TorchScript to trace the preprocessing steps and export the entire pipeline as a single scripted model
Why this is correct
TorchScript tracing captures the tokenisation and normalisation operations as graph nodes, fusing them with the PyTorch model into one serialised artefact. Vertex AI Prediction then serves this single scripted model, eliminating the separate preprocessing service call and satisfying the stem's latency-reduction constraint.
- ✗
Use TensorFlow Transform to convert preprocessing into a SavedModel and call it from the PyTorch model
Why it's wrong here
TensorFlow Transform emits a SavedModel for TensorFlow serving signatures, so a PyTorch model cannot invoke it inline during forward passes. It suits TF-based pipelines needing consistent train/serve feature engineering. Embedding tokenisation and normalisation directly in the PyTorch module avoids the extra service call the stem requires.
Go deeper
Related to this question
About these practice questions
This PMLE question is part of Courseiva's 775-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official Google Cloud exam blueprint
This PMLE practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the PMLE exam.