PMLE Serving and Scaling Models Practice Question
You have a TensorFlow model that you want to deploy on edge devices for real-time inference. The model was trained in Vertex AI. You need to convert it to a format suitable for on-device inference. Which approach should you use?
⚠ Common exam trap
A common misconception is that Vertex AI services like Edge Manager or Model Optimization directly produce a deployable edge format, when in fact TensorFlow Lite conversion is the required final step for on-device inference.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Convert the model to TensorFlow Lite using the TensorFlow Lite converter.
TensorFlow Lite is specifically designed for on-device inference on edge devices, offering optimized performance and reduced model size. The TensorFlow Lite converter transforms a TensorFlow model (e.g., from a SavedModel) into the FlatBuffer format (.tflite), which is lightweight and compatible with mobile and embedded platforms. This directly addresses the requirement for real-time inference on edge devices.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Export the model as a serialized TFX pipeline.
Why it's wrong here
Exporting a serialized TFX pipeline preserves the training and transformation graph, not a deployable inference artefact, so edge devices cannot load it for real-time scoring. TFX pipelines are tempting because they orchestrate training, validation and deployment workflows in Vertex AI; that orchestration role is correct when you need reproducible production pipelines, not on-device conversion.
- ✗
Export the model to a SavedModel and deploy it using Vertex AI Edge Manager.
Why it's wrong here
Vertex AI Edge Manager was retired and never performed format conversion; a SavedModel is a serving artefact, not an on-device format. Edge Manager suited managing fleets of deployed models, not converting them for mobile or embedded inference.
- ✓
Convert the model to TensorFlow Lite using the TensorFlow Lite converter.
Why this is correct
TensorFlow Lite is Google's runtime for on-device inference, and the TensorFlow Lite converter transforms a trained TensorFlow SavedModel into the compact .tflite format with optimisations such as quantisation. This suits edge devices with limited compute and memory.
- ✗
Use Vertex AI Model Optimization to compile the model for edge devices.
Why it's wrong here
Vertex AI Model Optimization performs quantisation-aware training and pruning to shrink models, but it does not emit a TensorFlow Lite file for on-device runtimes. It suits reducing model size and latency before conversion, not the conversion step itself.
Go deeper
Related to this question
About these practice questions
Courseiva writes every PMLE question from scratch — 775 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →
Same concept, more angles
2 more ways this is tested on PMLE
These questions test the same concept from different angles. Work through them to make sure you can recognise it however the exam phrases it.
Variation 1. An organization wants to deploy a model on edge devices (e.g., Android phones) for offline inference. They trained a model using TensorFlow. Which THREE steps should they take to prepare and deploy the model?
medium- ✓ A.Convert the model to TensorFlow Lite format.
- B.Deploy the model to a Vertex AI endpoint for online inference.
- ✓ C.Use Vertex AI Edge Manager to package and deploy the model.
- D.Export the model to ONNX format.
- ✓ E.Deploy the TFLite model to the edge devices.
Why A: Option A is correct because TensorFlow Lite is Google's lightweight runtime format designed for on-device/offline inference on mobile and edge hardware like Android phones, so the trained TensorFlow model must be converted (e.g., via the TFLite Converter) to a .tflite file. Option C is correct because Vertex AI Edge Manager (part of the Vertex AI/Edge AI suite) packages the model and manages its distribution and deployment to edge devices, providing the tooling to push and monitor models at the edge. Option E is correct because after conversion the resulting TFLite model is what actually gets deployed and executed on the edge devices for offline inference, completing the workflow. Option B is not appropriate because deploying to a Vertex AI endpoint provides cloud-hosted online inference, which contradicts the offline, on-device requirement. Option D is not appropriate because ONNX is a different interchange format and is not the target runtime for TensorFlow models on Android edge devices in this scenario.
Variation 2. An organization wants to deploy a TensorFlow model on edge devices such as smartphones and IoT devices for offline inference. Which format should they export the model to?
medium- A.ONNX format
- ✓ B.TensorFlow Lite (TFLite)
- C.SavedModel format
- D.HDF5 format
Why B: TensorFlow Lite (TFLite) is Google's lightweight runtime and model format specifically designed for on-device inference on mobile and embedded/IoT hardware. It produces a compact .tflite flatbuffer optimized for low latency, small binary size, and minimal memory, and supports hardware acceleration via delegates (GPU, NNAPI, Edge TPU). Exporting to TFLite is the standard path for offline inference on smartphones and IoT devices.
JA
Written by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
This PMLE practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the PMLE exam.