PMLE Scaling Prototypes into ML Models Practice Question
An ML engineer is training a PyTorch model on Vertex AI using a custom training job. The dataset is stored in a Cloud Storage bucket with 500,000 small JPEG images. The training job is configured with an n1-standard-8 machine and a single NVIDIA T4 GPU. The engineer observes that GPU utilization is very low (around 15%) and training is slow. The model code is not the bottleneck. What is the most likely cause and the best solution?
⚠ Common exam trap
The trap here is assuming that low GPU utilization always means insufficient GPU compute and upgrading the GPU, instead of diagnosing the data input pipeline.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
The data loading pipeline is inefficient due to many small files; convert the dataset to TFRecord or WebDataset format and use a larger batch size with prefetching.
Low GPU utilization with many small files typically points to an I/O-bound data pipeline. Cloud Storage has high per-file access latency, so reading hundreds of thousands of small JPEGs individually creates a bottleneck. Converting to a sharded format like TFRecord or WebDataset enables efficient sequential reads, and combining with prefetching and larger batches ensures the GPU is continuously fed, improving utilization and training speed.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
The NVIDIA T4 GPU is not powerful enough for this model; upgrade to an NVIDIA A100 GPU.
Why it's wrong here
While an A100 is more powerful, the low GPU utilization (15%) indicates the GPU is idle waiting for data, not that it lacks compute power. Upgrading the GPU would not resolve the data loading bottleneck and would increase cost unnecessarily.
- ✗
The training job is using a single worker and not distributing the load; switch to distributed training with multiple workers.
Why it's wrong here
Distributed training would increase complexity and may not improve GPU utilization if the data pipeline is the bottleneck. The issue is likely data loading, not compute capacity. Adding workers without fixing data input would not solve the low GPU utilization.
- ✗
The Cloud Storage bucket is in a different region than the training job; move the bucket to the same region.
Why it's wrong here
While cross-region data transfer can add latency, the primary bottleneck with many small files is per-file overhead, not region. Moving the bucket might help slightly but does not address the fundamental issue of reading thousands of tiny files individually over the network.
- ✓
The data loading pipeline is inefficient due to many small files; convert the dataset to TFRecord or WebDataset format and use a larger batch size with prefetching.
Why this is correct
Reading many small JPEGs individually from Cloud Storage introduces high per-file latency and I/O overhead, starving the GPU. Converting to a sharded, sequential format like TFRecord or WebDataset allows efficient streaming, and using prefetching and larger batches keeps the GPU fed, increasing utilization.
Go deeper
Related to this question
About these practice questions
Courseiva writes every PMLE question from scratch — 775 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official Google Cloud exam blueprint
This PMLE practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the PMLE exam.