PMLE Scaling Prototypes into ML Models Practice Question
You are training a model on Vertex AI using a custom training job. The training data is stored in a Cloud Storage bucket in the us-central1 region, and the training job runs in the us-central1 region. You notice that the training job takes significantly longer than expected due to data loading. You want to improve data loading performance without changing the model architecture. What should you do?
⚠ Common exam trap
The trap here is assuming that simply increasing buffer size or copying data locally will solve the data loading bottleneck without addressing the format and parallelization of reads.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Convert the training data to TFRecord format and use the tf.data.TFRecordDataset with parallel interleave and prefetching.
Converting data to TFRecord and using tf.data with parallel interleave and prefetching optimizes the input pipeline by reducing the number of read operations and overlapping data loading with computation. This is a standard best practice for improving training performance on Vertex AI when data is in Cloud Storage.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Copy the training data to the local SSD of the training VM before training begins, and read from there.
Why it's wrong here
Copying data to local SSD can improve I/O performance, but it adds a startup delay and requires enough disk space. For large datasets, this may not be feasible. Moreover, the copy operation itself takes time and may not be faster than reading directly from Cloud Storage if the data is already in the same region. This is not the optimal solution.
- ✓
Convert the training data to TFRecord format and use the tf.data.TFRecordDataset with parallel interleave and prefetching.
Why this is correct
TFRecord is a binary format that is more efficient to read than many small files. Using tf.data.TFRecordDataset with parallel interleave and prefetching allows overlapping data loading with model training, significantly improving throughput. This is a recommended practice for optimizing input pipelines on Vertex AI, especially when data is stored in Cloud Storage.
- ✗
Enable streaming reads from Cloud Storage by using the tf.data.experimental.make_csv_dataset function with a large buffer size.
Why it's wrong here
Streaming reads from Cloud Storage can help for large datasets, but increasing the buffer size alone does not address the fundamental latency of accessing many small files. This approach may improve throughput slightly, but it is not the most effective solution for reducing data loading time in this scenario. It also does not leverage Vertex AI's built-in optimizations for data loading.
- ✗
Use the Vertex AI Training reduction server to accelerate data loading.
Why it's wrong here
The reduction server is designed to accelerate distributed training by reducing gradient communication overhead, not to improve data loading from Cloud Storage. It is irrelevant to the data loading bottleneck described. Using it would not address the slow data loading and could introduce unnecessary complexity.
Go deeper
Related to this question
About these practice questions
This PMLE question is part of Courseiva's 775-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official Google Cloud exam blueprint
This PMLE practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the PMLE exam.