Courseiva

PMLE Scaling Prototypes into ML Models Practice Question

You have a Python training script that reads a 500 GB CSV dataset from a Cloud Storage bucket. You submit a Vertex AI custom training job using a pre-built container, specifying a machine with 16 vCPUs and 60 GB RAM. The job fails after a few minutes with an out-of-memory error. You need to scale the prototype to handle this dataset without changing the model architecture. What should you do?

⚠ Common exam trap

The trap here is assuming that simply increasing machine memory or splitting files will solve the problem, when the real fix is to stream data instead of loading it all at once.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Convert the CSV files to TFRecord format and use the tf.data API to stream the data in batches during training.

The out-of-memory error occurs because the training script attempts to load the entire 500 GB dataset into memory. Using TFRecord and tf.data allows streaming data in batches, so memory usage remains bounded. This is the standard method for scaling data input on Vertex AI and requires minimal changes to the training code.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    Increase the machine type to one with 128 GB RAM and retry the job.

    Why it's wrong here

    Increasing RAM may delay the failure but does not solve the underlying issue: the dataset is far larger than any practical memory size. Eventually the job will still exhaust memory as data grows, and this approach is not cost-effective. It does not enable scaling to arbitrarily large datasets, which is required here.

  • ✓

    Convert the CSV files to TFRecord format and use the tf.data API to stream the data in batches during training.

    Why this is correct

    Converting to TFRecord and using tf.data enables efficient streaming from Cloud Storage, so the full dataset is never loaded into memory. This directly addresses the out-of-memory error by decoupling data size from machine memory. It is the recommended approach for large-scale training on Vertex AI and requires only changing the input pipeline, not the model.

  • ✗

    Split the CSV files into smaller files and use a larger number of workers with data parallelism.

    Why it's wrong here

    Splitting files alone does not reduce per-worker memory if each worker still loads its entire shard into memory. Without a streaming input pipeline, each worker will still attempt to load its assigned data fully. Data parallelism increases compute but does not fix the memory bottleneck caused by loading all data at once.

  • ✗

    Use Vertex AI Pipelines to preprocess the data and store it in BigQuery, then read from BigQuery during training.

    Why it's wrong here

    Reading from BigQuery during training still requires loading data into memory unless a streaming approach is used. This adds complexity and does not inherently solve the out-of-memory error. While BigQuery can handle large data, the training script must still stream it, which is not addressed by this option.

About these practice questions

This PMLE question is part of Courseiva's 775-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official Google Cloud exam blueprint

This PMLE practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the PMLE exam.