PMLE Scaling Prototypes into ML Models Practice Question
You are training a scikit-learn model on Vertex AI using a custom training job. The training dataset is a 2 TB CSV file stored in Cloud Storage, and the job must run on a single CPU-only VM. Loading the entire file into memory fails because the machine has only 32 GB of RAM. You need to train the model without increasing the VM size and without rewriting the training code to use a distributed framework. What should you do?
⚠ Common exam trap
The trap here is assuming that changing the file format or mounting Cloud Storage reduces the memory required by scikit-learn, when the real fix is an estimator that supports incremental partial_fit updates.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Use the scikit-learn partial_fit method on an estimator that supports incremental learning, reading the CSV in chunks with pandas.read_csv and feeding each chunk to partial_fit.
Incremental learning with partial_fit is the correct approach because it lets a scikit-learn estimator update its parameters one chunk at a time. Reading the CSV in chunks with pandas.read_csv bounds memory to the chunk size rather than the full 2 TB. This keeps the single-model semantics, avoids distributed training, and fits within the 32 GB VM, directly satisfying all the stated constraints.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Use Vertex AI Pipelines to split the CSV into many smaller files, then run a separate training job for each file and average the model weights.
Why it's wrong here
Splitting the data and training independent models on each shard produces an ensemble or an average of models, not a single model trained on the full dataset. Averaging weights from models trained on disjoint shards is not mathematically equivalent to training on all rows, and it requires extra orchestration. Vertex AI Pipelines is useful for repeatable workflows, but it does not solve the memory limit of loading a 2 TB CSV on a 32 GB VM.
- ✗
Convert the CSV to TFRecord format and use tf.data to stream batches, then train the scikit-learn model on the streamed batches.
Why it's wrong here
TFRecord and tf.data are TensorFlow technologies. A scikit-learn estimator expects a NumPy array or a pandas DataFrame, not a tf.data.Dataset, and cannot consume streamed TFRecord batches directly. Converting the file format also does not reduce the RAM needed to materialize the full training matrix for a scikit-learn fit call. This option mixes frameworks and does not address the actual memory bottleneck.
- ✓
Use the scikit-learn partial_fit method on an estimator that supports incremental learning, reading the CSV in chunks with pandas.read_csv and feeding each chunk to partial_fit.
Why this is correct
Many scikit-learn estimators such as SGDClassifier and MiniBatchKMeans implement partial_fit, which updates the model incrementally on small batches. Reading the 2 TB CSV with pandas.read_csv in chunks keeps only one chunk in memory at a time, so the 32 GB VM is sufficient. This preserves a single model trained over all rows and requires no distributed framework, which matches the constraint exactly.
- ✗
Mount the Cloud Storage bucket as a local filesystem on the training VM and call pandas.read_csv on the mounted path, relying on the OS page cache to keep memory usage low.
Why it's wrong here
Mounting Cloud Storage as a filesystem does not make pandas.read_csv memory-efficient. The read still materializes the entire 2 TB DataFrame in RAM, and the OS page cache does not prevent that allocation. Page cache helps with repeated reads of the same file blocks, but it cannot satisfy a single large in-memory allocation. The job would still fail with an out-of-memory error.
Go deeper
Related to this question
About these practice questions
Courseiva writes every PMLE question from scratch — 775 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official Google Cloud exam blueprint
This PMLE practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the PMLE exam.