You are training a scikit-learn model on Vertex AI using a custom training job. The training dataset is a 2 TB CSV file stored in Cloud Storage, and the job must run on a single CPU-only VM. Loading the entire file into memory fails because the machine has only 32 GB of RAM. You need to train the model without increasing the VM size and without rewriting the training code to use a distributed framework. What should you do?
Many scikit-learn estimators such as SGDClassifier and MiniBatchKMeans implement partial_fit, which updates the model incrementally on small batches. Reading the 2 TB CSV with pandas.read_csv in chunks keeps only one chunk in memory at a time, so the 32 GB VM is sufficient. This preserves a single model trained over all rows and requires no distributed framework, which matches the constraint exactly.
Why this answer
Incremental learning with partial_fit is the correct approach because it lets a scikit-learn estimator update its parameters one chunk at a time. Reading the CSV in chunks with pandas.read_csv bounds memory to the chunk size rather than the full 2 TB. This keeps the single-model semantics, avoids distributed training, and fits within the 32 GB VM, directly satisfying all the stated constraints.
Exam trap
The trap here is assuming that changing the file format or mounting Cloud Storage reduces the memory required by scikit-learn, when the real fix is an estimator that supports incremental partial_fit updates.