Courseiva

PMLE Scaling Prototypes into ML Models Practice Question

You are training a PyTorch model on Vertex AI using a custom container. The training script uses DistributedDataParallel (DDP) with NCCL backend across 4 nodes, each with 8 GPUs. You notice that training throughput is low and GPUs are underutilized. After profiling, you find that the data loading is the bottleneck. You need to improve throughput without changing the model. What should you do?

⚠ Common exam trap

The trap here is assuming that more vCPUs or a different communication backend will solve the issue, when the real fix is to tune the DataLoader to feed data faster.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Increase the number of DataLoader worker processes and use a larger prefetch factor.

The data loading bottleneck in PyTorch distributed training is often alleviated by increasing the number of DataLoader worker processes and prefetch factor. This enables parallel data fetching and preloading of batches, keeping GPUs busy. Other options either do not address the bottleneck or could worsen performance.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✓

    Increase the number of DataLoader worker processes and use a larger prefetch factor.

    Why this is correct

    Increasing DataLoader workers and prefetch factor allows more parallel data loading and preloading of batches, reducing GPU idle time. This directly addresses the data loading bottleneck. It is a standard PyTorch optimization that requires minimal code changes and can significantly improve throughput in distributed training.

  • ✗

    Reduce the batch size per GPU to decrease memory pressure and allow more concurrent kernels.

    Why it's wrong here

    Reducing batch size may decrease memory usage but often leads to lower GPU utilization and longer training time because more iterations are needed. It does not address the data loading bottleneck. In fact, smaller batches may increase the relative overhead of data loading, worsening the issue.

  • ✗

    Switch from NCCL to Gloo backend for inter-node communication.

    Why it's wrong here

    Gloo is typically slower than NCCL for GPU-to-GPU communication and is not recommended for multi-GPU training. The bottleneck is data loading, not communication, so changing the backend would not help and could degrade performance. NCCL is optimized for NVIDIA GPUs and should be kept.

  • ✗

    Use a larger machine type with more vCPUs to increase data loading throughput.

    Why it's wrong here

    While more vCPUs can help data loading, simply increasing machine size without optimizing the DataLoader configuration may not fully utilize the extra cores. The DataLoader's num_workers parameter controls parallelism; if it is too low, additional vCPUs remain idle. Thus, tuning the DataLoader is more direct and cost-effective.

About these practice questions

Courseiva writes every PMLE question from scratch — 775 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official Google Cloud exam blueprint

This PMLE practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the PMLE exam.