Courseiva

AI0-001 AI Infrastructure and Technologies Practice Question

A data scientist is using PyTorch to train a custom NLP model. The training is slow on a single GPU. They want to speed up training by using multiple GPUs on a single machine. Which PyTorch feature should they use?

⚠ Common exam trap

AI0-001 often tests the confusion between DataParallel and DistributedDataParallel — candidates pick DDP for single-machine multi-GPU, but the question emphasizes minimal code changes, and DataParallel requires only wrapping the model, while DDP requires initializing a process group and launching multiple processes.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

torch.nn.DataParallel

torch.nn.DataParallel is PyTorch's built-in module for single-machine, multi-GPU training. It wraps the model and automatically splits the input batch across available GPUs, replicating the model on each device and gathering outputs on the primary GPU — requiring only a one-line code change (wrapping the model).

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    TorchScript tracing

    Why it's wrong here

    TorchScript tracing converts eager models into serialisable graphs for deployment, not for distributing training across devices. It is tempting because it optimises inference performance, and would be the right choice when exporting a trained model to a production runtime. It cannot shard gradients or data across multiple GPUs.

  • ✓

    torch.nn.DataParallel

    Why this is correct

    DataParallel splits each batch across all GPUs on one machine and gathers gradients back to the primary device, giving single-process, multi-GPU acceleration with a one-line wrapper. It directly satisfies the single-machine, multiple-GPU constraint, unlike DistributedDataParallel, which targets multi-node scaling.

  • ✗

    torch.optim.SGD

    Why it's wrong here

    torch.optim.SGD is an optimiser implementing stochastic gradient descent weight updates, not a mechanism for distributing work across devices. It is tempting because it appears in every training loop, and would be correct when choosing an update algorithm. It provides no data or model parallelism across GPUs.

  • ✗

    PyTorch Lightning's zero_grad function

    Why it's wrong here

    zero_grad clears accumulated gradients before each backward pass; it performs no device distribution. It is tempting because Lightning abstracts multi-GPU training, and would be correct when resetting gradients in a training step. It cannot shard the model or data across several GPUs.

About these practice questions

This AI0-001 question is part of Courseiva's 962-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official CompTIA exam blueprint

This AI0-001 practice question is part of Courseiva's free CompTIA certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the AI0-001 exam.