Courseiva

AI0-001 AI Infrastructure and Technologies Practice Question

A startup is training a recommendation model on a single workstation with one GPU. The dataset has grown to 2 TB, and training now takes several days. The team wants to reduce training time by adding more GPUs to the same workstation. Which technology should they use to enable efficient multi-GPU training with minimal code changes?

⚠ Common exam trap

The trap here is assuming that any parallel computing framework, such as Hadoop, can accelerate GPU training, when it lacks the necessary communication primitives.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

NVIDIA NCCL with DistributedDataParallel

NCCL with DistributedDataParallel is the standard approach for multi-GPU training within a single node. It provides efficient gradient synchronization and requires minimal code changes in PyTorch. The other options are either for data processing, single-GPU sharing, or edge inference, none of which solve the training time issue.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    Apache Hadoop MapReduce

    Why it's wrong here

    Hadoop MapReduce is designed for batch data processing, not for iterative GPU-accelerated model training. It does not provide the low-latency communication required for gradient synchronization. Using it for deep learning would introduce significant overhead and is not a practical solution for multi-GPU training.

  • ✗

    CUDA Multi-Process Service (MPS)

    Why it's wrong here

    CUDA MPS allows multiple processes to share a single GPU, improving utilization for concurrent workloads. However, it does not distribute a single training job across multiple GPUs. It is not a mechanism for data-parallel training and would not reduce the training time of one model.

  • ✓

    NVIDIA NCCL with DistributedDataParallel

    Why this is correct

    NCCL provides optimized inter-GPU communication, and DistributedDataParallel in PyTorch enables data-parallel training across multiple GPUs with minimal code changes. This combination is the standard for scaling training on a single node with multiple GPUs, directly reducing training time.

  • ✗

    TensorFlow Lite

    Why it's wrong here

    TensorFlow Lite is an inference framework for mobile and edge devices, not a training solution. It cannot train large models on multiple GPUs. Deploying it here would not address the training time problem and is entirely mismatched to the scenario.

About these practice questions

This AI0-001 question is part of Courseiva's 962-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official CompTIA exam blueprint

This AI0-001 practice question is part of Courseiva's free CompTIA certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the AI0-001 exam.