Courseiva

AI0-001 AI Infrastructure and Technologies Practice Question

A research lab is training a large language model on a cluster of GPUs. They notice that training throughput decreases significantly when scaling from 8 to 16 GPUs. The model uses data parallelism with synchronous updates. Which factor is most likely causing the decreased throughput?

⚠ Common exam trap

The trap here is attributing throughput drops to model or hyperparameter issues rather than inter-GPU communication overhead.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Increased communication overhead for gradient synchronization across GPUs.

Synchronous data parallelism requires all-reduce operations to synchronize gradients. As the number of GPUs increases, the communication volume and frequency grow, and if the interconnect bandwidth is limited, this becomes a bottleneck. This is a common scaling challenge in distributed training, leading to sublinear speedup or even decreased throughput.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    The learning rate is too high for the larger effective batch size.

    Why it's wrong here

    A high learning rate can cause instability or divergence, but it typically does not reduce throughput; it affects convergence. Throughput is about how many samples are processed per second. The scenario describes a decrease in throughput, not a training failure, so learning rate is unlikely to be the cause.

  • ✗

    The model is not using mixed precision training.

    Why it's wrong here

    Mixed precision can improve throughput by using tensor cores, but its absence would affect baseline performance, not necessarily cause a drop when scaling from 8 to 16 GPUs. The issue is specifically about scaling, so communication overhead is a more direct explanation.

  • ✗

    Insufficient GPU memory causing out-of-memory errors.

    Why it's wrong here

    Out-of-memory errors would typically cause training to fail rather than just decrease throughput. If memory were insufficient, the batch size would need to be reduced, but the scenario does not mention errors. The issue is more likely related to communication overhead when scaling.

  • ✓

    Increased communication overhead for gradient synchronization across GPUs.

    Why this is correct

    In synchronous data parallelism, gradients must be averaged across all GPUs after each backward pass. As the number of GPUs increases, the all-reduce communication cost grows, potentially becoming a bottleneck. This overhead can reduce throughput if the network bandwidth or latency is insufficient, especially when scaling from 8 to 16 GPUs.

About these practice questions

Courseiva writes every AI0-001 question from scratch — 962 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official CompTIA exam blueprint

This AI0-001 practice question is part of Courseiva's free CompTIA certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the AI0-001 exam.