AI0-001 AI Infrastructure and Technologies Practice Question
An AI research group trains a large language model across a cluster of GPU nodes. They observe that training throughput drops sharply whenever gradient synchronization occurs, and profiling shows GPUs idle while waiting for parameter updates to be exchanged. The model must remain mathematically identical to single-node training. Which change should the team make?
⚠ Common exam trap
The trap here is assuming that reducing communication frequency or volume is enough, when the real gain comes from overlapping communication with computation rather than shrinking it.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Switch to a distributed data-parallel strategy that overlaps gradient communication with backward computation, such as ring all-reduce with bucketing.
The idle time is caused by communication serialized after each backward pass. Ring all-reduce with bucketed, overlapped communication hides most of that transfer behind computation, so GPUs stay busy and the averaged gradient is unchanged. Because the arithmetic of the reduction is preserved, the resulting training run matches single-node behavior numerically.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Use lower-precision gradient compression that quantizes gradients to 8 bits before transmission.
Why it's wrong here
Quantizing gradients reduces communication volume and can speed synchronization, but it introduces numerical error into the updates. That directly violates the requirement that the model remain mathematically identical to single-node training. Compression schemes may be acceptable in some production settings, yet here the stated constraint rules them out.
- ✓
Switch to a distributed data-parallel strategy that overlaps gradient communication with backward computation, such as ring all-reduce with bucketing.
Why this is correct
Ring all-reduce exchanges gradients in chunks around the cluster so bandwidth is used evenly, and bucketing lets reduction of early layers begin while later layers are still computing backward. Communication then overlaps computation instead of serializing after it, cutting GPU idle time while producing the same averaged gradients, so the model remains mathematically identical to single-node training.
- ✗
Increase the number of gradient accumulation steps so synchronization happens less often.
Why it's wrong here
Accumulating gradients before synchronizing reduces the frequency of communication, but each synchronization still blocks until every worker contributes, and the effective batch changes, altering optimization behavior. The stall itself is not eliminated; it is merely deferred. The scenario requires removing idle time while preserving numerical equivalence, which this approach does not guarantee.
- ✗
Reduce the global batch size so that fewer gradients need to be exchanged per step.
Why it's wrong here
Shrinking the batch lowers the volume of data processed per step but does not remove the synchronization barrier that stalls GPUs. It also changes the optimization dynamics, potentially harming convergence and contradicting the requirement that training remain mathematically equivalent. The idle time stems from communication structure, not from batch volume alone.
Visual reference
About these practice questions
One of 962 original AI0-001 practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official CompTIA exam blueprint
This AI0-001 practice question is part of Courseiva's free CompTIA certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the AI0-001 exam.