Courseiva

AI0-001 AI Infrastructure and Technologies Practice Question

A machine learning team is training a large transformer model on a text corpus. They need to reduce training time while maintaining model accuracy. Which hardware configuration would be MOST effective for this task?

⚠ Common exam trap

AI0-001 often tests the misconception that a single powerful device (TPU or GPU) with model parallelism is better than a multi-GPU cluster; candidates may pick option D because TPUs are marketed for transformers, but the question emphasizes reducing training time, which favors data parallelism across many GPUs.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Use a cluster of GPUs with data parallelism

Training large transformer models is computationally intensive, and data parallelism across a cluster of GPUs allows the model to process multiple batches simultaneously, dramatically reducing training time. Each GPU holds a full copy of the model and processes a different subset of the data, with gradients synchronized across devices. This approach scales well and maintains accuracy as long as the effective batch size is tuned appropriately.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    Use a high-core-count CPU with large RAM

    Why it's wrong here

    CPUs execute transformer matrix multiplications serially across few vector units, so training time barely improves regardless of core count or RAM. High-core CPUs suit data preprocessing, feature engineering or small classical models, not the massively parallel tensor operations a large transformer requires.

  • ✓

    Use a cluster of GPUs with data parallelism

    Why this is correct

    Data parallelism distributes each batch across many GPUs, each holding a full model replica and synchronising gradients, which cuts wall-clock training time substantially. This scales effectively for large transformer models while preserving accuracy through equivalent gradient updates.

  • ✗

    Use a single GPU with model parallelism

    Why it's wrong here

    Model parallelism on a single GPU fails to reduce training time because it partitions the model across the device’s own memory, introducing sequential computation and communication overhead that does not scale throughput. It is tempting because it enables fitting a transformer that exceeds a single GPU’s VRAM, but the correct approach for speed is data parallelism across multiple GPUs, which processes distinct micro-batches concurrently.

  • ✗

    Use a single TPU with model parallelism

    Why it's wrong here

    A single TPU cannot hold a large transformer's parameters and optimiser states, so model parallelism across one device yields no throughput gain. TPUs excel when sharded across many chips via data and model parallelism together, which is the configuration this scenario actually demands.

About these practice questions

Courseiva writes every AI0-001 question from scratch — 962 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official CompTIA exam blueprint

This AI0-001 practice question is part of Courseiva's free CompTIA certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the AI0-001 exam.