AI0-001 AI Infrastructure and Technologies Practice Question
A data scientist is training a large language model on a custom dataset using PyTorch on AWS. The training is taking too long due to GPU memory constraints. The team wants to use multiple GPUs across instances with minimal code changes. Which AWS service should they use?
⚠ Common exam trap
The trap is confusing infrastructure services (EFA, Batch, ParallelCluster) with managed machine learning services; candidates may pick EFA because it sounds like a networking solution for distributed training, but it lacks the high-level libraries.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Amazon SageMaker with distributed training libraries
Amazon SageMaker with distributed training libraries provides built-in support for data and model parallelism across multiple GPUs and instances, requiring minimal code changes. It integrates with PyTorch and handles the orchestration of distributed training, making it the best choice for scaling training.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
AWS Elastic Fabric Adapter (EFA)
Why it's wrong here
EFA is a network adapter delivering low-latency inter-node transport, not a framework that shards a model across GPUs to relieve memory pressure. It is tempting because distributed training depends on fast interconnect, and EFA is the correct choice when networking, not memory, is the bottleneck.
- ✓
Amazon SageMaker with distributed training libraries
Why this is correct
SageMaker distributed training libraries handle data and model parallelism across multiple GPU instances with minimal PyTorch code changes, directly addressing the GPU memory constraint. It satisfies the stem's requirement to scale beyond a single instance without rewriting the training script.
- ✗
AWS Batch with GPU instances
Why it's wrong here
AWS Batch schedules containerised jobs onto GPU instances but provides no model-parallel or data-parallel sharding, so per-GPU memory limits persist. It is tempting because it manages GPU compute queues, which suits batch inference or embarrassingly parallel jobs rather than distributed training.
- ✗
AWS ParallelCluster with Slurm
Why it's wrong here
ParallelCluster with Slurm orchestrates HPC cluster provisioning but does not itself distribute a PyTorch model across GPUs or reduce per-GPU memory; it requires substantial scripting. It is tempting because it manages multi-instance GPU clusters, which suits tightly coupled simulation workloads rather than minimal-change distributed training.
About these practice questions
This AI0-001 question is part of Courseiva's 962-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official CompTIA exam blueprint
This AI0-001 practice question is part of Courseiva's free CompTIA certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the AI0-001 exam.