AIF-C01 Fundamentals of AI and ML Practice Question
A company is using Amazon SageMaker to train a large language model with hundreds of billions of parameters. The model does not fit into the memory of a single GPU. Which approach should they use to train the model efficiently?
⚠ Common exam trap
The AIF-C01 exam often tests the distinction between data parallelism and model parallelism, and the trap here is that candidates may confuse data parallelism (which splits data, not the model) as a solution for models that don't fit in memory, when in fact model parallelism is required for such cases.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Use SageMaker's model parallelism strategy with the SageMaker distributed training library
SageMaker's model parallelism strategy with the SageMaker distributed training library is specifically designed for training large models that do not fit into the memory of a single GPU. It partitions the model layers across multiple GPUs, enabling efficient training of models with hundreds of billions of parameters by overlapping computation and communication.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Use a larger instance with more GPU memory, such as p4d.24xlarge
Why it's wrong here
A single p4d.24xlarge provides roughly 320 GB of GPU memory across eight A100s, still far short of hundreds of billions of parameters plus optimiser states. It is tempting because scaling up is the obvious first instinct, and it would suffice for models fitting one device, but sharding across many GPUs is required.
- ✗
Use SageMaker's data parallelism strategy
Why it's wrong here
Data parallelism replicates the full model on every GPU, so each replica must still hold all parameters and optimiser states — impossible when the model exceeds one device. It is tempting because it is the standard SageMaker distributed strategy, and it would be correct for speeding up training of a model that already fits in single-GPU memory.
- ✓
Use SageMaker's model parallelism strategy with the SageMaker distributed training library
Why this is correct
Model parallelism shards the model's layers and parameters across multiple GPUs, so no single device must hold the full hundreds-of-billions-parameter model. This directly resolves the stem's constraint that the model does not fit into one GPU's memory.
- ✗
Reduce the model size by pruning layers until it fits into memory
Why it's wrong here
Pruning alters the model's architecture and accuracy, and it does not address the memory needed to train the original hundreds-of-billions-parameter model in the first place. It is tempting because reducing size sounds like a direct fix, and it would be valid for compressing an already-trained model for cheaper inference deployment.
Go deeper
Related to this question
About these practice questions
One of 862 original AIF-C01 practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →
JA
Written by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
This AIF-C01 practice question is part of Courseiva's free Amazon Web Services certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the AIF-C01 exam.