Courseiva

AIF-C01 Fundamentals of AI and ML Practice Question

A company is using Amazon SageMaker to train a large language model with hundreds of billions of parameters. The model does not fit into the memory of a single GPU. Which approach should they use to train the model efficiently?

⚠ Common exam trap

The AIF-C01 exam often tests the distinction between data parallelism and model parallelism, and the trap here is that candidates may confuse data parallelism (which splits data, not the model) as a solution for models that don't fit in memory, when in fact model parallelism is required for such cases.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Use SageMaker's model parallelism strategy with the SageMaker distributed training library

SageMaker's model parallelism strategy with the SageMaker distributed training library is specifically designed for training large models that do not fit into the memory of a single GPU. It partitions the model layers across multiple GPUs, enabling efficient training of models with hundreds of billions of parameters by overlapping computation and communication.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    Use a larger instance with more GPU memory, such as p4d.24xlarge

    Why it's wrong here

    A single p4d.24xlarge provides roughly 320 GB of GPU memory across eight A100s, still far short of hundreds of billions of parameters plus optimiser states. It is tempting because scaling up is the obvious first instinct, and it would suffice for models fitting one device, but sharding across many GPUs is required.

  • ✗

    Use SageMaker's data parallelism strategy

    Why it's wrong here

    Data parallelism replicates the full model on every GPU, so each replica must still hold all parameters and optimiser states — impossible when the model exceeds one device. It is tempting because it is the standard SageMaker distributed strategy, and it would be correct for speeding up training of a model that already fits in single-GPU memory.

  • ✓

    Use SageMaker's model parallelism strategy with the SageMaker distributed training library

    Why this is correct

    Model parallelism shards the model's layers and parameters across multiple GPUs, so no single device must hold the full hundreds-of-billions-parameter model. This directly resolves the stem's constraint that the model does not fit into one GPU's memory.

  • ✗

    Reduce the model size by pruning layers until it fits into memory

    Why it's wrong here

    Pruning alters the model's architecture and accuracy, and it does not address the memory needed to train the original hundreds-of-billions-parameter model in the first place. It is tempting because reducing size sounds like a direct fix, and it would be valid for compressing an already-trained model for cheaper inference deployment.

About these practice questions

One of 862 original AIF-C01 practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

This AIF-C01 practice question is part of Courseiva's free Amazon Web Services certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the AIF-C01 exam.