Courseiva
mediumMultiple Choice

MLA-C01 Practice Question: A data scientist is training a large model on…

A data scientist is training a large model on SageMaker and wants to reduce training time by using multiple GPUs. The model is small enough to fit on a single GPU but training is slow. Which SageMaker feature should be used?

⚠ Common exam trap

Many candidates confuse model parallelism (for large models) with data parallelism (for slow training of small models), or mistakenly think Elastic Inference can accelerate training when it is strictly for inference latency reduction.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Data parallelism using SageMaker's Distributed Data Parallel

SageMaker's Distributed Data Parallel (DDP) is the correct choice because it splits the mini-batch across multiple GPUs, allowing each GPU to hold a copy of the model and process a subset of the data simultaneously. This reduces training time for models that fit on a single GPU by leveraging data parallelism, where gradients are synchronized across GPUs after each step.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✓

    Data parallelism using SageMaker's Distributed Data Parallel

    Why this is correct

    Data parallelism replicates the small model across multiple GPUs, each processing a shard of the batch, then synchronises gradients via SageMaker's Distributed Data Parallel. This cuts training time by using several GPUs concurrently, matching the requirement.

  • ✗

    Use a larger instance with more vCPUs

    Why it's wrong here

    Additional vCPUs increase host CPU capacity, not GPU parallelism, so the training job still runs on one GPU and remains slow. It is tempting because larger instances often speed up workloads, but this choice is correct when the bottleneck is CPU-bound data preprocessing or input pipelines, not GPU compute.

  • ✗

    Model parallelism using SageMaker's Model Parallel

    Why it's wrong here

    Model parallelism splits a single model's layers across GPUs, which helps when the model cannot fit on one device; here it fits, so splitting adds inter-GPU communication overhead. Data parallelism replicates the model per GPU and is the correct choice for this scenario.

  • ✗

    Use Elastic Inference

    Why it's wrong here

    Elastic Inference attaches fractional GPU accelerators for inference only, so it cannot distribute the training workload across GPUs. It is tempting because it accelerates deep learning, but its purpose is cost-effective inference hosting, making it the right pick when serving predictions rather than shortening multi-GPU training jobs.

About these practice questions

One of 665 original MLA-C01 practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

This MLA-C01 practice question is part of Courseiva's free Amazon Web Services certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the MLA-C01 exam.