Courseiva
ML Model Development →hardMultiple Select

MLA-C01 ML Model Development Practice Question

A company is training a deep learning model for object detection using SageMaker. The training is very slow and the GPU memory is insufficient for the batch size. The team wants to scale across multiple GPUs efficiently. Which THREE actions should they take? (Choose THREE.)

⚠ Common exam trap

MLA-C01 often tests the distinction between cost-optimization features (spot instances), observability features (Debugger), and actual distributed-training mechanisms — candidates pick spot instances or Debugger thinking they 'help with scaling,' but only the distributed libraries and their SDK configuration address the memory and multi-GPU scaling requirement.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Use SageMaker distributed model parallelism

Option A is correct because SageMaker distributed model parallelism shards the model itself across multiple GPUs, which directly addresses the insufficient GPU memory problem by allowing a model too large for a single GPU to be split across devices, and it also speeds up training of large models. Option B is correct because SageMaker distributed data parallelism splits each mini-batch across GPUs and uses AllReduce for efficient gradient synchronization, enabling the team to scale the effective batch size and throughput across multiple GPUs efficiently. Option D is correct because the SageMaker distributed training configuration via the SageMaker SDK (e.g., the Distribution parameter with smdistributed settings in the estimator) is the required mechanism to actually enable and launch model or data parallelism on the training cluster. Option C is not correct because managed spot instances reduce cost, not training time or GPU memory constraints, and can even introduce interruptions. Option E is not correct because SageMaker Debugger only monitors and reports training bottlenecks; it does not itself scale training across multiple GPUs or resolve memory limitations.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✓

    Use SageMaker distributed model parallelism

    Why this is correct

    Model parallelism shards a single large model's layers across multiple GPUs, distributing parameters and activations so each device holds only a fraction. This directly addresses insufficient per-GPU memory for the batch size while scaling training across GPUs.

  • ✓

    Use SageMaker distributed data parallelism

    Why this is correct

    Distributed data parallelism shards each mini-batch across GPUs and synchronises gradients, so the effective batch size scales with GPU count. This directly addresses the insufficient per-GPU memory and slow training by using aggregate memory and compute across multiple GPUs in SageMaker.

  • ✗

    Use managed spot instances

    Why it's wrong here

    Managed spot instances cut compute cost through capacity reuse; they neither add GPU memory nor enable multi-GPU distribution, and can interrupt long training. They tempt cost-conscious teams, and would be correct when the priority is reducing spend on interruptible, checkpointed training jobs.

  • ✓

    Use a SageMaker distributed training configuration with the SageMaker SDK

    Why this is correct

    The SageMaker SDK's distributed training configuration activates the framework's distributed libraries and sets parameters such as process count per host. Without this configuration, the training job runs single-GPU, so it is the mechanism that actually enables multi-GPU scaling in SageMaker.

  • ✗

    Enable SageMaker Debugger to identify bottlenecks

    Why it's wrong here

    Debugger profiles metrics such as vanishing gradients and system bottlenecks; it does not increase GPU memory or distribute training across devices. It tempts teams wanting visibility into slow jobs, and would be correct when diagnosing training instability or resource utilisation rather than resolving memory and scaling limits.

Visual reference

Client Recursive Resolver Root DNS (13 root servers) TLD DNS (.com, .org, …) Authoritative example.com query IP addr answer

About these practice questions

Courseiva writes every MLA-C01 question from scratch — 665 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official Amazon Web Services exam blueprint

This MLA-C01 practice question is part of Courseiva's free Amazon Web Services certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the MLA-C01 exam.