Courseiva
hardMultiple Select

MLA-C01 Practice Question: Which THREE steps should be taken to optimize a…

Which THREE steps should be taken to optimize a large-scale distributed training job on SageMaker? (Choose 3.)

⚠ Common exam trap

Test-takers frequently confuse storage optimization (EBS) or inference features (batch transform) with training optimization, failing to recognize that distributed training performance hinges on compute, memory, and inter-node communication, not disk I/O or post-training steps.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Use GPU instances with high bandwidth and memory (e.g., ml.p4d.24xlarge).

Option B is correct because large-scale distributed training is compute- and memory-intensive, so GPU instances like ml.p4d.24xlarge provide high-bandwidth networking (up to 400 Gbps) and large GPU memory (A100 40 GB each) that accelerate training and reduce communication bottlenecks. Option D is correct because Elastic Fabric Adapter (EFA) enables low-latency, high-throughput inter-node communication using OS bypass, which is essential for scaling distributed training across many instances. Option E is correct because choosing the right distributed training strategy — Horovod for data parallelism, SageMaker distributed data parallel for optimized AllReduce, or model parallel for models too large for a single GPU — directly determines training efficiency and scalability. Option A is not appropriate because attaching multiple EBS volumes with provisioned throughput addresses storage I/O, not the inter-node GPU communication and compute parallelism that dominate large-scale distributed training. Option C is not appropriate because batch transform is an inference-time feature and does not optimize the training job itself.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    Attach multiple EBS volumes with throughput provisioning.

    Why it's wrong here

    Adding EBS volumes increases local disk throughput, which does not accelerate distributed training, where the bottleneck is inter-node network communication and GPU compute. It is tempting because storage provisioning helps data-heavy single-node workloads, and it would be correct when input pipelines are I/O bound rather than communication bound.

  • ✓

    Use GPU instances with high bandwidth and memory (e.g., ml.p4d.24xlarge).

    Why this is correct

    High-bandwidth GPU instances such as ml.p4d.24xlarge provide the inter-node communication throughput and per-device memory that distributed training demands, directly addressing the stem's large-scale constraint. Gradient synchronisation across many workers is bandwidth-bound, so faster interconnect and larger memory reduce communication stalls and prevent out-of-memory failures during training.

  • ✗

    Enable batch transform for offline inference after training.

    Why it's wrong here

    Batch transform runs inference after training completes, so it cannot affect the training job's speed or convergence. It is tempting because it is a genuine SageMaker optimisation feature, and it would be correct when optimising offline scoring throughput rather than the distributed training run itself.

  • ✓

    Use Elastic Fabric Adapter (EFA) for low-latency inter-node communication.

    Why this is correct

    Elastic Fabric Adapter provides kernel-bypass networking with hardware-level reliable transport, cutting inter-node latency and CPU overhead that standard TCP imposes during gradient exchange. For a large-scale distributed training job, this directly addresses the communication bottleneck across instances, satisfying the stem's optimisation requirement by scaling efficiently beyond a single node.

  • ✓

    Select the appropriate distributed training strategy (e.g., Horovod, SageMaker data parallel, or model parallel).

    Why this is correct

    Choosing the right distributed training strategy directly addresses the scaling constraint: Horovod and SageMaker data parallel split batches across instances for throughput, while model parallel shards layers too large for one GPU's memory. This matches the job's large-scale, multi-instance requirement.

About these practice questions

One of 665 original MLA-C01 practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

This MLA-C01 practice question is part of Courseiva's free Amazon Web Services certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the MLA-C01 exam.