Courseiva
mediumMultiple Choice

MLA-C01 Practice Question: Training a deep learning model on Amazon SageMaker

A company is training a deep learning model on Amazon SageMaker. The training job started but has been stuck in 'InProgress' state for an unusually long time with low CPU utilization. The data scientist suspects a bottleneck. What should be the first troubleshooting step?

⚠ Common exam trap

The trap here is that candidates often jump to scaling or instance changes (Options B and C) without first checking logs, assuming a performance issue is hardware-related when it is almost always a software or configuration issue inside the container.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Review CloudWatch Logs for the training container to identify errors or warnings.

When a SageMaker training job is stuck in 'InProgress' with low CPU utilization, the most common cause is a bottleneck in data loading or preprocessing within the training container. Reviewing CloudWatch Logs for the training container is the first troubleshooting step because it provides direct visibility into container-level errors, warnings, or stalls (e.g., hanging on a file read, waiting for a dependency, or a misconfigured data channel) that would not be visible from instance-level metrics alone.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    Switch the training job to use Spot instances to reduce cost and potentially improve throughput.

    Why it's wrong here

    Spot instances change billing and can be interrupted; they do not diagnose why CPU utilisation is low, which typically indicates an input-pipeline or GPU-bound bottleneck. It is tempting because Spot capacity cuts training cost, and would be the right answer when the goal is cost reduction rather than identifying a stalled job's bottleneck.

  • ✗

    Increase the number of training instances to parallelize data loading.

    Why it's wrong here

    Adding instances parallelises compute, not the stalled data-loading path; low CPU with a long InProgress state points to input pipeline starvation, so more instances simply idle too. It is tempting because horizontal scaling genuinely accelerates distributed training once data feeds each worker fast enough.

  • ✗

    Stop and restart the training job with a different instance type.

    Why it's wrong here

    Swapping instance type changes compute capacity, yet the symptom is low CPU utilisation, meaning the job is waiting on data rather than lacking processing power. It is tempting because resizing is a familiar remedy for slow training, and it would be correct if CPU or GPU were saturated.

  • ✓

    Review CloudWatch Logs for the training container to identify errors or warnings.

    Why this is correct

    CloudWatch Logs capture the training container's stdout and stderr, revealing errors, warnings, or stalls such as failed data downloads or dependency issues. With low CPU utilisation and a stuck InProgress state, inspecting these logs first identifies the underlying bottleneck before deeper changes.

About these practice questions

Courseiva writes every MLA-C01 question from scratch — 665 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

This MLA-C01 practice question is part of Courseiva's free Amazon Web Services certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the MLA-C01 exam.