Courseiva

MLA-C01 Data Preparation for Machine Learning Practice Question

A machine learning engineer runs a SageMaker Processing job that must load a 200 GB dataset from S3, compute statistics, and write a small summary to S3. The job repeatedly fails with an out-of-disk-space error on the processing instance. Which change is MOST likely to resolve the failure?

⚠ Common exam trap

The trap here is assuming Pipe mode applies to every SageMaker job type, when streaming input is a training-job feature and Processing jobs always mount files locally.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Increase the processing instance's volume size and use ShardedByS3Key so each instance downloads only its share of objects.

Processing jobs download input channels to local storage, so a 200 GB dataset can exhaust the default volume. Enlarging the volume gives the container room, and ShardedByS3Key spreads objects across multiple instances so each holds only a fraction. Pipe mode is not applicable to Processing jobs, and the other options adjust runtime or cleanup behavior without changing disk consumption.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    Add a lifecycle configuration script that deletes files from the input directory after each epoch.

    Why it's wrong here

    A Processing job runs a single pass of the script rather than epochs, so there is no per-epoch cleanup point to hook into. Deleting input files mid-script would also break subsequent reads and is not a supported pattern. This does not address the volume capacity problem at its source.

  • ✗

    Raise the max_runtime_in_seconds parameter and enable network isolation on the processing job.

    Why it's wrong here

    max_runtime_in_seconds only controls how long the job may run before being stopped, and network isolation restricts outbound access. Neither changes the amount of local disk consumed by downloaded input, so the out-of-disk error would persist. These parameters are unrelated to storage capacity.

  • ✓

    Increase the processing instance's volume size and use ShardedByS3Key so each instance downloads only its share of objects.

    Why this is correct

    The failure is caused by the local volume filling with downloaded input data. Enlarging the attached EBS volume provides headroom, and ShardedByS3Key distributes the S3 objects across instances so no single instance downloads the entire dataset. Together these directly address the root cause rather than merely retrying the job.

  • ✗

    Switch the input channel to Pipe mode so the container streams records instead of downloading files.

    Why it's wrong here

    Pipe mode is available to SageMaker training jobs, not to Processing jobs, whose inputs are mounted as files under /opt/ml/processing/input. Choosing it reflects a misunderstanding of the Processing container contract and would not change how the script accesses data. It therefore cannot resolve the disk exhaustion.

Visual reference

Client Recursive Resolver Root DNS (13 root servers) TLD DNS (.com, .org, …) Authoritative example.com query IP addr answer

Quick reference

AWS S3 Storage Class Comparison

Storage ClassMin DurationRetrievalUse Case
S3 StandardNoneImmediateFrequently accessed data
S3 Standard-IA30 daysImmediateInfrequent access, rapid retrieval
S3 One Zone-IA30 daysImmediateNon-critical infrequent data
S3 Intelligent-TieringNoneImmediate–hoursUnknown or changing access patterns
S3 Glacier Instant90 daysMillisecondsArchive with instant retrieval
S3 Glacier Flexible90 daysMinutes–hoursArchive, flexible retrieval
S3 Glacier Deep Archive180 daysHoursLong-term compliance archive

About these practice questions

This MLA-C01 question is part of Courseiva's 665-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official Amazon Web Services exam blueprint

This MLA-C01 practice question is part of Courseiva's free Amazon Web Services certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the MLA-C01 exam.