Courseiva

MLA-C01 Data Preparation for Machine Learning Practice Question

A machine learning engineer is using Amazon SageMaker Processing to preprocess a large dataset. The processing job runs a custom Python script that uses the pandas library to read multiple CSV files from an S3 input prefix. The script must write the processed output to a different S3 prefix. Which configuration of the ProcessingInput and ProcessingOutput parameters is correct for this scenario?

⚠ Common exam trap

The trap here is mixing up the source and destination fields for ProcessingInput and ProcessingOutput; source always refers to the origin (S3 for input, local for output) and destination to the target (local for input, S3 for output).

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

ProcessingInput with source set to the S3 input prefix and destination set to '/opt/ml/processing/input'; ProcessingOutput with source set to '/opt/ml/processing/output' and destination set to the S3 output prefix.

In SageMaker Processing, ProcessingInput specifies the S3 source and the local destination path where the data will be mounted in the container. ProcessingOutput specifies the local source path where the script writes output and the S3 destination where it will be uploaded. This bidirectional mapping ensures that the container can access input data and persist output data to S3. The correct configuration uses S3 URIs for external locations and local paths for container locations.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    ProcessingInput with source set to the S3 input prefix and destination set to the S3 output prefix; ProcessingOutput with source set to the S3 input prefix and destination set to '/opt/ml/processing/output'.

    Why it's wrong here

    This configuration incorrectly uses S3 prefixes for local destinations and vice versa. ProcessingInput destination must be a local path inside the container, not an S3 URI. ProcessingOutput source must be a local path, not an S3 URI. This would result in errors because the container expects local filesystem paths for input and output.

  • ✗

    ProcessingInput with source set to '/opt/ml/processing/input' and destination set to '/opt/ml/processing/output'; ProcessingOutput with source set to the S3 input prefix and destination set to the S3 output prefix.

    Why it's wrong here

    This option uses local paths for both source and destination of ProcessingInput, which is invalid because the input data resides in S3. ProcessingInput source must be the S3 URI. Also, the ProcessingOutput source should be a local path, but here it is set to an S3 prefix, which is incorrect. The job would not be able to locate the input data.

  • ✓

    ProcessingInput with source set to the S3 input prefix and destination set to '/opt/ml/processing/input'; ProcessingOutput with source set to '/opt/ml/processing/output' and destination set to the S3 output prefix.

    Why this is correct

    This configuration correctly maps the S3 input prefix to a local directory inside the processing container and the local output directory back to S3. The source for ProcessingInput is the S3 URI, and destination is the local path where the data will be available. For ProcessingOutput, source is the local path where the script writes outputs, and destination is the S3 URI. This is the standard pattern for SageMaker Processing jobs.

  • ✗

    ProcessingInput with source set to '/opt/ml/processing/input' and destination set to the S3 input prefix; ProcessingOutput with source set to the S3 output prefix and destination set to '/opt/ml/processing/output'.

    Why it's wrong here

    This reverses the correct mapping. ProcessingInput source must be the S3 location (the external data source), and destination must be the local container path. Similarly, ProcessingOutput source is the local path where outputs are written, and destination is the S3 location. Reversing these would cause the job to fail because the container cannot read from a local path that doesn't exist initially.

Quick reference

AWS S3 Storage Class Comparison

Storage ClassMin DurationRetrievalUse Case
S3 StandardNoneImmediateFrequently accessed data
S3 Standard-IA30 daysImmediateInfrequent access, rapid retrieval
S3 One Zone-IA30 daysImmediateNon-critical infrequent data
S3 Intelligent-TieringNoneImmediate–hoursUnknown or changing access patterns
S3 Glacier Instant90 daysMillisecondsArchive with instant retrieval
S3 Glacier Flexible90 daysMinutes–hoursArchive, flexible retrieval
S3 Glacier Deep Archive180 daysHoursLong-term compliance archive

About these practice questions

One of 665 original MLA-C01 practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official Amazon Web Services exam blueprint

This MLA-C01 practice question is part of Courseiva's free Amazon Web Services certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the MLA-C01 exam.