MLA-C01 Data Preparation for Machine Learning Practice Question
A machine learning engineer is using Amazon SageMaker Processing to preprocess a large dataset. The processing job runs a custom Python script that uses the pandas library to read multiple CSV files from an S3 input prefix. The script must write the processed output to a different S3 prefix. Which configuration of the ProcessingInput and ProcessingOutput parameters is correct for this scenario?
⚠ Common exam trap
The trap here is mixing up the source and destination fields for ProcessingInput and ProcessingOutput; source always refers to the origin (S3 for input, local for output) and destination to the target (local for input, S3 for output).
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
ProcessingInput with source set to the S3 input prefix and destination set to '/opt/ml/processing/input'; ProcessingOutput with source set to '/opt/ml/processing/output' and destination set to the S3 output prefix.
In SageMaker Processing, ProcessingInput specifies the S3 source and the local destination path where the data will be mounted in the container. ProcessingOutput specifies the local source path where the script writes output and the S3 destination where it will be uploaded. This bidirectional mapping ensures that the container can access input data and persist output data to S3. The correct configuration uses S3 URIs for external locations and local paths for container locations.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
ProcessingInput with source set to the S3 input prefix and destination set to the S3 output prefix; ProcessingOutput with source set to the S3 input prefix and destination set to '/opt/ml/processing/output'.
Why it's wrong here
This configuration incorrectly uses S3 prefixes for local destinations and vice versa. ProcessingInput destination must be a local path inside the container, not an S3 URI. ProcessingOutput source must be a local path, not an S3 URI. This would result in errors because the container expects local filesystem paths for input and output.
- ✗
ProcessingInput with source set to '/opt/ml/processing/input' and destination set to '/opt/ml/processing/output'; ProcessingOutput with source set to the S3 input prefix and destination set to the S3 output prefix.
Why it's wrong here
This option uses local paths for both source and destination of ProcessingInput, which is invalid because the input data resides in S3. ProcessingInput source must be the S3 URI. Also, the ProcessingOutput source should be a local path, but here it is set to an S3 prefix, which is incorrect. The job would not be able to locate the input data.
- ✓
ProcessingInput with source set to the S3 input prefix and destination set to '/opt/ml/processing/input'; ProcessingOutput with source set to '/opt/ml/processing/output' and destination set to the S3 output prefix.
Why this is correct
This configuration correctly maps the S3 input prefix to a local directory inside the processing container and the local output directory back to S3. The source for ProcessingInput is the S3 URI, and destination is the local path where the data will be available. For ProcessingOutput, source is the local path where the script writes outputs, and destination is the S3 URI. This is the standard pattern for SageMaker Processing jobs.
- ✗
ProcessingInput with source set to '/opt/ml/processing/input' and destination set to the S3 input prefix; ProcessingOutput with source set to the S3 output prefix and destination set to '/opt/ml/processing/output'.
Why it's wrong here
This reverses the correct mapping. ProcessingInput source must be the S3 location (the external data source), and destination must be the local container path. Similarly, ProcessingOutput source is the local path where outputs are written, and destination is the S3 location. Reversing these would cause the job to fail because the container cannot read from a local path that doesn't exist initially.
Quick reference
AWS S3 Storage Class Comparison
| Storage Class | Min Duration | Retrieval | Use Case |
|---|---|---|---|
| S3 Standard | None | Immediate | Frequently accessed data |
| S3 Standard-IA | 30 days | Immediate | Infrequent access, rapid retrieval |
| S3 One Zone-IA | 30 days | Immediate | Non-critical infrequent data |
| S3 Intelligent-Tiering | None | Immediate–hours | Unknown or changing access patterns |
| S3 Glacier Instant | 90 days | Milliseconds | Archive with instant retrieval |
| S3 Glacier Flexible | 90 days | Minutes–hours | Archive, flexible retrieval |
| S3 Glacier Deep Archive | 180 days | Hours | Long-term compliance archive |
Go deeper
Related to this question
About these practice questions
One of 665 original MLA-C01 practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official Amazon Web Services exam blueprint
This MLA-C01 practice question is part of Courseiva's free Amazon Web Services certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the MLA-C01 exam.