MLA-C01 Deployment and Orchestration of ML Workflows Practice Question
A retailer runs a nightly batch scoring job that processes 40 GB of transaction data and writes predictions to S3. Occasionally a single partition is corrupt, causing the entire job to fail after several hours. The team wants the job to skip the corrupt partition, log which partition failed, and still complete processing of the remaining data with minimal changes to their existing SageMaker Processing job. Which change should they make?
⚠ Common exam trap
The trap here is reaching for infrastructure-level retries or resizing, when a deterministic data corruption must be handled in application code to isolate the bad partition.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Wrap the per-partition processing logic in try/except inside the Processing container script, record the failing partition key to CloudWatch Logs, and continue to the next partition.
Adding per-partition exception handling inside the Processing container isolates corruption to the affected partition, logs its key for follow-up, and allows the job to finish the rest of the data. This is a localized code change that preserves the existing SageMaker Processing job structure and S3 outputs. It meets fault isolation, failure visibility, and completion requirements with the least disruption.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✓
Wrap the per-partition processing logic in try/except inside the Processing container script, record the failing partition key to CloudWatch Logs, and continue to the next partition.
Why this is correct
Handling exceptions per partition inside the container lets the job skip a corrupt input, emit the partition key to CloudWatch Logs, and proceed with the remaining data. It requires only a code change to the existing Processing job, preserving the current orchestration and S3 output path. This directly satisfies fault isolation, failure logging, and job completion with minimal disruption.
- ✗
Configure the Processing job with a retry policy and a maximum retry count of three so transient failures are retried automatically.
Why it's wrong here
A retry policy re-runs the entire job, and a deterministic data corruption error will fail again on each attempt. It does not isolate the bad partition, does not log which partition is corrupt, and wastes hours of compute per retry. Maximum retries are also capped by the service, so this cannot guarantee completion across many bad partitions.
- ✗
Convert the Processing job to a SageMaker Training job with checkpointing enabled so it can resume after the corrupt partition.
Why it's wrong here
Training job checkpointing saves model state for resumption, not data partition progress, and does not skip corrupt inputs. Converting a batch scoring workflow into a training job is a major architectural change that conflicts with the minimal-change requirement. It also does not produce per-partition failure logs, so it fails the logging and fault-isolation goals.
- ✗
Increase the Processing job's instance count and volume size so the corrupt partition is retried on a different instance.
Why it's wrong here
Adding instances or storage does not make a corrupt partition valid; the same parsing error will occur on whichever instance processes it. The job would still fail unless the code handles the exception. This increases cost without addressing fault isolation, and it does not produce the required log of which partition failed, so it does not meet the stated goal.
Quick reference
AWS S3 Storage Class Comparison
| Storage Class | Min Duration | Retrieval | Use Case |
|---|---|---|---|
| S3 Standard | None | Immediate | Frequently accessed data |
| S3 Standard-IA | 30 days | Immediate | Infrequent access, rapid retrieval |
| S3 One Zone-IA | 30 days | Immediate | Non-critical infrequent data |
| S3 Intelligent-Tiering | None | Immediate–hours | Unknown or changing access patterns |
| S3 Glacier Instant | 90 days | Milliseconds | Archive with instant retrieval |
| S3 Glacier Flexible | 90 days | Minutes–hours | Archive, flexible retrieval |
| S3 Glacier Deep Archive | 180 days | Hours | Long-term compliance archive |
Go deeper
Related to this question
About these practice questions
Courseiva writes every MLA-C01 question from scratch — 665 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official Amazon Web Services exam blueprint
This MLA-C01 practice question is part of Courseiva's free Amazon Web Services certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the MLA-C01 exam.