MLA-C01 Deployment and Orchestration of ML Workflows Practice Question
A team runs a SageMaker Pipeline that trains a model and registers it in the Model Registry. Compliance requires that the pipeline run automatically every time new labeled data lands in S3, and that each run record the exact S3 data prefix, the training image URI, and the git commit hash as lineage metadata. The engineer wants the least operational overhead. Which approach meets these requirements?
⚠ Common exam trap
The trap here is treating a scheduled poll or cached pipeline as equivalent to event-driven execution, which misses both trigger immediacy and the lineage metadata requirements.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Configure an S3 event notification to invoke a Lambda function that calls StartPipelineExecution with the data prefix and commit hash passed as pipeline parameters.
S3 event notifications driving a Lambda that invokes StartPipelineExecution provides immediate, event-driven execution with parameter overrides for the data prefix and commit hash. Those parameters propagate into pipeline steps and are captured as lineage metadata in the Model Registry, satisfying compliance. It avoids polling, scheduling waste, and manual gates, giving the lowest operational overhead of the listed designs.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Enable SageMaker Data Wrangler scheduled jobs to export data and manually start the pipeline after reviewing the export in SageMaker Studio.
Why it's wrong here
Manual start after human review defeats the requirement for automatic execution on new data and adds significant operational overhead. Data Wrangler is a data preparation tool, not a pipeline trigger, and its scheduled exports do not propagate the git commit hash or image URI into the Model Registry lineage. This approach also creates a human bottleneck that scales poorly with data arrival frequency.
- ✓
Configure an S3 event notification to invoke a Lambda function that calls StartPipelineExecution with the data prefix and commit hash passed as pipeline parameters.
Why this is correct
S3 event notifications can trigger Lambda on object creation, and the Lambda handler can call StartPipelineExecution with parameter overrides carrying the data prefix and commit hash. Pipeline parameters flow into processing and training steps, where they are recorded as lineage metadata via the Model Registry. This is event-driven, requires no polling infrastructure, and keeps operational overhead low while satisfying automatic triggering and full lineage capture.
- ✗
Create an EventBridge scheduled rule that runs the pipeline every 15 minutes and relies on the pipeline's caching to skip runs when data is unchanged.
Why it's wrong here
A fixed schedule is not event-driven and introduces latency up to the interval, and it wastes executions when no new data exists. Pipeline caching skips step execution based on input signatures, but it does not prevent StartPipelineExecution calls and does not, by itself, capture the git commit hash or S3 prefix as lineage metadata. This adds unnecessary orchestration overhead and does not guarantee a run per data arrival.
- ✗
Use AWS Step Functions with a Wait state polling S3 every minute, then call StartPipelineExecution once new objects are detected.
Why it's wrong here
Polling every minute is inefficient compared to native S3 event notifications and requires maintaining a Step Functions state machine plus custom comparison logic for object listings. It adds latency and cost, and the Step Functions definition would still need to pass the data prefix and commit hash into the pipeline. This is more operational overhead than an event-driven Lambda trigger for the same outcome.
Quick reference
AWS S3 Storage Class Comparison
| Storage Class | Min Duration | Retrieval | Use Case |
|---|---|---|---|
| S3 Standard | None | Immediate | Frequently accessed data |
| S3 Standard-IA | 30 days | Immediate | Infrequent access, rapid retrieval |
| S3 One Zone-IA | 30 days | Immediate | Non-critical infrequent data |
| S3 Intelligent-Tiering | None | Immediate–hours | Unknown or changing access patterns |
| S3 Glacier Instant | 90 days | Milliseconds | Archive with instant retrieval |
| S3 Glacier Flexible | 90 days | Minutes–hours | Archive, flexible retrieval |
| S3 Glacier Deep Archive | 180 days | Hours | Long-term compliance archive |
Go deeper
Related to this question
About these practice questions
Courseiva writes every MLA-C01 question from scratch — 665 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official Amazon Web Services exam blueprint
This MLA-C01 practice question is part of Courseiva's free Amazon Web Services certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the MLA-C01 exam.