Courseiva

MLA-C01 Deployment and Orchestration of ML Workflows Practice Question

A financial services company needs to deploy a machine learning model for real-time fraud detection. The model must be highly available across multiple Availability Zones and must support automatic scaling based on request volume. The company also needs to perform canary deployments to test new model versions with a small percentage of traffic before full rollout. Which SageMaker feature should they use?

⚠ Common exam trap

AWS often tests the distinction between real-time endpoints with production variants and Multi-Model Endpoints, where candidates mistakenly think Multi-Model Endpoints support canary deployments because they can host multiple models, but they lack traffic splitting and weighted routing capabilities.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

SageMaker real-time endpoint with production variants

SageMaker real-time endpoints with production variants enable canary deployments by routing a small percentage of traffic to a new model version while the majority goes to the current version. This feature also supports multi-AZ deployment for high availability and automatic scaling based on request volume via Application Auto Scaling, meeting all the stated requirements.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✓

    SageMaker real-time endpoint with production variants

    Why this is correct

    A SageMaker real-time endpoint with production variants hosts multiple model versions behind one endpoint and supports weighted traffic splitting, enabling canary deployments. Multi-AZ deployment and automatic scaling satisfy the availability and request-volume scaling constraints in the stem.

  • ✗

    SageMaker Multi-Model Endpoint

    Why it's wrong here

    Multi-Model Endpoint hosts many models behind one endpoint to cut hosting cost, and it lacks native traffic-splitting for canary rollouts. The option tempts because it serves multiple models concurrently, but the requirement here is a production variant with weighted traffic distribution, which is the production variant feature.

  • ✗

    SageMaker Batch Transform

    Why it's wrong here

    Batch Transform runs offline scoring jobs against stored datasets and provides no real-time endpoint, no multi-AZ hosting and no traffic splitting. The option tempts because it processes fraud data in bulk, but the scenario demands low-latency inference with canary deployments, which requires a real-time endpoint.

  • ✗

    SageMaker Serverless Inference

    Why it's wrong here

    Serverless Inference scales to zero and cannot span multiple Availability Zones, so it fails the high-availability requirement and offers no canary traffic splitting. The option tempts because it scales automatically with request volume, but that scaling is cold-start based and single-AZ only.

About these practice questions

One of 665 original MLA-C01 practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

Same concept, more angles

2 more ways this is tested on MLA-C01

These questions test the same concept from different angles. Work through them to make sure you can recognise it however the exam phrases it.

Variation 1. A company needs to deploy a new model version to a SageMaker real-time endpoint. They want to route 5% of traffic to the new version initially to monitor for errors before full rollout. Which deployment strategy should they use?

medium
  • A.Blue/green deployment
  • B.Shadow testing
  • ✓ C.Canary deployment with production variants
  • D.Multi-model endpoint

Why C: A canary deployment with production variants allows you to route a specific percentage of traffic (e.g., 5%) to the new model version by adjusting the `InitialVariantWeight` parameter in the production variant configuration. This enables gradual traffic shifting while monitoring errors, and you can later increase the weight to 100% for full rollout. SageMaker real-time endpoints support this natively by hosting multiple model variants behind the same endpoint.

Variation 2. A company uses SageMaker to deploy a model and wants to perform A/B testing by splitting traffic between two model variants. Which TWO actions should they take? (Select TWO.)

medium
  • ✓ A.Configure two production variants on the endpoint, each with an initial weight
  • B.Use SageMaker Model Registry to approve both variants
  • ✓ C.Use the UpdateEndpointWeightsAndCapacity API to adjust traffic after analysis
  • D.Deploy each variant to a separate endpoint and use Route53 weighted routing
  • E.Enable shadow testing on the endpoint

Why A: Option A is correct because SageMaker A/B testing is implemented by configuring a single endpoint with two production variants, each assigned an InitialVariantWeight so the endpoint can split inference traffic between them. Option C is correct because after analyzing the variants' performance, you can shift traffic by calling UpdateEndpointWeightsAndCapacity to change each variant's VariantWeight and DesiredInstanceCount without redeploying the endpoint. Option B is not required: Model Registry approval is a governance step and does not itself create or split endpoint traffic. Option D is wrong because deploying separate endpoints with Route53 weighted routing is a DNS-level workaround, not SageMaker's native A/B testing mechanism. Option E is wrong because shadow testing mirrors live traffic to a new variant for evaluation without serving its responses, so it does not split production traffic between two variants.

JA

Written by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

This MLA-C01 practice question is part of Courseiva's free Amazon Web Services certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the MLA-C01 exam.