Courseiva

CCNA Deployment and Orchestration of ML Workflows Questions

75 of 94 questions · Page 1/2 · Deployment and Orchestration of ML Workflows · Answers revealed

1
MCQhard

A team needs to deploy a model that has compliance requirements to log all inference requests and responses for auditing. The model will be served using a real-time endpoint. How can they achieve this without custom code?

A.Enable SageMaker Data Capture on the endpoint
B.Add a custom Lambda function using a container
C.Use SageMaker Debugger to monitor inference
D.Enable CloudTrail for the endpoint
AnswerA

SageMaker Data Capture on a real-time endpoint automatically logs inference requests and responses to Amazon S3, satisfying the audit requirement without custom code. It captures input and output payloads at the endpoint level, so compliance logging happens transparently for every invocation.

Why this answer

SageMaker Data Capture is the native, no-code feature that automatically logs inference requests and responses for real-time endpoints. It captures payload data to an S3 bucket without requiring any custom code, directly meeting the compliance requirement for audit logging.

Exam trap

The trap here is that candidates often confuse CloudTrail (which logs API calls) with Data Capture (which logs payloads), or they assume Debugger can be repurposed for inference logging, but Debugger only works during training.

How to eliminate wrong answers

Option B is wrong because adding a custom Lambda function using a container introduces custom code, which the question explicitly states should be avoided. Option C is wrong because SageMaker Debugger is designed for monitoring training jobs and debugging model performance, not for capturing inference request/response logs for auditing. Option D is wrong because AWS CloudTrail logs API calls to the SageMaker endpoint (e.g., InvokeEndpoint actions) but does not capture the actual inference request and response payloads.

2
Multi-Selecthard

An MLOps team is designing a SageMaker Pipeline to automate model retraining. The pipeline must: (1) run training only if new training data is available, (2) register the model in SageMaker Model Registry only if evaluation metrics exceed a threshold, (3) deploy the approved model to a staging endpoint automatically. Which THREE steps should they include? (Choose THREE.)

Select 3 answers
A.ConditionStep to check if evaluation metrics exceed the threshold
B.RegisterModel step to register the model in the Model Registry
C.TuningStep to perform hyperparameter optimization
D.TransformStep to deploy the model to a staging endpoint
E.ProcessingStep to check for new training data availability
AnswersA, B, E

ConditionStep evaluates a property from the evaluation step against a threshold and branches accordingly, so RegisterModel executes only on the true branch. This enforces the stem's requirement that registration happens only when evaluation metrics exceed the threshold.

Why this answer

Option E (ProcessingStep to check for new training data availability) is correct because a ProcessingStep can run a script or job that inspects the input data source and outputs a boolean/property indicating whether new data exists, which is exactly what requirement (1) needs before training runs. Option A (ConditionStep to check if evaluation metrics exceed the threshold) is correct because ConditionStep is the pipeline construct that evaluates a condition (e.g., a metric from a previous evaluation step against a threshold) and gates downstream steps, satisfying requirement (2). Option B (RegisterModel step to register the model in the Model Registry) is correct because RegisterModel is the dedicated SageMaker Pipelines step that creates a model package in the Model Registry, which is required to register the model once the metric condition passes.

Option C (TuningStep) is not correct because hyperparameter optimization is not required by any of the stated requirements. Option D (TransformStep) is not correct because TransformStep runs batch transform inference jobs, not endpoint deployment; deploying to a staging endpoint would use a different mechanism such as a LambdaStep or a model deployment step, so it does not satisfy requirement (3).

Exam trap

MLA-C01 often tests the misconception that TransformStep deploys models to endpoints, when it actually runs batch transform jobs; endpoint deployment requires a separate deployment step or Lambda invocation.

3
MCQmedium

A company wants to run inference on a large dataset stored in S3 using a pre-trained model. The inference can tolerate latency from minutes to hours, and they want a fully managed solution that autoscales to handle large volumes. Which SageMaker inference option is most suitable?

A.Batch transform
B.Real-time endpoint
C.Asynchronous inference
D.Serverless inference
AnswerA

Batch transform runs asynchronous inference over an entire S3 dataset, writing predictions back to S3, so latency of minutes to hours is acceptable. It provisions and autoscales compute for the job, then terminates it, satisfying the fully managed requirement without a persistent endpoint.

Why this answer

Batch transform is the most suitable option because the company needs to run inference on a large dataset stored in S3 with latency tolerance from minutes to hours, and requires a fully managed, autoscaling solution. SageMaker Batch Transform processes the entire dataset as a single job, automatically provisions and scales compute resources, and writes results back to S3 without the need for persistent endpoints.

Exam trap

AWS often tests the distinction between latency tolerance and workload type, where candidates mistakenly choose asynchronous inference because they see 'tolerates latency' and 'autoscaling' without recognizing that batch transform is the only option designed for processing entire datasets stored in S3 as a single job, not individual requests.

How to eliminate wrong answers

Option B (Real-time endpoint) is wrong because it is designed for low-latency, synchronous inference (milliseconds to seconds) and requires a persistent endpoint that incurs ongoing costs, not suitable for large batch jobs with hour-long latency tolerance. Option C (Asynchronous inference) is wrong because it is intended for near-real-time requests with payloads up to 1 GB and queues requests for processing, but it still maintains a persistent endpoint and is optimized for workloads with latency in seconds to minutes, not hours-long batch processing of large datasets. Option D (Serverless inference) is wrong because it is designed for intermittent, short-lived inference requests with automatic scaling to zero, but it has a maximum timeout of 15 minutes per request and is not suitable for processing large datasets as a single job.

4
MCQeasy

A data scientist wants to compare the performance of two model versions (V1 and V2) in production by splitting traffic between them. They want to gradually increase the percentage of traffic to the new version while monitoring metrics. Which SageMaker feature enables this?

A.SageMaker production variants with traffic splitting
B.SageMaker shadow testing
C.SageMaker blue/green deployment
D.SageMaker canary deployment
AnswerA

SageMaker production variants let a single endpoint host multiple model versions with weighted traffic distribution, so the data scientist can shift percentages gradually while monitoring metrics. This satisfies the stem's requirement for controlled A/B comparison and progressive rollout between V1 and V2.

Why this answer

SageMaker production variants with traffic splitting let you host multiple model versions behind a single endpoint and assign a percentage of invocations to each variant. To compare V1 and V2 and gradually shift traffic, you create an endpoint configuration with two production variants and set the initial traffic distribution (e.g., 90/10), then update the weights as you gain confidence. This is the native SageMaker mechanism for A/B comparison and gradual rollout.

Exam trap

The trap is that 'canary deployment' sounds like a distinct SageMaker feature, but in SageMaker it is implemented via production variants and traffic splitting, so candidates who pick the canary-named option miss the actual configuration mechanism.

How to eliminate wrong answers

Option B is wrong because shadow testing mirrors production traffic to a new variant without returning its responses to callers, so it cannot be used to serve a percentage of real traffic to V2. Option C is wrong because blue/green deployment shifts all traffic from the old fleet to the new fleet at once (with a bake period), not a gradual percentage-based split. Option D is wrong because SageMaker does not expose a feature literally named 'canary deployment'; canary behavior is implemented using production variants and traffic splitting, so this option describes the outcome rather than the feature.

5
MCQeasy

A data scientist needs to deploy a single ML model that will serve real-time predictions with low latency (under 10 ms) for a high-traffic web application. The model fits in memory and requires GPU acceleration. Which SageMaker inference option is MOST suitable?

A.Real-time endpoint on ml.m5 instances
B.Batch Transform
C.Real-time endpoint on ml.g4dn instances
D.Serverless Inference
AnswerC

A real-time endpoint on ml.g4dn instances provides GPU acceleration with persistent, low-latency inference, satisfying the sub-10 ms requirement for a high-traffic web application. Serverless inference lacks GPU support and cold starts, while batch transform cannot serve real-time requests.

Why this answer

Real-time endpoints on GPU instances (ml.g4dn) provide low latency and GPU acceleration, ideal for high-traffic, latency-sensitive workloads.

6
MCQeasy

A machine learning engineer has trained a model using SageMaker and wants to deploy it to a real-time endpoint. The engineer needs to specify the model artifacts, the inference code, and the environment. Which SageMaker resource should the engineer create first?

A.An endpoint configuration
B.A SageMaker pipeline
C.A SageMaker model
D.A SageMaker endpoint
AnswerC

A SageMaker model is the resource that encapsulates the model artifacts, inference code (as a Docker image), and environment variables. It is a prerequisite for creating an endpoint configuration and then an endpoint. Creating the model first is the correct initial step in the deployment process.

Why this answer

The deployment sequence in SageMaker begins with creating a model, which specifies the model artifacts and the inference container. This model is then referenced in an endpoint configuration, which defines the deployment settings. Finally, the endpoint is created from that configuration.

Therefore, the model is the first resource to create.

Exam trap

The trap here is thinking that an endpoint configuration or endpoint can be created without a model, but they both depend on the model resource.

7
Multi-Selectmedium

A team wants to deploy a new model using a canary deployment strategy on SageMaker. Which TWO configurations are necessary? (Choose two.)

Select 2 answers
A.Create a CloudWatch alarm to automatically rollback
B.Set the initial traffic distribution (e.g., 90% old, 10% new)
C.Enable data capture on the endpoint
D.Create a new endpoint configuration with two production variants, each pointing to a different model
E.Use SageMaker Model Registry to approve the new model
AnswersB, D

Setting the initial traffic distribution defines how much live inference traffic the new model variant receives from the outset. SageMaker's canary strategy requires this split to route a small percentage to the new variant while the remainder stays on the current one, satisfying the gradual, controlled rollout the stem demands.

Why this answer

Option B is correct because a SageMaker canary deployment is defined by specifying the initial traffic split between the existing (old) variant and the new variant, for example 90% to the old model and 10% to the new model, via the endpoint configuration's variant weights. Option D is correct because canary deployment requires a new endpoint configuration that contains two production variants, each referencing a different model, so traffic can be shifted between the old and new versions. Option A is not required, since CloudWatch alarms and automatic rollback are optional safeguards rather than mandatory canary configuration elements.

Option C is not required, as data capture is an optional monitoring feature and not part of the canary deployment definition. Option E is not required, because Model Registry approval is a governance step and not a necessary configuration for performing a canary deployment.

Exam trap

The trap is that several options are good practices (alarms, data capture, registry approval) but not necessary configurations, so candidates who conflate best practices with requirements select the wrong two.

8
MCQeasy

A data scientist wants to version and manage trained models, require approval before deployment, and enable cross-account deployment. Which SageMaker feature provides these capabilities?

A.SageMaker Neo
B.SageMaker Pipelines
C.Amazon Elastic Inference
D.SageMaker Model Registry
AnswerD

SageMaker Model Registry versions trained models through model package groups, enforces deployment approval via a manual approval status workflow, and supports cross-account deployment by sharing model packages with other AWS accounts. It directly satisfies all three stem requirements: versioning, approval gating, and cross-account deployment.

Why this answer

SageMaker Model Registry is the correct choice because it provides a centralized catalog for versioning trained models, supports approval workflows (e.g., pending, approved, rejected) to gate deployment, and enables cross-account deployment by sharing model package ARNs across AWS accounts via AWS Resource Access Manager (RAM) or cross-account IAM roles. This directly satisfies all three requirements: versioning, approval before deployment, and cross-account deployment.

Exam trap

The trap here is that candidates may confuse SageMaker Pipelines (which orchestrates the ML workflow) with Model Registry (which manages model versions and approvals), but Pipelines lacks native versioning and approval gatekeeping, while Model Registry is specifically designed for those governance tasks.

How to eliminate wrong answers

Option A is wrong because SageMaker Neo is a model optimization and compilation service that converts trained models into efficient runtime code for target hardware (e.g., ARM, Intel, NVIDIA), but it does not provide model versioning, approval workflows, or cross-account deployment capabilities. Option B is wrong because SageMaker Pipelines is a CI/CD orchestration service for building, training, and deploying ML pipelines, but it does not natively include a model registry with approval gates or cross-account deployment features; while it can integrate with Model Registry, Pipelines itself does not offer versioning or approval management. Option C is wrong because Amazon Elastic Inference is a service that attaches low-cost GPU-powered inference acceleration to SageMaker endpoints, but it has no role in model versioning, approval workflows, or cross-account deployment.

9
MCQhard

A company uses SageMaker Pipelines to orchestrate their ML workflow. They notice that if a pipeline step fails due to a transient error (e.g., a brief network issue), the entire pipeline fails and they must manually rerun from the beginning. They want to automatically retry failed steps a few times before failing. What should they do?

A.Use a Lambda function to catch step failures and re-invoke the step
B.Use the CreatePipelineExecution API with a flag to ignore failures
C.Configure a RetryPolicy in the pipeline step definition to specify the number of retry attempts and backoff
D.Use AWS Step Functions to orchestrate the workflow instead of SageMaker Pipelines
AnswerC

A RetryPolicy attached to the step definition lets SageMaker Pipelines automatically re-execute that step after a transient failure, using the configured max attempts and backoff interval, without restarting the whole pipeline. This directly satisfies the requirement to retry failed steps a few times before failing, removing manual reruns.

Why this answer

SageMaker Pipelines supports retry policies for steps. By setting a RetryPolicy in the step definition with a maximum number of retry attempts, the pipeline will automatically retry the step on failure. The other options do not achieve automatic retry within the pipeline: Step Functions would require rebuilding the pipeline, Lambda cannot retry pipeline steps, and CreatePipelineExecution does not handle retries.

10
Multi-Selectmedium

A company is using SageMaker to serve a model for real-time predictions. They want to test a new model version by routing a small percentage of live traffic to it while the rest goes to the current model. They also need to compare performance metrics. Which TWO actions should they take? (Select TWO.)

Select 2 answers
A.Deploy the new model to a separate endpoint and use Route 53 to split traffic
B.Compile the new model with SageMaker Neo before deployment
C.Use SageMaker Batch Transform to evaluate the new model
D.Monitor the performance of both variants using SageMaker CloudWatch metrics
E.Configure a production variant with the new model and set initial traffic weight to a small percentage
AnswersD, E

SageMaker publishes per-variant invocation, latency and error metrics to CloudWatch, enabling direct comparison of the new model against the current one. This satisfies the requirement to compare performance metrics across both variants during the live traffic test.

Why this answer

Amazon CloudWatch provides built-in metrics for SageMaker endpoints, including latency, invocation counts, and error rates, which can be monitored per production variant. This allows the company to compare the performance of the new model version against the current model in real time. Option E is correct because SageMaker endpoints support multiple production variants, and you can set an initial traffic weight (e.g., 5%) to route a small percentage of live traffic to the new model while the rest goes to the existing variant.

Exam trap

A common misconception is that you need an external load balancer or DNS service (like Route 53) to split traffic between model versions, but SageMaker's built-in production variant feature handles this natively.

11
MCQmedium

A machine learning engineer has trained a scikit-learn model and saved it as model.joblib in Amazon S3. The engineer wants SageMaker to host the model for real-time inference without writing a custom container or inference script, because the model uses only standard predict behavior. Which deployment approach should the engineer use?

A.Create a custom Docker image that installs scikit-learn, push it to Amazon ECR, and write an inference handler for the /invocations route.
B.Use SageMaker Batch Transform with the built-in Scikit-learn container and invoke the endpoint for each real-time request.
C.Deploy the model with the SageMaker Scikit-learn built-in framework container using the SageMaker Python SDK SKLearnModel class, pointing model_data to the S3 artifact.
D.Upload model.joblib to SageMaker Model Registry and rely on automatic endpoint creation when the model package is approved.
AnswerC

The Scikit-learn built-in framework container supports joblib and pickle artifacts and provides a default inference handler that loads the model and calls predict, so no custom script is needed. Supplying the S3 model artifact through SKLearnModel lets SageMaker extract the tarball and start a compliant real-time endpoint.

Why this answer

Because the model is a standard scikit-learn artifact and no custom inference logic is required, the built-in Scikit-learn framework container is the correct fit. Passing the S3 model artifact to the SKLearnModel class lets SageMaker load the joblib file and use the container's default predict handler, producing a real-time endpoint without custom code or image maintenance.

Exam trap

The trap here is assuming that any deployment requires a custom container, when SageMaker built-in framework containers already provide default inference handlers for supported libraries.

12
MCQmedium

An ML engineer needs to orchestrate a multi-step workflow that includes data preprocessing on Spark, model training on SageMaker, and deployment to a production endpoint. They require tight integration with other AWS services and the ability to add custom logic. Which AWS service should they use alongside SageMaker?

A.AWS Step Functions
B.AWS CloudFormation
C.SageMaker Pipelines
D.Amazon EventBridge
AnswerA

AWS Step Functions orchestrates the workflow as a state machine, invoking Spark preprocessing, SageMaker training jobs and endpoint deployment as discrete steps. Its native AWS service integrations plus Lambda and Activity tasks satisfy the requirement for custom logic, while retries, error handling and visual tracking coordinate the multi-step pipeline that SageMaker alone cannot sequence.

Why this answer

AWS Step Functions is the correct choice because it provides a serverless workflow orchestration service that can coordinate multi-step ML pipelines involving Spark on AWS Glue or EMR, SageMaker training jobs, and endpoint deployments. It offers tight integration with over 200 AWS services via direct SDK integrations, supports custom logic through Lambda functions, and includes built-in error handling, retries, and parallel execution — making it ideal for complex, heterogeneous ML workflows that extend beyond SageMaker's native capabilities.

Exam trap

The trap here is that candidates confuse SageMaker Pipelines (a SageMaker-native orchestrator) with a general-purpose orchestrator, overlooking the requirement for tight integration with non-SageMaker services like Spark and custom logic — Step Functions is the correct choice for heterogeneous, multi-service ML workflows.

How to eliminate wrong answers

Option B (AWS CloudFormation) is wrong because it is an Infrastructure as Code (IaC) service for provisioning and managing AWS resources declaratively, not a workflow orchestrator — it cannot sequence steps like 'run Spark job, then train model, then deploy endpoint' with conditional logic or dynamic state management. Option C (SageMaker Pipelines) is wrong because while it can orchestrate SageMaker-native steps (training, tuning, batch transform), it lacks direct integration with external services like Spark on EMR or Glue and cannot easily incorporate custom logic outside the SageMaker ecosystem — the question explicitly requires tight integration with other AWS services and custom logic beyond SageMaker. Option D (Amazon EventBridge) is wrong because it is an event bus service for routing events between services based on rules, not a workflow orchestrator — it cannot manage sequential dependencies, retries, or stateful execution of a multi-step pipeline.

13
MCQhard

A team runs a SageMaker Pipeline that trains a model and registers it in the Model Registry. Compliance requires that the pipeline run automatically every time new labeled data lands in S3, and that each run record the exact S3 data prefix, the training image URI, and the git commit hash as lineage metadata. The engineer wants the least operational overhead. Which approach meets these requirements?

A.Enable SageMaker Data Wrangler scheduled jobs to export data and manually start the pipeline after reviewing the export in SageMaker Studio.
B.Configure an S3 event notification to invoke a Lambda function that calls StartPipelineExecution with the data prefix and commit hash passed as pipeline parameters.
C.Create an EventBridge scheduled rule that runs the pipeline every 15 minutes and relies on the pipeline's caching to skip runs when data is unchanged.
D.Use AWS Step Functions with a Wait state polling S3 every minute, then call StartPipelineExecution once new objects are detected.
AnswerB

S3 event notifications can trigger Lambda on object creation, and the Lambda handler can call StartPipelineExecution with parameter overrides carrying the data prefix and commit hash. Pipeline parameters flow into processing and training steps, where they are recorded as lineage metadata via the Model Registry. This is event-driven, requires no polling infrastructure, and keeps operational overhead low while satisfying automatic triggering and full lineage capture.

Why this answer

S3 event notifications driving a Lambda that invokes StartPipelineExecution provides immediate, event-driven execution with parameter overrides for the data prefix and commit hash. Those parameters propagate into pipeline steps and are captured as lineage metadata in the Model Registry, satisfying compliance. It avoids polling, scheduling waste, and manual gates, giving the lowest operational overhead of the listed designs.

Exam trap

The trap here is treating a scheduled poll or cached pipeline as equivalent to event-driven execution, which misses both trigger immediacy and the lineage metadata requirements.

14
MCQeasy

A machine learning engineer needs to optimize a trained TensorFlow model for deployment on edge devices with limited compute. Which SageMaker feature should they use to compile the model for target hardware?

A.SageMaker Model Monitor
B.SageMaker Neo
C.SageMaker Debugger
D.SageMaker Elastic Inference
AnswerB

SageMaker Neo compiles trained models into optimised executables for specific target hardware, reducing compute and memory footprint on constrained edge devices. It satisfies the stem's requirement to compile a TensorFlow model for the target edge hardware without manual re-engineering.

Why this answer

SageMaker Neo is the correct choice because it is specifically designed to compile trained machine learning models into an optimized format for target hardware architectures, such as ARM, Intel, or NVIDIA, enabling efficient inference on edge devices with limited compute resources. It uses a compiler to apply hardware-specific optimizations like operator fusion and memory layout tuning, reducing latency and memory footprint without requiring manual code changes.

Exam trap

The trap here is that candidates confuse SageMaker Neo with SageMaker Elastic Inference, mistakenly thinking Elastic Inference compiles models for edge devices, when in fact Elastic Inference only accelerates cloud inference by attaching a fractional GPU and does not perform compilation or target edge hardware.

How to eliminate wrong answers

Option A is wrong because SageMaker Model Monitor is used for detecting data drift and model quality degradation in production, not for compiling or optimizing models for hardware. Option C is wrong because SageMaker Debugger is a tool for monitoring training jobs, capturing tensors and metrics to debug issues like vanishing gradients, not for post-training compilation or hardware-specific optimization. Option D is wrong because SageMaker Elastic Inference attaches a separate accelerator to an endpoint for low-cost GPU acceleration, but it does not compile or optimize the model for edge hardware; it is a runtime acceleration service for cloud inference, not for edge deployment.

15
MCQmedium

An ML engineer wants to use MLflow on SageMaker to track experiments and log metrics. They have set up MLflow on an EC2 instance. How can they best integrate MLflow tracking with SageMaker training jobs?

A.Install MLflow on the SageMaker notebook instance only
B.Use the SageMaker Experiments integration with MLflow
C.Set the MLFLOW_TRACKING_URI environment variable in the training job and use the mlflow library in the training script
D.Use SageMaker Processing to run MLflow after training
AnswerC

Setting MLFLOW_TRACKING_URI in the training job's environment points the mlflow client inside the training container at the EC2-hosted tracking server, so metrics and artefacts log remotely during training. This satisfies the requirement to integrate the existing MLflow setup without altering SageMaker's managed training infrastructure.

Why this answer

The ML engineer can set the `MLFLOW_TRACKING_URI` environment variable in the SageMaker training job definition and use the `mlflow` library inside the training script to log parameters, metrics, and artifacts directly to the MLflow tracking server running on the EC2 instance. This approach allows the training job to communicate with the external MLflow server over HTTP/HTTPS without requiring any additional SageMaker integrations.

Exam trap

A common misconception is that SageMaker Experiments is required for tracking with MLflow, but the correct approach is to directly configure the MLflow tracking URI and use the mlflow library in the training script, as SageMaker does not natively block outbound HTTP connections to an external MLflow server.

How to eliminate wrong answers

Option A is wrong because installing MLflow only on the SageMaker notebook instance does not enable tracking from SageMaker training jobs, which run on separate, ephemeral compute instances that do not have access to the notebook instance's local MLflow server. Option B is wrong because SageMaker Experiments is a separate tracking service that does not natively integrate with an external MLflow server; using it would require additional custom code to bridge the two systems, and it does not replace the need to set the tracking URI. Option D is wrong because SageMaker Processing is designed for data preprocessing and postprocessing, not for real-time metric logging during training; running MLflow after training would miss the ability to log metrics incrementally during the training run.

16
MCQmedium

A startup wants to deploy a model that has variable traffic patterns, with some periods of no traffic and occasional spikes. They want to pay only for what they use and do not want to manage instances. Which SageMaker inference option should they choose?

A.Batch transform
B.Real-time endpoint with auto-scaling
C.Serverless inference
D.Multi-model endpoint
AnswerC

Serverless inference provisions compute automatically and scales to zero during idle periods, so the startup pays only per invocation and never manages instances. This matches the variable, spiky traffic pattern and the stated requirement to avoid instance management entirely.

Why this answer

Serverless inference is the correct choice because it automatically scales to zero during periods of no traffic and scales up to handle spikes, charging only for the compute time used. This eliminates the need to manage underlying instances, making it ideal for variable and intermittent traffic patterns.

Exam trap

The trap here is that candidates often confuse auto-scaling with the ability to scale to zero, but real-time endpoints with auto-scaling still maintain a minimum number of instances, incurring costs during idle periods, whereas serverless inference truly scales to zero.

How to eliminate wrong answers

Option A is wrong because batch transform is designed for offline, asynchronous predictions on large datasets, not for real-time or variable traffic patterns with occasional spikes. Option B is wrong because real-time endpoints with auto-scaling still require provisioning and managing underlying instances, and they cannot scale to zero, meaning you incur costs even during no traffic. Option D is wrong because multi-model endpoints reduce hosting costs by sharing instances across models but still require managing instances and cannot scale to zero, so you pay for idle capacity.

17
MCQmedium

A machine learning team needs to deploy a new model version for A/B testing, gradually shifting traffic from the old version to the new version over 24 hours. Which deployment strategy should they use?

A.Blue/green deployment
B.Shadow testing
C.Direct deployment with immediate full traffic
D.Canary deployment
AnswerD

Canary deployment routes a small percentage of live traffic to the new model version first, then incrementally increases that share as metrics stay healthy. This directly satisfies the stem's requirement to shift traffic gradually over 24 hours while limiting blast radius if the new version underperforms.

Why this answer

Canary deployment is the correct strategy because it allows gradual traffic shifting from the old model version to the new one over a specified time period (e.g., 24 hours) while monitoring for errors or performance degradation. This approach minimizes risk by exposing only a small percentage of users to the new version initially, then incrementally increasing traffic as confidence grows, which aligns perfectly with the A/B testing requirement.

Exam trap

AWS often tests the distinction between canary and blue/green deployment, where candidates mistakenly choose blue/green because both involve two versions, but blue/green is an instant switch, not a gradual traffic shift.

How to eliminate wrong answers

Option A is wrong because blue/green deployment involves switching all traffic from the old environment (blue) to the new environment (green) in a single cutover, not gradual traffic shifting over 24 hours. Option B is wrong because shadow testing runs the new model version in parallel with the old one but sends traffic only to the old version, comparing outputs offline without affecting live users, so it does not gradually shift traffic. Option C is wrong because direct deployment with immediate full traffic replaces the old version instantly, providing no gradual rollout or A/B testing capability.

18
MCQhard

An ML team uses AWS Step Functions to orchestrate a retraining pipeline triggered by EventBridge when new training data arrives. The pipeline includes a SageMaker training job and a model evaluation. If evaluation fails, the team wants to send an alert. How should they implement this?

A.Use SQS dead-letter queue for failed training jobs
B.Add a Catch rule in the Step Functions state machine to invoke a Lambda alert function
C.Configure SageMaker training job to publish to SNS on failure
D.Use EventBridge to monitor the training job status
AnswerB

A Catch rule on the evaluation state captures the failure and transitions to a Lambda task that sends the alert. This satisfies the stem's requirement to notify the team when evaluation fails, since Step Functions otherwise terminates the execution without invoking downstream alerting.

Why this answer

Step Functions supports error handling via Catch rules; a Catch on the training or evaluation task can transition to a Lambda function that sends an alert.

19
MCQhard

A company needs to deploy a large language model (LLM) on SageMaker with the Triton Inference Server to maximize GPU utilization and reduce latency. They have an NVIDIA A100 GPU. Which SageMaker inference option supports Triton?

A.SageMaker Batch Transform with Triton
B.SageMaker real-time endpoint using a Triton Inference Server container
C.SageMaker Serverless Inference with a custom container
D.SageMaker Neo compiled model on a CPU endpoint
AnswerB

Why this answer

SageMaker real-time endpoints support the Triton Inference Server through a pre-built container that integrates with NVIDIA A100 GPUs, enabling dynamic batching and concurrent model execution to maximize GPU utilization and reduce latency. Triton is designed for high-throughput inference on GPU hardware, making it the correct choice for this scenario.

Exam trap

The trap here is that candidates may confuse SageMaker Batch Transform with real-time endpoints, assuming Triton can be used for batch processing, but Triton is specifically designed for real-time, low-latency inference and is not supported in Batch Transform jobs.

How to eliminate wrong answers

Option A is wrong because SageMaker Batch Transform does not support the Triton Inference Server; it is designed for offline, asynchronous inference on large datasets without real-time GPU optimization features. Option C is wrong because SageMaker Serverless Inference does not support GPU instances or custom containers with Triton; it is limited to CPU-based inference and automatically managed scaling. Option D is wrong because SageMaker Neo compiles models for CPU or edge devices, not for GPU inference with Triton, and using a CPU endpoint would not leverage the A100 GPU's capabilities.

20
MCQmedium

A data science team uses SageMaker Pipelines for automated training. They need to conditionally register a model only if evaluation metrics exceed a threshold. Which pipeline step type should they use after the evaluation step?

A.Condition step
B.Processing step
C.Transform step
D.RegisterModel step
AnswerA

A condition step evaluates a JSON condition against the evaluation step's output and branches execution accordingly, so registration only proceeds when metrics exceed the threshold. This satisfies the requirement for conditional model registration within SageMaker Pipelines, unlike a processing or callback step.

Why this answer

The Condition step evaluates a condition and branches the pipeline; if the condition is met, the pipeline proceeds to register the model.

21
MCQmedium

A company runs a batch inference job on 10 TB of image data stored in S3. Each image needs to be processed by a GPU-accelerated model. The job is not time-sensitive and cost is the primary concern. Which SageMaker option is MOST appropriate?

A.SageMaker Serverless Inference
B.SageMaker Batch Transform with GPU instance and spot instances
C.SageMaker Async Inference with GPU
D.SageMaker real-time endpoint on GPU instances
AnswerB

Batch Transform runs inference over S3 data without a persistent endpoint, and pairing GPU instances with spot capacity cuts cost substantially for a non-time-sensitive job. This satisfies the cost-primary constraint, since managed spot training applies to training jobs, not batch inference.

Why this answer

Batch Transform with GPU spot instances is the most cost-effective choice for a non-time-sensitive, large-scale batch inference job on 10 TB of data. Spot instances offer up to 90% cost savings over on-demand, and Batch Transform natively handles splitting the dataset, distributing work across instances, and writing results to S3 without requiring a persistent endpoint.

Exam trap

The trap here is that candidates confuse 'batch inference' with 'async inference' and choose Option C, not realizing that Async Inference still requires a running endpoint and is designed for near-real-time processing, not cost-optimized offline batch jobs.

How to eliminate wrong answers

Option A is wrong because SageMaker Serverless Inference is designed for intermittent, low-latency workloads with a maximum payload size of 6 MB and a maximum concurrency of 200, making it unsuitable for processing 10 TB of image data. Option C is wrong because SageMaker Async Inference is optimized for near-real-time requests with large payloads (up to 1 GB) and requires a persistent endpoint, incurring higher costs than a batch job that can use spot instances. Option D is wrong because SageMaker real-time endpoints are provisioned 24/7 and designed for low-latency, high-throughput serving, which is wasteful and expensive for a non-time-sensitive batch job that can tolerate startup delays and interruptions.

22
MCQeasy

A company wants to version and track ML models, with an approval workflow for promoting models from staging to production. Which SageMaker feature should they use?

A.SageMaker Model Monitor
B.SageMaker Experiments
C.SageMaker Pipelines
D.SageMaker Model Registry
AnswerD

SageMaker Model Registry stores versioned model groups with metadata and approval status, letting you gate promotion from staging to production through an explicit approval workflow. It directly satisfies the versioning and approval constraint, unlike raw S3 artefacts or plain endpoints, which lack governance states.

Why this answer

SageMaker Model Registry is the correct choice because it provides a centralized repository to catalog, version, and manage ML models, and it supports approval workflows (e.g., PendingApproval, Approved, Rejected) to promote models from staging to production. This directly addresses the requirement for version tracking and an approval gate for model promotion.

Exam trap

The trap here is that candidates confuse SageMaker Pipelines (which orchestrates the workflow) with SageMaker Model Registry (which manages the model versions and approvals), but the question specifically asks for the feature that handles versioning and approval workflow, not the orchestration of the pipeline itself.

How to eliminate wrong answers

Option A is wrong because SageMaker Model Monitor is designed for detecting data and model quality drift in production, not for versioning or approval workflows. Option B is wrong because SageMaker Experiments is used for tracking and comparing training runs (e.g., hyperparameters, metrics), not for managing model versions or approval states. Option C is wrong because SageMaker Pipelines orchestrates end-to-end ML workflows (e.g., data processing, training, deployment) but does not natively provide a model version registry or approval workflow; it can integrate with Model Registry for that purpose.

23
MCQmedium

A company is deploying a large NLP model on SageMaker for real-time inference. They want to reduce inference latency and cost by optimizing the model for the target hardware. The model is trained in PyTorch. Which SageMaker feature should they use to compile the model for best performance on the chosen instance?

A.SageMaker Neo
B.AWS Step Functions
C.Amazon Elastic Inference
D.SageMaker Triton Inference Server
AnswerA

SageMaker Neo compiles PyTorch models into optimised executables tuned to the target instance's specific processor architecture, cutting inference latency and cost. It satisfies the stem's requirement to compile the trained model for best performance on the chosen hardware, unlike generic deployment or autoscaling features that leave the model graph unoptimised.

Why this answer

SageMaker Neo is the correct choice because it is specifically designed to compile trained models (including PyTorch models) into an optimized binary for a target hardware instance, reducing inference latency and improving throughput. Neo applies hardware-specific optimizations such as operator fusion, memory layout tuning, and quantization, which directly address the need for best performance on the chosen SageMaker instance.

Exam trap

The trap here is that candidates confuse model compilation (Neo) with inference serving (Triton) or hardware acceleration (Elastic Inference), leading them to pick a service that addresses a different part of the inference pipeline.

How to eliminate wrong answers

Option B is wrong because AWS Step Functions is a serverless workflow orchestration service, not a model compilation tool; it cannot optimize model performance for hardware. Option C is wrong because Amazon Elastic Inference attaches a separate accelerator to an instance for cost-effective inference, but it does not compile or optimize the model itself; it only provides additional compute resources. Option D is wrong because SageMaker Triton Inference Server is a high-performance inference server that supports multiple frameworks and model formats, but it does not compile the model for the target hardware; it serves models as-is, relying on the underlying framework's runtime.

24
MCQmedium

A data scientist wants to train a model on SageMaker using a custom PyTorch script, then register the best model in the SageMaker Model Registry. The training job is part of a SageMaker Pipeline. Which pipeline step should be used to register the model?

A.RegisterModelStep
B.CreateModelStep
C.TrainingStep
D.TransformStep
AnswerA

RegisterModelStep is the dedicated SageMaker Pipelines step that packages the trained model artefact and its metadata into a model package group in the Model Registry, chaining directly from the training step within the pipeline definition.

Why this answer

The `RegisterModelStep` is specifically designed to create a model resource and register it in the SageMaker Model Registry as part of a pipeline. It takes the training output (e.g., model artifacts from a `TrainingStep`) and packages it with the specified inference image and metadata, then creates a model package group version. This is the correct step for registering a model after training, as it directly integrates with the Model Registry for versioning and approval workflows.

Exam trap

The trap here is that candidates confuse `CreateModelStep` (which creates a deployable model resource) with `RegisterModelStep` (which creates a model package version in the registry), assuming both serve the same purpose of model registration.

How to eliminate wrong answers

Option B is wrong because `CreateModelStep` only creates a SageMaker model resource (for deployment or batch inference) but does not register it in the Model Registry; it lacks the versioning and metadata capabilities needed for registry management. Option C is wrong because `TrainingStep` is used to run a training job and produce model artifacts, but it has no built-in functionality to register the model into the Model Registry; registration requires a separate step. Option D is wrong because `TransformStep` is used for batch inference (transform jobs) on existing models, not for registering models into the registry.

25
MCQmedium

A company is using SageMaker Pipelines to orchestrate their ML workflow. They have a Condition step that checks if a model's accuracy exceeds 0.9. If true, they want to register the model in the model registry; otherwise, they want to run a retraining step. Which step type should they use for the decision?

A.Condition step
B.Transform step
C.Processing step
D.Tuning step
AnswerA

A Condition step evaluates a boolean expression against the evaluation step's accuracy metric and branches accordingly: the true branch registers the model, the false branch triggers retraining. It is the only step type providing conditional branching logic in SageMaker Pipelines.

Why this answer

A Condition step in SageMaker Pipelines evaluates a boolean expression against a property (such as a model accuracy metric from a preceding Evaluation step) and branches the pipeline into a true or false path. This is exactly the construct needed to register the model when accuracy > 0.9 and otherwise trigger retraining.

Exam trap

The trap is assuming any step that 'makes a decision' must be a Processing step because it runs code — candidates forget that SageMaker Pipelines has a dedicated Condition step type specifically for branching logic.

How to eliminate wrong answers

Option B is wrong because a Transform step runs batch transform jobs for inference against a dataset, not conditional branching. Option C is wrong because a Processing step runs a containerized data-processing or evaluation job; it produces outputs but does not branch the DAG. Option D is wrong because a Tuning step runs a hyperparameter tuning job that launches multiple training jobs — it has no conditional logic capability.

26
MCQmedium

A company wants to deploy 50 small models (each ~100 MB) for real-time inference. They need to minimize hosting costs while maintaining low latency. Which SageMaker hosting option is most cost-effective?

A.SageMaker Serverless Inference
B.SageMaker Asynchronous Inference
C.SageMaker Multi-Model Endpoint (MME)
D.SageMaker real-time endpoint with one instance per model
AnswerC

Multi-model endpoint loads models on demand into shared memory and caches them, so 50 models share one instance rather than each needing its own endpoint. This satisfies the cost-minimisation constraint while preserving real-time, low-latency inference.

Why this answer

SageMaker Multi-Model Endpoint (MME) hosts many models on a single endpoint and dynamically loads them into memory/disk on invocation, sharing the underlying instance. For 50 small models (~100 MB each), this dramatically reduces hosting cost versus one endpoint per model while still providing real-time, low-latency inference. Serverless Inference is pay-per-invoke but has cold starts and memory limits that can hurt latency for frequent calls.

Exam trap

MLA-C01 often tests the misconception that Serverless Inference is always cheapest, ignoring cold starts and the fact that MME shares one instance across many models for steady real-time traffic.

How to eliminate wrong answers

Option A is wrong because Serverless Inference incurs cold-start latency and is billed per invocation with memory/time caps, which is not ideal for consistently low-latency, high-volume real-time inference across 50 models. Option B is wrong because Asynchronous Inference is designed for large payloads and long processing times with queued requests, not low-latency real-time responses. Option D is wrong because one instance per model for 50 models multiplies hosting costs by 50, directly violating the cost-minimization requirement.

27
Multi-Selecthard

A team is building a SageMaker Pipeline that trains a model and then registers it in the SageMaker Model Registry. They want the pipeline to automatically deploy the model to a real-time endpoint only after a human approves the model package. Which two actions should the team take to implement this approval-gated deployment? (Choose two.)

Select 2 answers
A.Set the model package status to Approved directly in the RegisterModel step configuration so the endpoint deploys without human intervention.
B.Add a RegisterModel step to the pipeline that creates a model package with a PendingManualApproval status.
C.Use a SageMaker Clarify processing step to approve the model package when bias metrics fall within thresholds.
D.Add a ConditionStep to the pipeline that checks the model package status immediately after RegisterModel and branches to a deployment step.
E.Configure an Amazon EventBridge rule that matches the model package state change to Approved and invokes a target that starts the deployment.
AnswersB, E

RegisterModel creates a model package group entry and sets the package status to PendingManualApproval when manual approval is configured. This produces the artifact that a reviewer evaluates and approves, which is the trigger point the deployment automation watches for.

Why this answer

Manual approval gating relies on the model package lifecycle: RegisterModel creates a package in PendingManualApproval, and a reviewer later approves it. Because the pipeline execution does not pause for review, an EventBridge rule watching for the Approved state change is what triggers the downstream deployment automation, keeping humans in the loop.

Exam trap

The trap here is assuming a pipeline step can wait for human approval, when pipelines are finite executions and approval events must be handled by an external event-driven mechanism.

28
MCQeasy

A data science team has trained a PyTorch model for real-time inference and needs to deploy it on AWS with GPU acceleration while minimizing cold-start latency. Which SageMaker inference option should they choose?

A.Serverless inference
B.Batch transform
C.Asynchronous inference endpoint
D.Real-time endpoint with ml.g4dn instance
AnswerD

A real-time endpoint keeps the model loaded on a persistent ml.g4dn GPU instance, so inference requests avoid the container initialisation delay that serverless inference incurs. This satisfies the GPU acceleration and minimal cold-start latency constraints simultaneously.

Why this answer

Real-time endpoints with GPU instances (e.g., ml.g4dn) provide low latency and support GPU acceleration, suitable for interactive inference. Serverless inference does not support GPU instances, asynchronous inference is for non-real-time, and batch transform is for offline predictions.

29
MCQhard

A data science team wants to host 50 different models for a recommendation engine. Each model is small (under 100 MB) and traffic patterns are unpredictable. They need to minimize cost and operational overhead. Which approach should they take?

A.Deploy each model to its own real-time endpoint
B.Use SageMaker serverless inference for each model
C.Use a single multi-model endpoint (MME)
D.Use a single multi-container endpoint
AnswerC

A multi-model endpoint hosts many models behind one container, so 50 small models share a single endpoint rather than 50 separate ones. This directly minimises cost and operational overhead while handling unpredictable traffic through shared autoscaling.

Why this answer

A single multi-model endpoint (MME) allows hosting multiple models (up to thousands) on the same endpoint, sharing the underlying compute instance. This minimizes cost and operational overhead for small models (under 100 MB) with unpredictable traffic, as the endpoint dynamically loads and unloads models from Amazon S3 into memory based on incoming requests, eliminating the need for separate endpoints or idle compute.

Exam trap

AWS often tests the distinction between multi-model endpoints (for multiple independent models) and multi-container endpoints (for a single model with multiple containers), leading candidates to confuse the two and incorrectly choose option D.

How to eliminate wrong answers

Option A is wrong because deploying each model to its own real-time endpoint would require 50 separate endpoints, each with its own compute instance, leading to high cost and operational overhead due to idle resources during unpredictable traffic patterns. Option B is wrong because SageMaker serverless inference is designed for infrequent or sporadic traffic, but it still requires a separate serverless endpoint per model, incurring per-request costs and cold-start latency for each model, which does not minimize cost or overhead for 50 models. Option D is wrong because a single multi-container endpoint is intended for hosting multiple containers that serve a single model (e.g., pre-processing and inference), not for hosting multiple independent models; it does not support dynamic model loading and would require separate containers for each model, defeating the purpose of consolidation.

30
MCQmedium

A media company runs a real-time recommendation model on a SageMaker endpoint. Traffic triples every evening between 18:00 and 22:00 and drops to near zero overnight, and the team wants to cut costs without underserving evening users. They want the endpoint to scale out automatically as invocations rise and scale back in when traffic falls. Which solution should they implement?

A.Create a second endpoint and use an Application Load Balancer to route evening traffic between the two endpoints.
B.Deploy the model to a SageMaker Serverless Inference endpoint and set the maximum concurrency to a fixed value equal to the evening peak.
C.Enable automatic scaling on the endpoint with a target-tracking scaling policy based on the SageMakerVariantInvocationsPerInstance metric.
D.Increase the instance count of the endpoint to the evening peak permanently and rely on Savings Plans to offset the extra cost.
AnswerC

Target tracking on SageMakerVariantInvocationsPerInstance lets Application Auto Scaling add instances when per-instance invocation load exceeds the target and remove them when it drops, matching the nightly peak-and-trough pattern without manual intervention. This is the documented approach for provisioning real-time endpoint capacity dynamically, and it preserves latency by scaling out before instances saturate.

Why this answer

Target-tracking automatic scaling based on SageMakerVariantInvocationsPerInstance is the native mechanism for elastic real-time endpoints. It registers the variant as a scalable target with Application Auto Scaling, then adds or removes instances to hold the metric near the target. This directly matches a recurring evening peak followed by an overnight trough while keeping latency low, unlike fixed serverless concurrency or static over-provisioning.

Exam trap

The trap here is assuming Serverless Inference is always the cheapest elastic option, when a fixed maximum concurrency reintroduces the same idle cost as a permanently over-provisioned endpoint.

31
MCQhard

A company wants to serve a large ensemble of models using NVIDIA Triton Inference Server on SageMaker for high throughput GPU inference. Which SageMaker inference option supports this?

A.Asynchronous Inference
B.Multi-model endpoint
C.Serverless Inference
D.Real-time endpoint with a custom container running Triton
AnswerD

A real-time endpoint with a custom container lets you run NVIDIA Triton Inference Server directly, enabling ensemble execution, dynamic batching and concurrent model execution on GPU instances. This satisfies the high-throughput GPU inference requirement, which single-model SageMaker containers cannot provide.

Why this answer

NVIDIA Triton Inference Server is a custom inference server that supports multiple frameworks, model ensembles, and dynamic batching. To use it on SageMaker, you deploy a real-time endpoint with a custom container that runs Triton, which supports GPU inference and high throughput. This is the only option that explicitly supports Triton.

Exam trap

MLA-C01 often tests the assumption that any SageMaker hosting option can run Triton, when only a real-time endpoint with a custom container supports arbitrary inference servers like Triton.

How to eliminate wrong answers

Option A is wrong because Asynchronous Inference is a request-queueing mode for large payloads/long processing, not a mechanism for running Triton. Option B is wrong because Multi-Model Endpoint is a SageMaker hosting feature for many models on one endpoint, but it does not natively run Triton; Triton requires a custom container. Option C is wrong because Serverless Inference does not support custom containers with GPU and has cold-start/memory limits incompatible with high-throughput Triton ensembles.

32
MCQmedium

A retailer runs a nightly batch scoring job that processes 40 GB of transaction data and writes predictions to S3. Occasionally a single partition is corrupt, causing the entire job to fail after several hours. The team wants the job to skip the corrupt partition, log which partition failed, and still complete processing of the remaining data with minimal changes to their existing SageMaker Processing job. Which change should they make?

A.Wrap the per-partition processing logic in try/except inside the Processing container script, record the failing partition key to CloudWatch Logs, and continue to the next partition.
B.Configure the Processing job with a retry policy and a maximum retry count of three so transient failures are retried automatically.
C.Convert the Processing job to a SageMaker Training job with checkpointing enabled so it can resume after the corrupt partition.
D.Increase the Processing job's instance count and volume size so the corrupt partition is retried on a different instance.
AnswerA

Handling exceptions per partition inside the container lets the job skip a corrupt input, emit the partition key to CloudWatch Logs, and proceed with the remaining data. It requires only a code change to the existing Processing job, preserving the current orchestration and S3 output path. This directly satisfies fault isolation, failure logging, and job completion with minimal disruption.

Why this answer

Adding per-partition exception handling inside the Processing container isolates corruption to the affected partition, logs its key for follow-up, and allows the job to finish the rest of the data. This is a localized code change that preserves the existing SageMaker Processing job structure and S3 outputs. It meets fault isolation, failure visibility, and completion requirements with the least disruption.

Exam trap

The trap here is reaching for infrastructure-level retries or resizing, when a deterministic data corruption must be handled in application code to isolate the bad partition.

33
MCQhard

A team uses SageMaker Pipelines to train and register a model. They want to conditionally run a hyperparameter tuning step only if the data quality check passes. Which pipeline step type should they use to branch the execution?

A.TuningStep
B.TrainingStep
C.ConditionStep
D.TransformStep
AnswerC

ConditionStep evaluates a Boolean condition against a pipeline property, such as the data quality check's output, and branches execution accordingly. It gates the tuning step so it runs only when the check passes, satisfying the conditional-execution requirement without altering the underlying training logic.

Why this answer

The ConditionStep allows comparing values and branching to different steps. If data quality passes, the tuning step runs; otherwise, the pipeline stops or runs an alternative step. Other steps do not provide conditional branching.

34
Multi-Selecthard

A machine learning team is deploying a model to a SageMaker endpoint and needs to implement A/B testing between two model versions. They want to split traffic 80/20 and monitor performance metrics for each variant. Which two actions should they take? (Choose two.)

Select 2 answers
A.Configure an Application Load Balancer to distribute traffic between two separate endpoints.
B.Use SageMaker Model Monitor to automatically compare the accuracy of the two variants.
C.Create a SageMaker endpoint configuration with two production variants, each specifying a different model and initial weight.
D.Create two separate SageMaker endpoints and use AWS Lambda to route requests based on a random number.
E.Enable data capture on the endpoint to log request and response data for analysis.
AnswersC, E

An endpoint configuration with multiple production variants allows you to deploy multiple models to a single endpoint. Each variant can have a different model and an initial weight that determines the traffic split. This is the foundational step for A/B testing, as it enables simultaneous serving of both model versions.

Why this answer

To perform A/B testing on SageMaker, you create an endpoint configuration with multiple production variants, each with a different model and weight to split traffic. Enabling data capture logs the requests and responses for each variant, allowing you to analyze performance metrics. Other options either do not provide native A/B testing or add unnecessary complexity.

Exam trap

The trap here is assuming that Model Monitor automatically compares variant performance, but it only monitors individual models for drift and quality.

35
MCQmedium

A company needs to deploy a new model version to a SageMaker real-time endpoint. They want to route 5% of traffic to the new version initially to monitor for errors before full rollout. Which deployment strategy should they use?

A.Blue/green deployment
B.Shadow testing
C.Canary deployment with production variants
D.Multi-model endpoint
AnswerC

Canary deployment with production variants lets SageMaker split endpoint traffic, sending 5% to the new model variant while the old version serves the rest. This directly satisfies the requirement to monitor errors before full rollout.

Why this answer

A canary deployment with production variants allows you to route a specific percentage of traffic (e.g., 5%) to the new model version by adjusting the `InitialVariantWeight` parameter in the production variant configuration. This enables gradual traffic shifting while monitoring errors, and you can later increase the weight to 100% for full rollout. SageMaker real-time endpoints support this natively by hosting multiple model variants behind the same endpoint.

Exam trap

The trap here is that candidates confuse canary deployment with shadow testing, mistakenly thinking shadow testing also routes live user traffic, when in fact shadow testing only duplicates traffic for validation without affecting the user experience.

How to eliminate wrong answers

Option A is wrong because blue/green deployment switches all traffic from the old version to the new version at once, not a gradual 5% routing, which defeats the purpose of initial error monitoring. Option B is wrong because shadow testing sends a copy of live traffic to the new version but does not serve responses to users; it is used for validation without impacting production traffic, not for routing a percentage of user-facing traffic. Option D is wrong because a multi-model endpoint hosts multiple models on the same endpoint but does not provide traffic splitting or weighted routing between model versions; it is designed for cost efficiency with many models, not gradual rollout.

36
MCQeasy

A machine learning engineer needs to run a one-time scoring job over 500 GB of data stored in Amazon S3 using a trained model, and the results must be written back to S3. There is no requirement for a persistent HTTPS endpoint. Which SageMaker feature should the engineer use?

A.Run a SageMaker Batch Transform job that reads input objects from S3 and writes inference output back to S3.
B.Create a real-time endpoint and send each S3 object as an HTTPS request through the InvokeEndpoint API.
C.Use a SageMaker Processing job with a custom script that loads the model and writes predictions to S3.
D.Deploy a serverless inference endpoint and invoke it repeatedly until all 500 GB has been processed.
AnswerA

Batch Transform is purpose-built for offline scoring of large datasets in S3. It provisions the needed compute, distributes the work across the input objects, and writes output back to a specified S3 location, then tears the resources down, matching the one-time job with no persistent endpoint.

Why this answer

Batch Transform exists precisely for offline, high-volume inference where inputs and outputs live in Amazon S3 and no persistent endpoint is required. It manages the compute fleet for the duration of the job, parallelizes across input data, writes results to S3, and then releases resources, which fits a one-time 500 GB scoring task.

Exam trap

The trap here is reaching for a real-time or serverless endpoint for bulk scoring, when those are optimized for interactive request/response rather than large offline datasets.

37
MCQhard

An ML team uses SageMaker Pipelines to automate model retraining. They want to skip redundant training steps when input data has not changed. Which feature should they enable?

A.Pipeline caching
B.Pipeline variable expressions
C.Model registry approval
D.Step parallelism
AnswerA

Pipeline caching reuses a step's outputs when its inputs, code and parameters are unchanged, so the training step is skipped entirely rather than re-executed. This directly satisfies the stem's requirement to avoid redundant training when input data has not changed, saving compute time and cost.

Why this answer

SageMaker Pipelines caching stores the output of a step keyed by the step's inputs (code, data, hyperparameters, etc.). When a pipeline run executes and the inputs are unchanged, the cached output is reused and the step is skipped, avoiding redundant training. This directly addresses the requirement to skip training when input data hasn't changed.

Exam trap

MLA-C01 often tests the confusion between caching (skip unchanged steps) and parallelism (run steps concurrently) — candidates must match the optimization goal to the correct feature.

How to eliminate wrong answers

Option B is wrong because pipeline variable expressions are for parameterizing pipeline definitions (e.g., passing values between steps or at runtime), not for detecting unchanged inputs and skipping execution. Option C is wrong because model registry approval is a governance workflow for promoting models to production, unrelated to step execution optimization. Option D is wrong because step parallelism runs independent steps concurrently to reduce wall-clock time, but it does not skip steps whose inputs are unchanged.

38
MCQmedium

A machine learning engineer is deploying a real-time inference endpoint on Amazon SageMaker AI for a fraud detection model. The model must serve predictions with consistent latency under 50 ms and the team expects traffic to fluctuate unpredictably, with occasional bursts. The engineer wants to automatically adjust the number of instances based on actual workload while minimizing cost during idle periods. Which SageMaker AI feature should the engineer configure?

A.SageMaker AI multi-model endpoints with a shared inference container
B.SageMaker AI automatic scaling with a target-tracking policy based on the SageMakerVariantInvocationsPerInstance metric
C.SageMaker AI Inference Recommender to select the optimal instance type and count
D.SageMaker AI Asynchronous Inference with an auto-scaling policy on queue depth
AnswerB

Target-tracking automatic scaling uses CloudWatch metrics like SageMakerVariantInvocationsPerInstance to adjust instance count dynamically, matching capacity to demand. This directly addresses unpredictable traffic and minimizes cost during low usage while maintaining latency. It is the native SageMaker AI scaling mechanism for real-time endpoints.

Why this answer

Automatic scaling with target-tracking on the invocations-per-instance metric is the correct approach because it continuously adjusts the number of instances to match actual traffic, ensuring latency targets are met while scaling down during idle periods to save cost. The other options either provide static recommendations or are designed for different inference patterns.

Exam trap

The trap here is confusing a one-time right-sizing recommendation from Inference Recommender with runtime automatic scaling that responds to live traffic.

39
MCQmedium

A company wants to deploy a machine learning model using infrastructure as code to ensure reproducibility. They need to define the SageMaker Studio domain, user profiles, and the endpoint configuration. Which tool should they use?

A.AWS CloudFormation or AWS CDK
B.SageMaker Pipelines
C.AWS Step Functions
D.SageMaker Studio
AnswerA

AWS CloudFormation and AWS CDK both declare SageMaker resources — Studio domains, user profiles, endpoint configurations — as versioned templates, so the same stack redeploys identically across environments. This satisfies the reproducibility constraint, unlike console-based provisioning or imperative scripts, which drift and cannot be diffed or rolled back.

Why this answer

AWS CloudFormation and AWS CDK are infrastructure-as-code (IaC) tools that allow you to define, provision, and manage AWS resources declaratively. For this use case, they can model the entire SageMaker Studio domain, user profiles, and endpoint configuration in templates or code, ensuring reproducibility and version control. This aligns directly with the requirement to deploy ML infrastructure as code.

Exam trap

The trap here is that candidates confuse SageMaker Pipelines (a CI/CD service for ML steps) with infrastructure-as-code tools, forgetting that Pipelines does not manage underlying infrastructure resources like Studio domains or endpoint configurations.

How to eliminate wrong answers

Option B (SageMaker Pipelines) is wrong because it is a purpose-built CI/CD service for ML workflows (training, tuning, batch transforms), not for defining and provisioning infrastructure resources like Studio domains or endpoints. Option C (AWS Step Functions) is wrong because it is a serverless workflow orchestration service for coordinating distributed applications and microservices, not for defining infrastructure resources declaratively. Option D (SageMaker Studio) is wrong because it is the web-based IDE for ML development, not a tool for defining or deploying infrastructure as code.

40
MCQeasy

Which SageMaker feature compiles a trained model into an optimized binary for a specific hardware target (e.g., Intel, ARM, NVIDIA, or edge devices) to improve inference performance?

A.SageMaker Model Monitor
B.SageMaker Neo
C.Amazon Elastic Inference
D.SageMaker Clarify
AnswerB

SageMaker Neo compiles trained models into optimised executables for specific hardware targets, including Intel, ARM, NVIDIA and edge devices. This directly satisfies the stem's requirement for a hardware-specific binary that improves inference performance, unlike training or hosting features that leave the model format unchanged.

Why this answer

SageMaker Neo is a model compilation service that optimizes models for specific hardware targets. Amazon Elastic Inference attaches GPU acceleration to endpoints, but does not compile models. Model Monitor monitors quality.

SageMaker Clarify explains predictions.

41
MCQmedium

A company needs to serve real-time predictions from a large ensemble of three deep learning models, each requiring different inference environments (PyTorch, TensorFlow, MXNet). Which SageMaker endpoint type supports running multiple inference containers together?

A.Multi-model endpoint
B.Real-time endpoint with a single container
C.Multi-container endpoint
D.Asynchronous endpoint
AnswerC

Multi-container endpoints run up to fifteen containers together on one instance, letting each model use its own PyTorch, TensorFlow or MXNet environment. This satisfies the stem's need for multiple inference containers co-hosted, which single-model and serverless options cannot provide.

Why this answer

Amazon SageMaker multi-container endpoints allow you to run multiple inference containers (e.g., PyTorch, TensorFlow, MXNet) within a single endpoint, each handling different models or inference environments. This is achieved by deploying multiple containers behind a single endpoint with a serial or direct invocation pattern, enabling real-time predictions from the ensemble without managing separate endpoints.

Exam trap

The trap here is that candidates often confuse 'multi-model endpoint' (multiple models in one container) with 'multi-container endpoint' (multiple containers with different environments), leading them to incorrectly select Option A.

How to eliminate wrong answers

Option A is wrong because a multi-model endpoint hosts multiple models within a single container, not multiple containers with different inference environments; it uses a shared serving container and loads models dynamically from Amazon S3. Option B is wrong because a real-time endpoint with a single container can only run one inference environment, making it impossible to serve the three different deep learning frameworks required by the ensemble. Option D is wrong because an asynchronous endpoint is designed for large payloads and long processing times, not for real-time predictions, and it still uses a single container per endpoint.

42
MCQmedium

A fraud-detection team trains a model in SageMaker and wants to shift 10 percent of live prediction traffic to a newly retrained model to compare accuracy before a full cutover. The endpoint already serves the current model on one production variant. They need the endpoint to route a controlled fraction of requests to the new model without changing the client application. What should they do?

A.Update the existing endpoint to add a second production variant for the retrained model and set its initial variant weight to 10 while the original variant keeps 90.
B.Enable an inference pipeline on the endpoint and place the retrained model as the second container in the pipeline.
C.Use a SageMaker batch transform job against the live traffic stream to score 10 percent of requests with the retrained model.
D.Create a new endpoint for the retrained model and use Route 53 weighted routing to send 10 percent of DNS queries to the new endpoint.
AnswerA

SageMaker production variants let one endpoint host multiple models with assigned weights, and InvokeEndpoint distributes traffic according to those weights. Setting the new variant to 10 and the existing one to 90 shifts a controlled fraction of requests without any client change, since the client still calls the same endpoint name. This is the built-in mechanism for canary-style traffic shifting.

Why this answer

Production variants with weights are the native SageMaker mechanism for sending a percentage of live traffic to a new model on the same endpoint. Adding a variant for the retrained model and assigning it a weight of 10 while the incumbent holds 90 yields a canary deployment that clients see as a single endpoint. Weights can be adjusted as confidence grows and the old variant removed at full cutover.

Exam trap

The trap here is reaching for DNS-level or batch mechanisms for traffic splitting, when SageMaker production variant weights operate per invocation on a single endpoint.

43
MCQmedium

A company wants to serve 200 different PyTorch models. Each model is small (under 1 GB) and only a fraction are used at any time. To minimize cost and management overhead, which SageMaker inference option should be used?

A.Use batch transform for all models
B.Create a separate real-time endpoint for each model
C.Use a multi-container endpoint
D.Use a multi-model endpoint
AnswerD

Multi-model endpoints load models on demand from S3 into shared container memory and unload idle ones, so hosting 200 small PyTorch models costs far less than 200 endpoints. This satisfies the minimise-cost and management-overhead constraint given only a fraction are used concurrently.

Why this answer

A multi-model endpoint (MME) is the correct choice because it allows you to host multiple small PyTorch models (under 1 GB each) on a single endpoint, sharing the underlying compute instance. This minimizes cost by only paying for the active instances, and reduces management overhead since you don't need to create or manage separate endpoints for each model. SageMaker dynamically loads and unloads models from the container's memory based on invocation patterns, which is ideal for a scenario where only a fraction of the 200 models are used at any time.

Exam trap

In AWS SageMaker, the common trap is confusing multi-container endpoints (used for multi-step inference pipelines with different containers) with multi-model endpoints (used for hosting multiple independent models). Candidates often select multi-container endpoints thinking they can serve multiple models, but each container still hosts one model; multi-model endpoints are designed to host many models on a single instance, dynamically loading them as needed.

How to eliminate wrong answers

Option A is wrong because batch transform is designed for offline, asynchronous inference on a complete dataset, not for serving real-time predictions on-demand, and it would incur cost for processing all models even when not needed. Option B is wrong because creating a separate real-time endpoint for each of the 200 models would be prohibitively expensive and introduce significant management overhead, as each endpoint requires its own compute instance and incurs hourly charges regardless of usage. Option C is wrong because a multi-container endpoint is used to run multiple containers (e.g., for pre-processing, inference, post-processing) within a single endpoint, not to host multiple independent models; it does not support dynamic loading/unloading of different model artifacts.

44
MCQeasy

A company needs to deploy a model that processes large payloads (up to 1 GB) asynchronously. The results should be written to S3, and the team needs SNS notifications upon completion. Which SageMaker inference option is MOST suitable?

A.Asynchronous Inference
B.Batch Transform
C.Real-time endpoint
D.Serverless Inference
AnswerA

Asynchronous Inference queues requests and supports payloads up to 1 GB, writing results to Amazon S3 and publishing completion notifications via Amazon SNS. This directly satisfies the stem's large-payload, asynchronous and notification constraints, unlike real-time endpoints, which cap payloads far smaller.

Why this answer

SageMaker Asynchronous Inference is purpose-built for large payloads (up to 1 GB) and long processing times (up to 15 minutes), queuing requests and writing results to S3. It natively supports SNS notifications on completion or failure, matching the requirement exactly. This makes it the correct fit for asynchronous, large-payload processing with S3 output and SNS alerts.

Exam trap

MLA-C01 often tests the payload/timeout limits of each inference type — candidates who don't memorize the 1 GB/15 min (async), 6 MB/60 s (real-time), and 4 MB/60 s (serverless) limits pick the wrong option.

How to eliminate wrong answers

Option B is wrong because Batch Transform is for offline batch scoring of an entire dataset, not for individual asynchronous requests with per-request SNS notifications. Option C is wrong because Real-time endpoints have a 6 MB payload limit and a 60-second timeout, so they cannot handle 1 GB payloads. Option D is wrong because Serverless Inference has a 4 MB payload limit and a 60-second timeout, and does not natively write results to S3 or send SNS notifications.

45
MCQhard

A company deploys a large NLP model on a SageMaker real-time endpoint using an ml.p3.2xlarge instance. To reduce inference cost without sacrificing throughput, they want to compile the model for their target hardware. Which service should they use?

A.SageMaker Neo
B.Triton Inference Server on SageMaker
C.Amazon Elastic Inference
D.SageMaker Inference Recommender
AnswerA

SageMaker Neo compiles the model for the specific target instance's instruction set, producing an optimised artefact that runs faster on ml.p3.2xlarge. This satisfies the reduce-cost-without-losing-throughput constraint by improving hardware utilisation rather than shrinking the instance.

Why this answer

SageMaker Neo is the correct service because it compiles trained machine learning models into an optimized binary for a specific target hardware (e.g., ml.p3.2xlarge with NVIDIA GPUs). This reduces inference latency and cost by applying hardware-specific optimizations such as kernel fusion and memory layout tuning, while preserving the original model's throughput. The compilation process uses Apache TVM under the hood to generate efficient code for the target instance type.

Exam trap

AWS often tests the distinction between model compilation (Neo) and runtime serving optimizations (Triton) or hardware acceleration (Elastic Inference), leading candidates to confuse a compile-time optimization service with a runtime serving framework or a hardware add-on.

How to eliminate wrong answers

Option B is wrong because Triton Inference Server on SageMaker is a model serving framework that supports multiple backends (e.g., TensorRT, ONNX Runtime) and dynamic batching, but it does not perform ahead-of-time model compilation for a specific hardware target; it optimizes runtime execution, not compile-time optimization. Option C is wrong because Amazon Elastic Inference attaches a separate accelerator to a CPU instance for cost savings, but it is not a model compilation service; it is a hardware attachment that does not compile the model for the target instance. Option D is wrong because SageMaker Inference Recommender is a tool for benchmarking and recommending instance types and endpoint configurations based on load tests, not for compiling or optimizing the model binary for a specific hardware target.

46
Multi-Selectmedium

A company wants to deploy a new model using a canary deployment strategy on SageMaker. Which two actions should they take? (Select TWO.)

Select 2 answers
A.Register both models in the Model Registry with 'Approved' status
B.Use SageMaker Model Monitor to compare model performance
C.Create a new endpoint with two production variants
D.Enable data capture on the endpoint
E.Set the initial traffic weights for the variants (e.g., 95% and 5%)
AnswersC, E

A canary deployment on SageMaker requires a single endpoint hosting two production variants, with traffic initially weighted heavily toward the existing model. Creating that endpoint satisfies the stem's requirement, letting a small percentage of inference requests shift to the new model.

Why this answer

Option C is correct because a SageMaker canary deployment is implemented by creating an endpoint configuration with two production variants (the existing/stable model and the new model) and deploying them behind a single endpoint, which is the foundational mechanism for splitting traffic between versions. Option E is correct because the canary pattern requires assigning initial traffic weights to those variants—typically a small percentage (e.g., 5%) to the new model and the remainder (e.g., 95%) to the stable model—so the new model receives limited live traffic before being promoted. Option A is not required for the deployment itself; Model Registry 'Approved' status is a governance/approval step, not a technical prerequisite for creating canary variants.

Option B is not required because Model Monitor is used for detecting data/quality drift and bias, not for comparing model performance during a canary rollout. Option D is not required because data capture is an optional monitoring feature that logs request/response data, not a necessary action to perform a canary deployment.

Exam trap

MLA-C01 often tests whether candidates confuse the mechanics of canary deployment (two variants + traffic weights) with supporting features like Model Monitor or Model Registry — the trap is selecting monitoring/governance options as if they were required to create the canary.

47
MCQmedium

A team has a SageMaker Pipeline that trains a model and registers it in the Model Registry. They want to automate the deployment of the approved model to a staging environment. Which event-driven approach should they use?

A.Use an SQS queue to store approval messages and have a cron job process them
B.Set up a CloudWatch alarm on the Model Registry's ApprovalStatus metric
C.Use Amazon EventBridge to listen for Model Registry approval events and trigger an AWS Lambda function that deploys the model
D.Configure an AWS Step Functions state machine to poll the Model Registry every minute
AnswerC

Amazon EventBridge subscribes to Model Registry approval state-change events, and a rule invokes a Lambda function that performs the deployment. This event-driven pattern reacts automatically to approval, satisfying the requirement to deploy the approved model to staging without polling.

Why this answer

Amazon EventBridge natively integrates with SageMaker Model Registry and emits events such as 'SageMaker Model Package State Change' when a model package's approval status changes to Approved. An EventBridge rule can match that event pattern and invoke a Lambda function (or Step Functions, CodePipeline, etc.) to deploy the model to staging. This is the canonical event-driven pattern for automating post-approval deployment in SageMaker MLOps workflows.

Exam trap

MLA-C01 often tests whether candidates confuse CloudWatch metrics/alarms with EventBridge events — Model Registry approval is an event, not a metric, so any answer referencing a CloudWatch alarm on ApprovalStatus is a distractor.

How to eliminate wrong answers

Option A is wrong because SQS is a pull-based queue with no native awareness of Model Registry state changes, and a cron job introduces polling latency and unnecessary infrastructure. Option B is wrong because CloudWatch does not expose an 'ApprovalStatus' metric for the Model Registry — approval is an event, not a metric, so an alarm cannot be created on it. Option D is wrong because polling every minute is inefficient, adds cost and latency, and Step Functions is not the recommended trigger mechanism when EventBridge delivers the event directly.

48
MCQmedium

A startup wants to deploy a containerized ML application that includes both a model inference server and a preprocessing component in the same endpoint. Which SageMaker endpoint type supports running multiple containers?

A.Asynchronous Inference
B.Multi-container endpoint
C.Multi-model endpoint
D.Real-time endpoint
AnswerB

Multi-container endpoints run up to fifteen containers behind one endpoint, letting the inference server and preprocessing component be packaged and invoked together. This satisfies the stem's requirement to host both components within a single endpoint.

Why this answer

SageMaker multi-container endpoints allow you to run up to 15 containers on a single endpoint, with containers invoked in a defined sequence (inference pipeline) or directly. This is the correct choice when a preprocessing component and a model inference server must coexist in one endpoint, since the preprocessing container can transform the request before passing it to the inference container. This pattern is ideal for encapsulating feature engineering with the model.

Exam trap

MLA-C01 often tests the confusion between multi-container endpoints (multiple containers, one endpoint) and multi-model endpoints (multiple models, one container) — candidates who conflate the two pick the wrong option.

How to eliminate wrong answers

Option A is wrong because Asynchronous Inference is a deployment mode for long-running or large-payload requests with queuing, not a mechanism for running multiple containers. Option C is wrong because Multi-model endpoints host many models behind a single container using a shared serving stack — they do not run multiple distinct containers. Option D is wrong because Real-time endpoint is a hosting mode (synchronous, low-latency) that can host a single container or a multi-container pipeline, but 'Real-time endpoint' alone does not describe the multi-container capability being asked about.

49
MCQeasy

A company wants to update an existing SageMaker real-time endpoint to serve a new model version. They need to route a small percentage of traffic to the new version initially and monitor for errors before switching fully. Which deployment pattern supports this?

A.Shadow testing
B.A/B testing with traffic splitting
C.Canary deployment with weighted production variants
D.Blue/green deployment
AnswerC

Canary deployment shifts a small percentage of endpoint traffic to the new model variant while the remainder serves the existing version, letting you monitor error metrics before full cutover. Weighted production variants provide exactly this gradual traffic routing.

Why this answer

SageMaker real-time endpoints support canary deployments by configuring multiple production variants with weighted traffic distribution. You can assign a small weight (e.g., 5%) to the new model version variant and 95% to the existing one, then monitor CloudWatch metrics for errors before shifting all traffic to the new variant. This matches the requirement for a gradual, monitored rollout.

Exam trap

Candidates often mistakenly choose blue/green deployment because it sounds like a safe rollout, but it lacks the gradual traffic shifting required for monitoring a small percentage first.

How to eliminate wrong answers

Option A is wrong because shadow testing (also called mirroring) sends a copy of live traffic to the new model without affecting the live response, but SageMaker does not natively support shadow testing for real-time endpoints; it is typically used for testing without routing any user-facing traffic. Option B is wrong because A/B testing with traffic splitting is a broader concept that could be implemented via weighted variants, but the specific pattern described in the question (routing a small percentage of traffic to a new version and monitoring before switching fully) is precisely a canary deployment, not just any A/B test. Option D is wrong because blue/green deployment involves switching all traffic at once from the old (blue) to the new (green) environment, which does not allow for a small percentage of traffic to be routed initially for monitoring.

50
MCQmedium

An organization wants to ensure that only approved model versions can be deployed to production. They use the SageMaker Model Registry to track model versions. How can they enforce that only approved models are deployed?

A.Manually review each model before deployment
B.Use SageMaker Model Monitor to check model quality after deployment
C.Use IAM policies to restrict deployment to only Approved model versions
D.Store model metadata in a DynamoDB table and check it before deployment
AnswerC

IAM policies can be written to allow SageMaker CreateEndpoint only for models with an Approved approval status, which is best practice.

Why this answer

AWS IAM policies can be used to conditionally restrict SageMaker API actions (e.g., CreateEndpointConfig, CreateModel) based on the model version's approval status. By evaluating the `sagemaker:ModelPackageApprovalStatus` condition key in an IAM policy, you can enforce that only model versions with an `Approved` status can be deployed, providing a native, automated, and auditable enforcement mechanism without manual intervention or external dependencies.

Exam trap

The trap here is that candidates confuse SageMaker Model Monitor (post-deployment monitoring) with pre-deployment approval enforcement, or they assume custom external checks (DynamoDB) are necessary when SageMaker provides native IAM-based conditional enforcement.

How to eliminate wrong answers

Option A is wrong because manual review is not a technical enforcement mechanism; it introduces human error, lacks auditability, and does not prevent unauthorized deployments via API or automation. Option B is wrong because SageMaker Model Monitor is a post-deployment tool that detects data drift and quality issues after the model is already serving traffic; it cannot prevent the deployment of unapproved models. Option D is wrong because storing metadata in DynamoDB and checking it before deployment requires custom code, introduces latency, and is not a native SageMaker enforcement mechanism; it also bypasses the built-in approval tracking in the Model Registry.

51
MCQeasy

A company uses SageMaker Neo to compile a trained model for deployment on edge devices. What is the primary benefit of using Neo?

A.It monitors model drift in production
B.It reduces model size and improves inference speed on target hardware
C.It automatically retrains the model on new data
D.It provides a serverless inference endpoint
AnswerB

SageMaker Neo compiles models into optimised executables for specific target hardware, reducing model size and improving inference latency and throughput on edge devices. This satisfies the edge deployment constraint, where resource limits make unoptimised frameworks impractical.

Why this answer

SageMaker Neo compiles trained models into optimized executables for specific target hardware (CPU, GPU, or edge accelerators), producing smaller artifacts and faster inference by leveraging hardware-specific instruction sets. It is designed for edge and constrained deployments where latency, memory, and compute are limited. Neo does not handle monitoring, retraining, or endpoint provisioning.

Exam trap

MLA-C01 often tests confusion between SageMaker features — candidates pick Model Monitor or Serverless Inference when the question is specifically about compiling models for edge hardware with Neo.

How to eliminate wrong answers

Option A is wrong because model drift monitoring is handled by SageMaker Model Monitor, not Neo. Option C is wrong because automatic retraining is a pipeline concern (SageMaker Pipelines, Clarify, or custom MLOps), not a compilation feature. Option D is wrong because serverless inference endpoints are provided by SageMaker Serverless Inference, which is unrelated to Neo's compilation role.

52
MCQeasy

A data science team needs to deploy a trained PyTorch model for real-time inference with sub-100ms latency. The model fits on a single GPU. Which SageMaker inference option is MOST cost-effective while meeting the latency requirement?

A.SageMaker Batch Transform
B.SageMaker real-time endpoint on ml.g4dn.xlarge
C.SageMaker Async Inference
D.SageMaker Serverless Inference
AnswerB

A SageMaker real-time endpoint on ml.g4dn.xlarge provides a persistent, GPU-backed inference host with low single-digit millisecond overhead, meeting sub-100ms latency. Since the model fits one GPU, this single-instance option is more cost-effective than multi-GPU or serverless alternatives.

Why this answer

SageMaker real-time endpoints provide dedicated, persistent instances that can handle synchronous inference with sub-100ms latency. The ml.g4dn.xlarge instance includes a single NVIDIA T4 GPU, which is sufficient for the model size and offers the lowest cost among GPU instances that meet the latency requirement. This option balances performance and cost for real-time, low-latency inference.

Exam trap

The trap here is that candidates often choose SageMaker Serverless Inference for its cost-saving potential, but they overlook the cold start latency and lack of GPU support, which makes it unsuitable for real-time, sub-100ms inference with PyTorch models.

How to eliminate wrong answers

Option A is wrong because SageMaker Batch Transform is designed for asynchronous, offline inference on large datasets, not for real-time sub-100ms latency; it processes data in batches and returns results only after the job completes. Option C is wrong because SageMaker Async Inference queues inference requests and processes them asynchronously, which introduces unpredictable latency and is not suitable for sub-100ms real-time requirements. Option D is wrong because SageMaker Serverless Inference auto-scales from zero and has a cold start latency that can exceed 100ms, especially for GPU-based models, making it unsuitable for strict real-time latency demands.

53
MCQmedium

A company wants to deploy a scikit-learn model to a SageMaker AI real-time endpoint. The model must be loaded from a custom Python module that contains preprocessing logic not present in the built-in scikit-learn container. The team wants to minimize operational overhead and does not need to change system-level libraries. Which approach should the engineer take?

A.Package the preprocessing logic in an inference.py file and pass it as the entry_point to a SageMaker AI framework estimator using the scikit-learn framework.
B.Use the SageMaker AI built-in scikit-learn container without an entry point and rely on the default inference handler.
C.Build a fully custom Docker image with a Bring Your Own Container (BYOC) approach and push it to Amazon ECR.
D.Deploy the model with SageMaker AI batch transform and invoke it from the application on demand.
AnswerA

SageMaker AI framework estimators support an entry_point script that defines model_fn and input_fn or transform_fn. This lets the engineer add custom Python preprocessing while reusing the managed scikit-learn container, which minimizes operational overhead because no Docker image must be built or maintained.

Why this answer

A SageMaker AI framework estimator with an entry_point script lets the team inject custom Python preprocessing while reusing the managed scikit-learn container. This satisfies the functional requirement with far less operational effort than building and maintaining a custom Docker image.

Exam trap

The trap here is defaulting to a fully custom container whenever any custom code is needed, even when an entry point script on a managed framework container is sufficient.

54
MCQmedium

An ML engineer has a real-time SageMaker endpoint serving a fraud-detection model. The team wants to release a new model version to a small percentage of live traffic first, monitor CloudWatch metrics for accuracy regressions, and roll back quickly if performance degrades. They also want the production and candidate variants to share the same endpoint so latency comparisons are apples-to-apples. Which SageMaker deployment strategy should they use?

A.Update the existing endpoint in place by replacing the production variant's model with the new model artifact.
B.Configure the endpoint with two production variants (current and candidate) and set initial variant weights, using CloudWatch alarms to trigger rollback.
C.Create a second endpoint with the new model and use Route 53 weighted routing to split traffic.
D.Deploy the new model as a production variant on the existing endpoint and configure a shadow variant with zero traffic weight.
AnswerB

SageMaker supports multiple production variants on a single endpoint with per-variant traffic weights, so the current model keeps most traffic while the candidate receives a small percentage. Both variants share the endpoint's instances, making latency comparison fair. CloudWatch alarms on model or invocation metrics can invoke automatic rollback via Deployment Guardrails, satisfying the gradual rollout and quick-recovery requirements.

Why this answer

A single SageMaker endpoint can host multiple production variants, each with its own model artifact and a traffic weight that you adjust without redeploying infrastructure. This provides canary-style exposure, apples-to-apples latency on shared instances, and a fast rollback path when CloudWatch alarms fire. It is the native SageMaker mechanism for controlled model releases with live monitoring, matching every stated requirement.

Exam trap

The trap here is assuming that any two-model setup provides safe canary rollout, when separate endpoints or shadow variants cannot send a controlled fraction of live responses to the candidate.

55
MCQmedium

A machine learning team has a model that needs to serve predictions with very low latency (under 10 ms) for a real-time web application. The model is a small ensemble of three neural networks that fits in memory. Which SageMaker inference option is MOST appropriate?

A.SageMaker batch transform
B.SageMaker real-time endpoint
C.SageMaker asynchronous inference
D.SageMaker serverless inference
AnswerB

Real-time endpoints keep the model loaded on persistent instances and return predictions synchronously, avoiding the cold-start and queueing overhead of serverless inference. For a small in-memory ensemble needing sub-10 ms responses, this persistent hosting meets the latency requirement.

Why this answer

SageMaker real-time endpoints are designed for low-latency, synchronous inference, making them the best fit for a model that must serve predictions in under 10 ms. Since the ensemble of three neural networks fits in memory, a real-time endpoint can keep the model loaded and respond to each request with minimal overhead, typically using HTTPS and the SageMaker InvokeEndpoint API.

Exam trap

The trap here is that candidates confuse 'low latency' with 'serverless' or 'asynchronous' options, not realizing that serverless inference has cold starts and asynchronous inference adds queueing delays, both of which break the sub-10 ms requirement.

How to eliminate wrong answers

Option A is wrong because SageMaker batch transform is an asynchronous, offline inference option that processes large datasets in batches and does not provide real-time, low-latency responses. Option C is wrong because SageMaker asynchronous inference is designed for requests with large payloads or long processing times, and it introduces queueing and callback mechanisms that add latency beyond the 10 ms requirement. Option D is wrong because SageMaker serverless inference auto-scales from zero and has a cold-start latency that can exceed 10 ms, making it unsuitable for sub-10 ms real-time predictions.

56
MCQmedium

A team needs to deploy a new model version to production while minimizing risk. They want to route 5% of live traffic to the new model and 95% to the current model, and then gradually increase the new model's traffic. Which SageMaker deployment pattern should they use?

A.Shadow testing
B.Blue/green deployment
C.A/B testing with production variants
D.Canary deployment using production variants
AnswerD

Canary deployment using production variants lets you split live traffic across multiple model variants behind one endpoint, initially weighting 5% to the new version and 95% to the current one. You then shift the traffic distribution gradually, satisfying the low-risk, incremental rollout constraint.

Why this answer

Canary deployment using production variants in Amazon SageMaker allows you to route a small percentage of live traffic (e.g., 5%) to the new model version while the rest goes to the existing model. You can then gradually increase the traffic to the new model as confidence grows, minimizing risk. This pattern is specifically designed for progressive rollouts with the ability to roll back if issues arise.

Exam trap

The trap is confusing canary deployment with A/B testing; both use production variants, but canary is for gradual traffic shifting to minimize risk, while A/B testing is for comparing model performance with a fixed split.

How to eliminate wrong answers

Option A is wrong because shadow testing mirrors traffic to the new model without affecting production responses, so it does not route live traffic to the new model. Option B is wrong because blue/green deployment shifts all traffic at once after testing, not gradually. Option C is wrong because A/B testing with production variants is used to compare model performance by splitting traffic, but it typically involves a fixed split for experimentation, not a gradual increase for safe deployment.

57
MCQhard

An ML platform team must orchestrate a workflow that trains a model, evaluates it against a baseline, and only registers the model if evaluation passes. If evaluation fails, the workflow must notify the data science channel and stop without registering. The team wants the orchestration logic to be expressed as code, versioned in git, and integrated with SageMaker training jobs and Model Registry. Which approach best fits these requirements?

A.Author a SageMaker Pipeline with a Condition step that compares evaluation metrics to the baseline and branches to a RegisterModel step or a Fail step.
B.Write a Python script on an EC2 instance that calls the SageMaker APIs sequentially and uses if/else statements to decide whether to register.
C.Use AWS Glue workflows to chain the training and evaluation jobs and branch on the evaluation result.
D.Define an AWS Step Functions state machine with a Choice state that inspects evaluation output and calls SageMaker APIs directly.
AnswerA

SageMaker Pipelines are defined as code, can be versioned in git, and natively integrate training jobs, processing jobs, and the Model Registry. A Condition step compares the evaluation metric to the baseline and routes execution to either RegisterModel or a terminal Fail step, enforcing the gate automatically. This expresses the entire orchestration logic declaratively and reproducibly, matching all stated requirements.

Why this answer

SageMaker Pipelines express orchestration as versionable code and include a Condition step for metric-based branching, plus native RegisterModel and Fail steps. This enforces the evaluation gate before registration without custom glue, and the pipeline definition can live in git and be executed reproducibly. It best satisfies the code-as-orchestration, integration, and versioning requirements.

Exam trap

The trap here is assuming any general-purpose orchestrator with an if/else can substitute for native SageMaker pipeline gating and Model Registry integration.

58
MCQeasy

An ML engineer needs to compile a trained TensorFlow model to run efficiently on a target edge device with an ARM CPU. Which AWS service should they use?

A.SageMaker Debugger
B.AWS Inferentia
C.SageMaker Neo
D.Amazon Elastic Inference
AnswerC

SageMaker Neo compiles trained models for specific target hardware, including ARM CPUs, producing optimised executables that run efficiently on edge devices. It directly satisfies the stem's requirement to compile a TensorFlow model for an ARM-based edge target.

Why this answer

SageMaker Neo compiles trained models from frameworks like TensorFlow, PyTorch, and MXNet into optimized executables for specific target hardware, including ARM CPUs, Intel, and AWS Inferentia. It performs graph-level optimizations and generates a runtime that runs efficiently on the edge device.

Exam trap

MLA-C01 often tests the distinction between hardware accelerators (Inferentia, Elastic Inference) and the compilation/optimization service (Neo) — candidates pick the chip when the question asks for a service to compile a model.

How to eliminate wrong answers

Option A is wrong because SageMaker Debugger is a training-time tool for inspecting tensors, gradients, and metrics — it does not compile or optimize models for edge deployment. Option B is wrong because AWS Inferentia is a custom inference chip, not a compilation service; it is a hardware target, not a tool for compiling TensorFlow models. Option D is wrong because Amazon Elastic Inference attaches GPU-powered inference acceleration to EC2/SageMaker instances in the cloud, not to ARM edge devices, and it is being deprecated.

59
MCQmedium

A company runs a real-time fraud detection model on a SageMaker endpoint. The model is updated weekly, and each update must be validated against live traffic without affecting existing predictions. The team wants to compare the new model's performance against the current model using a small percentage of incoming requests, while ensuring that the current model continues to serve the majority of traffic. Which SageMaker deployment strategy should they use?

A.A/B testing with production variants, where the new model variant receives a small portion of traffic.
B.Canary deployment, where the new model is deployed to a separate endpoint and traffic is gradually shifted using a load balancer.
C.Blue/green deployment, where traffic is shifted all at once from the old model to the new model after validation.
D.Shadow testing, where the new model receives a copy of live traffic but its predictions are not returned to the caller.
AnswerA

A/B testing with production variants allows you to deploy multiple models to the same endpoint and distribute traffic between them. By assigning a small weight to the new variant, you can evaluate its performance on live traffic while the existing model handles the rest. This meets the requirement of validating the new model without impacting the majority of predictions.

Why this answer

SageMaker production variants enable A/B testing by allowing multiple models on a single endpoint with configurable traffic weights. This lets the team direct a small percentage of live traffic to the new model while the current model serves the rest, facilitating performance comparison without disrupting service. Other strategies either do not return predictions or switch all traffic at once.

Exam trap

The trap here is confusing shadow testing with A/B testing, as shadow testing does not return predictions to the caller and thus cannot be used for live performance validation.

60
MCQeasy

A hospital's ML team deploys a diagnostic model to a SageMaker real-time endpoint. Compliance requires that every inference request and response be recorded for later auditing, and the records must be retrievable months later. The team needs a low-effort way to capture this data. Which approach should they use?

A.Turn on AWS CloudTrail data events for the endpoint and deliver them to an S3 bucket.
B.Enable data capture on the endpoint with a sampling percentage of 100 and an S3 destination, and apply an S3 lifecycle policy to retain the objects for the required period.
C.Enable Amazon CloudWatch metrics on the endpoint and export the metrics to Amazon S3 for archival.
D.Add application code that writes each request and response to Amazon CloudWatch Logs and set the log group retention to the audit period.
AnswerB

Data capture records request and response payloads to Amazon S3 automatically, and setting the sampling percentage to 100 captures every invocation, satisfying an audit requirement for complete records. Pointing capture at an S3 prefix and applying a lifecycle policy retains the objects for the mandated duration. This requires minimal application changes and is the purpose-built feature for this need.

Why this answer

SageMaker data capture is the purpose-built feature for recording inference request and response payloads to Amazon S3, and a 100 percent sampling rate ensures every invocation is stored for audit. Pairing the capture destination with an S3 lifecycle policy retains records for the required period at low cost. Metrics, CloudTrail, and application-level logging do not capture full payloads as directly.

Exam trap

The trap here is confusing request-level API auditing via CloudTrail or numeric monitoring via CloudWatch metrics with full payload capture, which only data capture provides.

61
MCQhard

A data science team uses SageMaker Pipelines to orchestrate their ML workflow. They noticed that even when source data hasn't changed, the pipeline re-runs all steps, wasting compute time. What should they enable to avoid redundant runs?

A.Enable pipeline caching by setting the CacheConfig property for each step
B.Configure the pipeline to run on a schedule instead of on-demand
C.Use the Parameter step to pass previous execution ID
D.Use Lambda step to check data changes before running
AnswerA

CacheConfig on each pipeline step lets SageMaker Pipelines skip re-execution when inputs, code and parameters are unchanged, reusing prior outputs. Enabling it satisfies the stem's requirement to avoid redundant runs and wasted compute when source data has not changed.

Why this answer

SageMaker Pipelines supports step caching via the `CacheConfig` property. When enabled, the pipeline checks if the step's inputs (including source data, parameters, and code) have changed since the last successful run. If no changes are detected, the step is skipped and the previous output is reused, eliminating redundant compute.

Exam trap

The trap here is that candidates may think caching requires external logic (like a Lambda step) or scheduling, when SageMaker Pipelines has a native `CacheConfig` property that directly addresses redundant runs with minimal configuration.

How to eliminate wrong answers

Option B is wrong because scheduling the pipeline does not prevent redundant runs; it only triggers execution at fixed intervals, which could still re-run all steps even when data hasn't changed. Option C is wrong because passing a previous execution ID via a Parameter step does not enable caching; it merely provides a reference but does not automatically skip unchanged steps. Option D is wrong because using a Lambda step to check data changes before running adds custom logic but is not a built-in mechanism for step-level caching; SageMaker Pipelines already provides `CacheConfig` for this purpose, making a Lambda workaround unnecessary and less efficient.

62
MCQmedium

A team uses SageMaker Pipelines to automate retraining. They want to skip the training step if the data has not changed since the last run. Which feature should they enable?

A.Parameterized executions
B.Lineage tracking
C.Step caching
D.Condition step with a custom check
AnswerC

Step caching stores the step's inputs and outputs, so when the data and hyperparameters match a previous run, SageMaker reuses the cached artefacts and skips execution entirely. This directly satisfies the requirement to bypass retraining when the dataset is unchanged, avoiding redundant compute costs.

Why this answer

Step caching in SageMaker Pipelines allows you to reuse the output from a previous execution of a step if its input data and configuration parameters have not changed. By enabling caching on the training step, the pipeline automatically skips re-executing that step when the data is identical, saving time and cost. This directly addresses the requirement to skip retraining when data has not changed.

Exam trap

Candidates often confuse step caching (automatic, built-in) with a Condition step (manual, custom logic), leading them to overthink and choose the more complex option D when the simpler caching feature is the correct answer.

How to eliminate wrong answers

Option A is wrong because parameterized executions allow you to pass different parameters into a pipeline run, but they do not automatically skip steps based on data changes; they simply enable dynamic input values. Option B is wrong because lineage tracking records the relationships between artifacts and steps for governance and reproducibility, but it does not provide any mechanism to skip step execution. Option D is wrong because while a Condition step can branch pipeline execution based on a custom check, it requires you to implement the logic to compare data versions manually, whereas step caching provides built-in, automatic detection of unchanged inputs.

63
MCQhard

A machine learning engineer deploys a new model version to a SageMaker endpoint with production variants. They want to gradually shift traffic from the old model to the new model, monitoring for errors, and automatically roll back if the error rate exceeds 5%. Which deployment pattern should they use?

A.Canary deployment with CloudWatch alarms
B.A/B testing with traffic splitting
C.Blue/green deployment
D.Shadow testing
AnswerA

A canary deployment routes a small percentage of traffic to the new variant while CloudWatch alarms watch the error rate, triggering automatic rollback when it exceeds 5%. It satisfies the stem's gradual-shift and auto-rollback constraints, unlike all-at-once or shadow patterns.

Why this answer

Canary deployment with CloudWatch alarms is correct because SageMaker production variants allow you to split traffic between the old and new model (e.g., 90/10), and CloudWatch alarms can monitor the new variant's error rate and trigger automatic rollback when it exceeds 5%. This pattern is purpose-built for gradual, monitored traffic shifting with automated safety nets.

Exam trap

MLA-C01 often tests the distinction between canary (gradual shift + auto rollback) and A/B testing (statistical comparison) — candidates pick A/B testing because both involve traffic splitting, but only canary is designed for progressive rollout with automated rollback on error thresholds.

How to eliminate wrong answers

Option B is wrong because A/B testing with traffic splitting is designed to compare model performance statistically (e.g., conversion rates) rather than to gradually shift all traffic with automatic rollback on error thresholds. Option C is wrong because blue/green deployment shifts traffic all at once (or in a single cutover) between two identical environments, which does not provide the gradual, monitored ramp-up the scenario requires. Option D is wrong because shadow testing mirrors live traffic to the new model without serving its responses to users, so it cannot be used to gradually shift production traffic or trigger rollback based on user-facing error rates.

64
MCQmedium

A company has 200 small PyTorch models that are each used infrequently but need to be available for real-time inference. To minimize costs, they want to host all models on a single endpoint. Which SageMaker feature should they use?

A.Multi-model endpoint (MME)
B.Multi-container endpoint
C.Batch Transform job
D.Asynchronous inference endpoint
AnswerA

Multi-model endpoint hosts many models on one endpoint, loading each from S3 on invocation and caching it in memory. Because the PyTorch models are infrequent, this satisfies the single-endpoint and cost-minimisation constraints without provisioning per-model hosting.

Why this answer

SageMaker multi-model endpoints (MME) let a single endpoint host hundreds or thousands of models behind one container, loading each model into memory or disk on demand and unloading idle ones. This is purpose-built for the scenario of many small, infrequently used models that must still serve real-time inference, because you pay for one endpoint's worth of instances rather than one endpoint per model. The SageMaker SDK/API passes a TargetModel parameter on each InvokeEndpoint call so the endpoint knows which model to load and serve.

Exam trap

MLA-C01 often tests the confusion between multi-model endpoints (many models, one container, dynamic load) and multi-container endpoints (few containers, different frameworks, static), so candidates who see 'multiple models' and pick multi-container get it wrong.

How to eliminate wrong answers

Option B (Multi-container endpoint) is wrong because multi-container endpoints are designed to run a small number of containers (typically up to 15) that each host a different framework or pipeline stage, and they do not dynamically load/unload hundreds of models — they are for heterogeneous inference pipelines, not model fleets. Option C (Batch Transform) is wrong because it is an offline, batch-scoring service that spins up a job, processes a dataset, and tears down; it cannot provide the always-available real-time inference the company requires. Option D (Asynchronous inference endpoint) is wrong because async inference is for large payloads/long processing times with queued requests, not for hosting many small models on one endpoint — it still requires one endpoint per model or a single model per endpoint.

65
Multi-Selectmedium

An organization wants to automate ML retraining using an event-driven architecture. Which THREE services should they combine? (Select THREE.)

Select 3 answers
A.SageMaker (training jobs or pipelines)
B.Amazon EventBridge
C.AWS Lambda
D.AWS Glue
E.Amazon CloudWatch Logs
AnswersA, B, C

SageMaker training jobs or pipelines execute the actual retraining computation when triggered. Combined with event detection and orchestration services, it forms the event-driven retraining chain, satisfying the stem's automation requirement. It performs model training rather than storing data or routing events.

Why this answer

Amazon EventBridge [CORRECT] is the event-driven backbone: it receives events (e.g., new data in S3, a schedule, or a custom PutEvents call) and routes them to targets via rules, which is exactly what an event-driven retraining architecture requires. AWS Lambda [CORRECT] acts as the lightweight, serverless compute glue that responds to those EventBridge events and invokes the training workflow, so no servers need to be managed. SageMaker (training jobs or pipelines) [CORRECT] performs the actual model retraining and orchestration, since SageMaker Pipelines can be triggered programmatically to run training jobs and register updated models.

AWS Glue is a data-integration/ETL service and, while useful for preparing data, it is not required to deliver the event-driven retraining trigger or the training itself. Amazon CloudWatch Logs is for storing and monitoring log data and does not provide event routing or ML training capability.

Exam trap

The trap here is that candidates often confuse AWS Glue as a compute trigger for ML retraining, but Glue is designed for batch ETL and lacks the event-driven, low-latency invocation capabilities required for this architecture.

66
MCQeasy

A company wants to deploy a trained XGBoost model for batch inference on a large dataset stored in S3. The inference job should be cost-effective and does not require real-time responses. Which SageMaker inference option should they use?

A.SageMaker Batch Transform
B.SageMaker real-time endpoint
C.SageMaker Asynchronous Inference
D.SageMaker Serverless Inference
AnswerA

SageMaker Batch Transform runs inference over large S3 datasets on managed instances that terminate when the job completes, avoiding the cost of a persistently running endpoint. This satisfies the stem's cost-effectiveness requirement and its lack of any real-time response need.

Why this answer

SageMaker Batch Transform is designed for offline, high-throughput inference on large datasets stored in S3, with no persistent endpoint and no real-time requirement. It is the most cost-effective option for this scenario because you pay only for the duration of the batch job.

Exam trap

MLA-C01 often tests the cost/latency trade-off between Batch Transform and Asynchronous Inference — candidates pick Asynchronous because it sounds 'batch-like,' but Asynchronous Inference is for near-real-time large-payload requests, not bulk offline scoring.

How to eliminate wrong answers

Option B is wrong because real-time endpoints are always-on and billed continuously, making them cost-inefficient for non-real-time batch workloads. Option C is wrong because Asynchronous Inference is designed for large payloads or long processing times with near-real-time queuing, not for bulk offline scoring of an entire S3 dataset. Option D is wrong because Serverless Inference is for intermittent, unpredictable real-time requests with cold-start tolerance, not for large-scale batch jobs.

67
MCQeasy

A machine learning engineer needs to deploy a new version of a model gradually, initially sending 5% of traffic to the new version and 95% to the current version, while monitoring for errors. Which deployment pattern should they use?

A.Blue/green deployment
B.Canary deployment
C.Shadow testing
D.Rolling deployment
AnswerB

A canary deployment shifts a small percentage of traffic, here 5%, to the new model version while the remainder stays on the current version, allowing error monitoring before full rollout. This matches the gradual, low-risk traffic-splitting requirement in the stem.

Why this answer

Canary deployment is the correct pattern because it allows the ML engineer to route a small percentage of traffic (e.g., 5%) to the new model version while keeping the majority (95%) on the current version. This enables gradual rollout with real-time monitoring for errors, and if issues are detected, traffic can be instantly shifted back to the stable version.

Exam trap

Candidates often confuse canary deployment with blue/green deployment, thinking both involve gradual traffic shifting. However, blue/green is an all-or-nothing switch between environments, while canary allows incremental percentage-based routing (e.g., 5%) with real-time monitoring and automated rollback.

How to eliminate wrong answers

Option A is wrong because blue/green deployment involves switching all traffic at once from the current environment (blue) to the new environment (green), which does not support gradual traffic shifting or incremental error monitoring. Option C is wrong because shadow testing sends a copy of live traffic to the new model without affecting user-facing responses, but it does not route actual user traffic to the new version, so it cannot be used to gradually shift real traffic percentages. Option D is wrong because rolling deployment updates instances incrementally (e.g., replacing pods one by one), but it does not provide fine-grained traffic splitting like 5% vs 95% and lacks the instant rollback capability of canary deployments.

68
MCQhard

A team uses SageMaker real-time endpoints for inference. They want to deploy a new model version and compare its performance with the current version under live traffic without affecting user experience. Which method should they use?

A.A/B testing with production variant traffic splitting
B.Batch transform on a holdout test set
C.Blue/green deployment
D.Shadow testing with SageMaker
AnswerD

Shadow testing deploys the new model variant alongside the production variant, mirroring live traffic to it without returning its responses to users. This satisfies the stem's requirement to compare performance under live traffic while leaving user experience unaffected.

Why this answer

Shadow testing with SageMaker deploys the new model alongside the production model and mirrors live traffic to it without returning its responses to users. This lets the team compare performance under real traffic with zero impact on user experience.

Exam trap

MLA-C01 often tests the difference between shadow testing (no user impact, mirrored traffic) and A/B testing (real user impact, split traffic) — candidates pick A/B testing because both compare models under live traffic, but only shadow testing guarantees zero user exposure.

How to eliminate wrong answers

Option A is wrong because A/B testing with production variant traffic splitting sends a portion of real user traffic to the new model and returns its responses, which does affect user experience. Option B is wrong because Batch Transform on a holdout set evaluates offline on historical data, not live traffic, so it does not reflect real production conditions. Option C is wrong because blue/green deployment shifts live traffic to the new model, which directly affects users and is not a silent comparison.

69
MCQhard

A company is deploying a large language model (LLM) to a SageMaker endpoint. They want to minimize inference latency and cost by using GPU acceleration and model parallelism. The model is too large to fit on a single GPU. Which SageMaker feature should they use?

A.SageMaker Elastic Inference
B.SageMaker distributed data parallel library
C.SageMaker model parallelism library
D.SageMaker Inference Recommender
AnswerC

The SageMaker model parallelism library enables training and inference of large models that cannot fit on a single GPU by partitioning the model across multiple GPUs. It supports tensor parallelism and pipeline parallelism, and it can be used for inference to reduce latency and cost. This directly addresses the need for model parallelism with GPU acceleration.

Why this answer

The SageMaker model parallelism library is designed to partition large models across multiple GPUs, enabling inference for models that exceed a single GPU's memory. It supports various parallelism strategies and can reduce latency and cost. The other options are either for training, deprecated, or only provide recommendations, not the required model parallelism.

Exam trap

The trap here is confusing the distributed data parallel library with model parallelism; data parallel is for training and does not split the model.

70
MCQhard

A machine learning team runs a SageMaker AI Pipeline that trains a model and registers it in the SageMaker AI Model Registry. A separate deployment process must promote the model to production only after a human reviewer approves the model version. The team wants to automate the promotion so that approval in the Model Registry triggers the deployment without manual intervention. Which combination of steps should the engineer implement?

A.Use a SageMaker AI Projects template that automatically deploys every model version as soon as it is registered.
B.Schedule an AWS Lambda function every minute to call DescribeModelPackage and deploy any approved model version.
C.Add a ConditionStep to the existing pipeline that checks the model version status and deploys the model.
D.Configure an Amazon EventBridge rule for the SageMaker AI Model Registry model version state change to Approved, and use it to start a deployment workflow.
AnswerD

The SageMaker AI Model Registry emits state change events when a model version moves to Approved. An EventBridge rule can match that event and invoke a target such as AWS Step Functions or AWS CodePipeline to run the deployment. This automates promotion only after the human approval step, satisfying the requirement without polling.

Why this answer

Model approval in the SageMaker AI Model Registry changes the model version state and emits an event. Capturing that event with Amazon EventBridge and routing it to a deployment workflow links the human approval gate to automated promotion, which is exactly the required behavior.

Exam trap

The trap here is trying to enforce a post-pipeline human approval inside the pipeline itself instead of reacting to the Model Registry state change event.

71
MCQmedium

A machine learning engineer is deploying a model to a SageMaker endpoint that must handle occasional large payloads up to 1 GB. The inference time can take up to 10 minutes. The team wants to minimize cost and avoid idle compute. Which deployment option is most appropriate?

A.SageMaker real-time endpoint with automatic scaling
B.SageMaker Serverless Inference
C.SageMaker Asynchronous Inference
D.SageMaker batch transform
AnswerC

Asynchronous Inference supports payloads up to 1 GB and allows inference times up to 15 minutes. It queues requests and processes them asynchronously, returning results via Amazon S3. This makes it ideal for large payloads and long-running inference, and it can scale to zero when idle, minimizing cost.

Why this answer

Asynchronous Inference is designed for large payloads (up to 1 GB) and long inference times (up to 15 minutes). It queues requests and returns results via S3, and it can scale to zero when idle, reducing cost. Real-time and serverless inference have payload and timeout limits, and batch transform is for offline batch processing, not on-demand requests.

Exam trap

The trap here is assuming that real-time endpoints can handle any payload size, but they are limited to 6 MB and are cost-inefficient for long-running jobs.

72
MCQmedium

A company wants to deploy a PyTorch model on SageMaker using the NVIDIA Triton Inference Server for GPU acceleration. They have an existing Triton configuration. Which approach should they take?

A.Use SageMaker Neo to compile the model for Triton
B.Package Triton as a custom container and use SageMaker batch transform
C.Use the SageMaker Triton Inference Server container from the Deep Learning Containers
D.Use the standard SageMaker PyTorch container and install Triton at runtime
AnswerC

The SageMaker Triton Deep Learning Container ships NVIDIA Triton Inference Server preinstalled and configured, so the existing Triton model configuration can be deployed directly. This satisfies the GPU acceleration requirement without building or maintaining a custom container image.

Why this answer

AWS provides a pre-built SageMaker Triton Inference Server container as part of the Deep Learning Containers (DLCs), which is optimized for GPU acceleration and supports the existing Triton configuration without modification. This container integrates directly with SageMaker hosting endpoints, enabling seamless deployment of PyTorch models with Triton's features like dynamic batching and model concurrency.

Exam trap

The trap here is that candidates may assume SageMaker Neo is a universal compilation tool for any inference server, but Neo is specifically for hardware-specific optimization and does not support Triton's runtime environment, leading them to incorrectly select Option A.

How to eliminate wrong answers

Option A is wrong because SageMaker Neo compiles models for specific hardware targets (e.g., Intel, ARM) and does not support compilation for the NVIDIA Triton Inference Server; Neo is designed for edge devices and does not integrate with Triton's serving architecture. Option B is wrong because while packaging Triton as a custom container is possible, using SageMaker batch transform is not the recommended approach for real-time inference with GPU acceleration; batch transform is for offline, asynchronous processing, not for low-latency serving. Option D is wrong because installing Triton at runtime on the standard PyTorch container is inefficient and error-prone; it adds startup latency, may cause dependency conflicts, and bypasses the pre-optimized, tested Triton container that AWS provides.

73
MCQhard

A company needs to update a model in production without any downtime. They currently have a single real-time endpoint serving traffic. Which approach allows them to deploy a new model version and switch traffic gradually while being able to roll back quickly?

A.Use a canary deployment by creating a new production variant with the new model and shifting traffic incrementally
B.Use a multi-model endpoint and replace the model file
C.Stop the endpoint, update the model, and restart the endpoint
D.Update the existing endpoint's model directly using UpdateEndpoint
AnswerA

A canary deployment adds a new production variant holding the new model, then shifts traffic incrementally via variant weights. This satisfies the stem's no-downtime and gradual-shift requirements, and weights can be reverted instantly to roll back.

Why this answer

A canary deployment creates a new production variant on the existing endpoint with the new model and shifts a small percentage of traffic to it, allowing gradual validation and quick rollback by shifting traffic back to the original variant. This avoids downtime and provides controlled risk.

Exam trap

The trap is thinking 'update the endpoint' is sufficient — candidates must recognize that in-place updates cause downtime and lack gradual traffic control, whereas canary variants provide safe, reversible rollouts.

How to eliminate wrong answers

Option B is wrong because a multi-model endpoint hosts multiple models behind one endpoint but does not provide traffic-shifting or gradual rollout semantics for a new version. Option C is wrong because stopping the endpoint causes downtime, violating the no-downtime requirement. Option D is wrong because updating the existing endpoint's model directly replaces the model in place with no gradual traffic shift and no easy rollback path.

74
MCQeasy

An ML engineer has trained a model and stored the model artifacts in an Amazon S3 bucket in the same AWS Region as the planned SageMaker AI endpoint. During endpoint creation, the engineer must specify the S3 location of the model artifacts. Which permission must the SageMaker AI execution role have for the endpoint to load the model successfully?

A.s3:PutObject on the model artifact prefix in the bucket.
B.s3:DeleteObject on the model artifact prefix in the bucket.
C.s3:GetObject on the model artifact objects in the bucket.
D.s3:ListAllMyBuckets at the account level.
AnswerC

The SageMaker AI execution role is assumed by the hosting infrastructure to download model artifacts from Amazon S3 at container startup. Without s3:GetObject on the specific artifact objects, the container cannot retrieve the model and endpoint creation or invocation fails, so this permission is required.

Why this answer

The SageMaker AI execution role must be able to read the model artifacts from Amazon S3 when the container starts. Granting s3:GetObject on the artifact objects provides exactly that read access and follows least privilege.

Exam trap

The trap here is granting broad Amazon S3 permissions such as listing all buckets instead of the specific read access the endpoint actually needs.

75
MCQmedium

A data science team is using AWS Step Functions to orchestrate a machine learning workflow that includes a SageMaker training job followed by a model deployment. They want to ensure that if the training job fails, the workflow retries up to three times with exponential backoff before sending a notification to an Amazon SNS topic. Which Step Functions feature should they use to implement this?

A.A Map state that iterates over a list of retry attempts, invoking the training job each time until it succeeds.
B.A Choice state that checks the training job status and loops back to the training state if it failed, with a counter to limit attempts.
C.Retry and Catch fields on the training task state, with a retry policy specifying MaxAttempts and BackoffRate, and a Catch field that transitions to an SNS publish state.
D.A Parallel state that runs the training job and an SNS notification simultaneously, with a retry policy on the training job.
AnswerC

Step Functions allows you to define Retry and Catch on individual states. The Retry field can specify MaxAttempts, IntervalSeconds, and BackoffRate to implement exponential backoff. The Catch field can transition to a fallback state, such as an SNS publish task, when retries are exhausted. This directly meets the requirement.

Why this answer

Step Functions' Retry and Catch fields on a state provide built-in error handling. By configuring Retry with MaxAttempts and BackoffRate, the training task will be retried with exponential backoff. The Catch field can then route to an SNS publish state after all retries fail, ensuring notification only on persistent failure.

Exam trap

The trap here is using a Choice state or Map state to implement retries, which lacks native exponential backoff and complicates error handling compared to Retry and Catch.

Page 1 of 2 · 94 questions totalNext →

Ready to test yourself?

Try a timed practice session using only Deployment and Orchestration of ML Workflows questions.