MLA-C01 · domain
Deployment and Orchestration of ML Workflows
This domain covers building, automating, and operating ML workflows on AWS: SageMaker Pipelines for orchestration, Model Registry for versioning and approval, and deployment strategies like real-time endpoints, batch transform, and multi-variant traffic shifting. Questions present a concrete operational goal and ask which SageMaker feature, condition, or configuration achieves it.
Focused practice
Practice Deployment and Orchestration of ML Workflows questions
Scored sessions drawing only from this domain — pick a length below.
Start 20-question practice test →What this domain covers
What to know about Deployment and Orchestration of ML Workflows
Be able to wire a SageMaker Pipeline that trains, evaluates, and conditionally registers a model, then deploy it via a real-time endpoint using the right traffic-shifting and auto scaling configuration. The critical skill is matching each requirement to the correct SageMaker feature and parameter.
SageMaker Pipelines step conditions and PropertyFile-based conditional model registration in Model Registry
SageMaker endpoint deployment patterns: linear and canary traffic shifting across production variants
SageMaker Pipelines caching configuration to skip unchanged steps and reuse prior outputs
Multi-AZ highly available real-time endpoints with Application Auto Scaling on InvocationsPerInstance
Watch out for
Common Deployment and Orchestration of ML Workflows exam traps
- ▸Confusing Model Registry approval status with endpoint deployment; registering a model does not automatically serve it
- ▸Assuming Pipelines always re-execute steps; caching must be explicitly enabled per step to reuse outputs
- ▸Mixing up canary and linear traffic shifting, or forgetting that variant weights must sum to 100
Question index
All Deployment and Orchestration of ML Workflows questions (94)
Click any question to see the full explanation, or start a practice session above.
A team needs to deploy a model that has compliance requirements to log all inference requests and responses for auditing. The model will be served using a real-time endpoint. How can they achieve this without custom code?
Hard2An MLOps team is designing a SageMaker Pipeline to automate model retraining. The pipeline must: (1) run training only if new training data is available, (2) register the model in SageMaker Model Registry only if evaluation metrics exceed a threshold, (3) deploy the approved model to a staging endpoint automatically. Which THREE steps should they include? (Choose THREE.)
Hard3A company wants to run inference on a large dataset stored in S3 using a pre-trained model. The inference can tolerate latency from minutes to hours, and they want a fully managed solution that autoscales to handle large volumes. Which SageMaker inference option is most suitable?
Medium4A data scientist wants to compare the performance of two model versions (V1 and V2) in production by splitting traffic between them. They want to gradually increase the percentage of traffic to the new version while monitoring metrics. Which SageMaker feature enables this?
Easy5A data scientist needs to deploy a single ML model that will serve real-time predictions with low latency (under 10 ms) for a high-traffic web application. The model fits in memory and requires GPU acceleration. Which SageMaker inference option is MOST suitable?
Easy6A machine learning engineer has trained a model using SageMaker and wants to deploy it to a real-time endpoint. The engineer needs to specify the model artifacts, the inference code, and the environment. Which SageMaker resource should the engineer create first?
Easy7A team wants to deploy a new model using a canary deployment strategy on SageMaker. Which TWO configurations are necessary? (Choose two.)
Medium8A data scientist wants to version and manage trained models, require approval before deployment, and enable cross-account deployment. Which SageMaker feature provides these capabilities?
Easy9A company uses SageMaker Pipelines to orchestrate their ML workflow. They notice that if a pipeline step fails due to a transient error (e.g., a brief network issue), the entire pipeline fails and they must manually rerun from the beginning. They want to automatically retry failed steps a few times before failing. What should they do?
Hard10A company is using SageMaker to serve a model for real-time predictions. They want to test a new model version by routing a small percentage of live traffic to it while the rest goes to the current model. They also need to compare performance metrics. Which TWO actions should they take? (Select TWO.)
Medium11A machine learning engineer has trained a scikit-learn model and saved it as model.joblib in Amazon S3. The engineer wants SageMaker to host the model for real-time inference without writing a custom container or inference script, because the model uses only standard predict behavior. Which deployment approach should the engineer use?
Medium12An ML engineer needs to orchestrate a multi-step workflow that includes data preprocessing on Spark, model training on SageMaker, and deployment to a production endpoint. They require tight integration with other AWS services and the ability to add custom logic. Which AWS service should they use alongside SageMaker?
Medium13A team runs a SageMaker Pipeline that trains a model and registers it in the Model Registry. Compliance requires that the pipeline run automatically every time new labeled data lands in S3, and that each run record the exact S3 data prefix, the training image URI, and the git commit hash as lineage metadata. The engineer wants the least operational overhead. Which approach meets these requirements?
Hard14A machine learning engineer needs to optimize a trained TensorFlow model for deployment on edge devices with limited compute. Which SageMaker feature should they use to compile the model for target hardware?
Easy15An ML engineer wants to use MLflow on SageMaker to track experiments and log metrics. They have set up MLflow on an EC2 instance. How can they best integrate MLflow tracking with SageMaker training jobs?
Medium16A startup wants to deploy a model that has variable traffic patterns, with some periods of no traffic and occasional spikes. They want to pay only for what they use and do not want to manage instances. Which SageMaker inference option should they choose?
Medium17A machine learning team needs to deploy a new model version for A/B testing, gradually shifting traffic from the old version to the new version over 24 hours. Which deployment strategy should they use?
Medium18An ML team uses AWS Step Functions to orchestrate a retraining pipeline triggered by EventBridge when new training data arrives. The pipeline includes a SageMaker training job and a model evaluation. If evaluation fails, the team wants to send an alert. How should they implement this?
Hard19A company needs to deploy a large language model (LLM) on SageMaker with the Triton Inference Server to maximize GPU utilization and reduce latency. They have an NVIDIA A100 GPU. Which SageMaker inference option supports Triton?
Hard20A data science team uses SageMaker Pipelines for automated training. They need to conditionally register a model only if evaluation metrics exceed a threshold. Which pipeline step type should they use after the evaluation step?
Medium21A company runs a batch inference job on 10 TB of image data stored in S3. Each image needs to be processed by a GPU-accelerated model. The job is not time-sensitive and cost is the primary concern. Which SageMaker option is MOST appropriate?
Medium22A company wants to version and track ML models, with an approval workflow for promoting models from staging to production. Which SageMaker feature should they use?
Easy23A company is deploying a large NLP model on SageMaker for real-time inference. They want to reduce inference latency and cost by optimizing the model for the target hardware. The model is trained in PyTorch. Which SageMaker feature should they use to compile the model for best performance on the chosen instance?
Medium24A data scientist wants to train a model on SageMaker using a custom PyTorch script, then register the best model in the SageMaker Model Registry. The training job is part of a SageMaker Pipeline. Which pipeline step should be used to register the model?
Medium25A company is using SageMaker Pipelines to orchestrate their ML workflow. They have a Condition step that checks if a model's accuracy exceeds 0.9. If true, they want to register the model in the model registry; otherwise, they want to run a retraining step. Which step type should they use for the decision?
Medium26A company wants to deploy 50 small models (each ~100 MB) for real-time inference. They need to minimize hosting costs while maintaining low latency. Which SageMaker hosting option is most cost-effective?
Medium27A team is building a SageMaker Pipeline that trains a model and then registers it in the SageMaker Model Registry. They want the pipeline to automatically deploy the model to a real-time endpoint only after a human approves the model package. Which two actions should the team take to implement this approval-gated deployment? (Choose two.)
Hard28A data science team has trained a PyTorch model for real-time inference and needs to deploy it on AWS with GPU acceleration while minimizing cold-start latency. Which SageMaker inference option should they choose?
Easy29A data science team wants to host 50 different models for a recommendation engine. Each model is small (under 100 MB) and traffic patterns are unpredictable. They need to minimize cost and operational overhead. Which approach should they take?
Hard30A media company runs a real-time recommendation model on a SageMaker endpoint. Traffic triples every evening between 18:00 and 22:00 and drops to near zero overnight, and the team wants to cut costs without underserving evening users. They want the endpoint to scale out automatically as invocations rise and scale back in when traffic falls. Which solution should they implement?
Medium31A company wants to serve a large ensemble of models using NVIDIA Triton Inference Server on SageMaker for high throughput GPU inference. Which SageMaker inference option supports this?
Hard32A retailer runs a nightly batch scoring job that processes 40 GB of transaction data and writes predictions to S3. Occasionally a single partition is corrupt, causing the entire job to fail after several hours. The team wants the job to skip the corrupt partition, log which partition failed, and still complete processing of the remaining data with minimal changes to their existing SageMaker Processing job. Which change should they make?
Medium33A team uses SageMaker Pipelines to train and register a model. They want to conditionally run a hyperparameter tuning step only if the data quality check passes. Which pipeline step type should they use to branch the execution?
Hard34A machine learning team is deploying a model to a SageMaker endpoint and needs to implement A/B testing between two model versions. They want to split traffic 80/20 and monitor performance metrics for each variant. Which two actions should they take? (Choose two.)
Hard35A company needs to deploy a new model version to a SageMaker real-time endpoint. They want to route 5% of traffic to the new version initially to monitor for errors before full rollout. Which deployment strategy should they use?
Medium36A machine learning engineer needs to run a one-time scoring job over 500 GB of data stored in Amazon S3 using a trained model, and the results must be written back to S3. There is no requirement for a persistent HTTPS endpoint. Which SageMaker feature should the engineer use?
Easy37An ML team uses SageMaker Pipelines to automate model retraining. They want to skip redundant training steps when input data has not changed. Which feature should they enable?
Hard38A machine learning engineer is deploying a real-time inference endpoint on Amazon SageMaker AI for a fraud detection model. The model must serve predictions with consistent latency under 50 ms and the team expects traffic to fluctuate unpredictably, with occasional bursts. The engineer wants to automatically adjust the number of instances based on actual workload while minimizing cost during idle periods. Which SageMaker AI feature should the engineer configure?
Medium39A company wants to deploy a machine learning model using infrastructure as code to ensure reproducibility. They need to define the SageMaker Studio domain, user profiles, and the endpoint configuration. Which tool should they use?
Medium40Which SageMaker feature compiles a trained model into an optimized binary for a specific hardware target (e.g., Intel, ARM, NVIDIA, or edge devices) to improve inference performance?
Easy41A company needs to serve real-time predictions from a large ensemble of three deep learning models, each requiring different inference environments (PyTorch, TensorFlow, MXNet). Which SageMaker endpoint type supports running multiple inference containers together?
Medium42A fraud-detection team trains a model in SageMaker and wants to shift 10 percent of live prediction traffic to a newly retrained model to compare accuracy before a full cutover. The endpoint already serves the current model on one production variant. They need the endpoint to route a controlled fraction of requests to the new model without changing the client application. What should they do?
Medium43A company wants to serve 200 different PyTorch models. Each model is small (under 1 GB) and only a fraction are used at any time. To minimize cost and management overhead, which SageMaker inference option should be used?
Medium44A company needs to deploy a model that processes large payloads (up to 1 GB) asynchronously. The results should be written to S3, and the team needs SNS notifications upon completion. Which SageMaker inference option is MOST suitable?
Easy45A company deploys a large NLP model on a SageMaker real-time endpoint using an ml.p3.2xlarge instance. To reduce inference cost without sacrificing throughput, they want to compile the model for their target hardware. Which service should they use?
Hard46A company wants to deploy a new model using a canary deployment strategy on SageMaker. Which two actions should they take? (Select TWO.)
Medium47A team has a SageMaker Pipeline that trains a model and registers it in the Model Registry. They want to automate the deployment of the approved model to a staging environment. Which event-driven approach should they use?
Medium48A startup wants to deploy a containerized ML application that includes both a model inference server and a preprocessing component in the same endpoint. Which SageMaker endpoint type supports running multiple containers?
Medium49A company wants to update an existing SageMaker real-time endpoint to serve a new model version. They need to route a small percentage of traffic to the new version initially and monitor for errors before switching fully. Which deployment pattern supports this?
Easy50An organization wants to ensure that only approved model versions can be deployed to production. They use the SageMaker Model Registry to track model versions. How can they enforce that only approved models are deployed?
Medium51A company uses SageMaker Neo to compile a trained model for deployment on edge devices. What is the primary benefit of using Neo?
Easy52A data science team needs to deploy a trained PyTorch model for real-time inference with sub-100ms latency. The model fits on a single GPU. Which SageMaker inference option is MOST cost-effective while meeting the latency requirement?
Easy53A company wants to deploy a scikit-learn model to a SageMaker AI real-time endpoint. The model must be loaded from a custom Python module that contains preprocessing logic not present in the built-in scikit-learn container. The team wants to minimize operational overhead and does not need to change system-level libraries. Which approach should the engineer take?
Medium54An ML engineer has a real-time SageMaker endpoint serving a fraud-detection model. The team wants to release a new model version to a small percentage of live traffic first, monitor CloudWatch metrics for accuracy regressions, and roll back quickly if performance degrades. They also want the production and candidate variants to share the same endpoint so latency comparisons are apples-to-apples. Which SageMaker deployment strategy should they use?
Medium55A machine learning team has a model that needs to serve predictions with very low latency (under 10 ms) for a real-time web application. The model is a small ensemble of three neural networks that fits in memory. Which SageMaker inference option is MOST appropriate?
Medium56A team needs to deploy a new model version to production while minimizing risk. They want to route 5% of live traffic to the new model and 95% to the current model, and then gradually increase the new model's traffic. Which SageMaker deployment pattern should they use?
Medium57An ML platform team must orchestrate a workflow that trains a model, evaluates it against a baseline, and only registers the model if evaluation passes. If evaluation fails, the workflow must notify the data science channel and stop without registering. The team wants the orchestration logic to be expressed as code, versioned in git, and integrated with SageMaker training jobs and Model Registry. Which approach best fits these requirements?
Hard58An ML engineer needs to compile a trained TensorFlow model to run efficiently on a target edge device with an ARM CPU. Which AWS service should they use?
Easy59A company runs a real-time fraud detection model on a SageMaker endpoint. The model is updated weekly, and each update must be validated against live traffic without affecting existing predictions. The team wants to compare the new model's performance against the current model using a small percentage of incoming requests, while ensuring that the current model continues to serve the majority of traffic. Which SageMaker deployment strategy should they use?
Medium60A hospital's ML team deploys a diagnostic model to a SageMaker real-time endpoint. Compliance requires that every inference request and response be recorded for later auditing, and the records must be retrievable months later. The team needs a low-effort way to capture this data. Which approach should they use?
Easy61A data science team uses SageMaker Pipelines to orchestrate their ML workflow. They noticed that even when source data hasn't changed, the pipeline re-runs all steps, wasting compute time. What should they enable to avoid redundant runs?
Hard62A team uses SageMaker Pipelines to automate retraining. They want to skip the training step if the data has not changed since the last run. Which feature should they enable?
Medium63A machine learning engineer deploys a new model version to a SageMaker endpoint with production variants. They want to gradually shift traffic from the old model to the new model, monitoring for errors, and automatically roll back if the error rate exceeds 5%. Which deployment pattern should they use?
Hard64A company has 200 small PyTorch models that are each used infrequently but need to be available for real-time inference. To minimize costs, they want to host all models on a single endpoint. Which SageMaker feature should they use?
Medium65An organization wants to automate ML retraining using an event-driven architecture. Which THREE services should they combine? (Select THREE.)
Medium66A company wants to deploy a trained XGBoost model for batch inference on a large dataset stored in S3. The inference job should be cost-effective and does not require real-time responses. Which SageMaker inference option should they use?
Easy67A machine learning engineer needs to deploy a new version of a model gradually, initially sending 5% of traffic to the new version and 95% to the current version, while monitoring for errors. Which deployment pattern should they use?
Easy68A team uses SageMaker real-time endpoints for inference. They want to deploy a new model version and compare its performance with the current version under live traffic without affecting user experience. Which method should they use?
Hard69A company is deploying a large language model (LLM) to a SageMaker endpoint. They want to minimize inference latency and cost by using GPU acceleration and model parallelism. The model is too large to fit on a single GPU. Which SageMaker feature should they use?
Hard70A machine learning team runs a SageMaker AI Pipeline that trains a model and registers it in the SageMaker AI Model Registry. A separate deployment process must promote the model to production only after a human reviewer approves the model version. The team wants to automate the promotion so that approval in the Model Registry triggers the deployment without manual intervention. Which combination of steps should the engineer implement?
Hard71A machine learning engineer is deploying a model to a SageMaker endpoint that must handle occasional large payloads up to 1 GB. The inference time can take up to 10 minutes. The team wants to minimize cost and avoid idle compute. Which deployment option is most appropriate?
Medium72A company wants to deploy a PyTorch model on SageMaker using the NVIDIA Triton Inference Server for GPU acceleration. They have an existing Triton configuration. Which approach should they take?
Medium73A company needs to update a model in production without any downtime. They currently have a single real-time endpoint serving traffic. Which approach allows them to deploy a new model version and switch traffic gradually while being able to roll back quickly?
Hard74An ML engineer has trained a model and stored the model artifacts in an Amazon S3 bucket in the same AWS Region as the planned SageMaker AI endpoint. During endpoint creation, the engineer must specify the S3 location of the model artifacts. Which permission must the SageMaker AI execution role have for the endpoint to load the model successfully?
Easy75A data science team is using AWS Step Functions to orchestrate a machine learning workflow that includes a SageMaker training job followed by a model deployment. They want to ensure that if the training job fails, the workflow retries up to three times with exponential backoff before sending a notification to an Amazon SNS topic. Which Step Functions feature should they use to implement this?
Medium76A company uses SageMaker Pipelines to automate their ML workflow. They notice that the pipeline reruns all steps even when the input data has not changed. Which feature should they enable to avoid unnecessary recomputation?
Easy77A company uses SageMaker to deploy a model and wants to perform A/B testing by splitting traffic between two model variants. Which TWO actions should they take? (Select TWO.)
Medium78A team uses MLflow on SageMaker for experiment tracking. They want to automate the retraining of a model when new training data arrives in an S3 bucket. Which combination of services should they use?
Medium79A machine learning engineer needs to deploy a TensorFlow model that requires a custom inference environment with specific system libraries. The model will be used in a real-time application with variable traffic. They want to minimize cold start latency. Which SageMaker hosting option should they choose?
Medium80A company wants to deploy a model using a serverless inference endpoint that can automatically scale to zero when not in use and has a configurable maximum concurrency. Which SageMaker inference option meets these requirements?
Easy81A machine learning engineer wants to automatically trigger a retraining pipeline whenever new training data arrives in an S3 bucket. The pipeline uses SageMaker Pipelines. Which AWS service should be used to detect the S3 event and start the pipeline?
Easy82A team built a SageMaker Pipeline that includes a training step and a model evaluation step. They want to automatically register a model in SageMaker Model Registry only if the evaluation metric (accuracy) exceeds 0.9. Which pipeline step should be used to implement this conditional logic?
Medium83An ML engineer is designing a SageMaker Pipeline for model training and registration. They need to ensure that the pipeline can be re-run with different datasets without manual intervention, and that the steps are only re-executed if inputs have changed. Which THREE features should they configure? (Select THREE.)
Hard84A team is migrating their ML infrastructure to AWS and wants to use infrastructure as code to manage SageMaker Studio domains, user profiles, and associated resources. Which services can they use for this purpose? (Select THREE.)
Medium85A team needs to deploy a PyTorch model that uses custom CUDA kernels. They want to use NVIDIA Triton Inference Server on SageMaker for high-performance serving. Which SageMaker configuration is required to use Triton?
Hard86A machine learning engineer is using SageMaker Pipelines to orchestrate a training workflow. The pipeline includes a processing step that outputs a dataset, which is then used by a training step. The engineer notices that the processing step runs every time the pipeline executes, even when the input data has not changed. The engineer wants to avoid re-running the processing step if the input data and code are unchanged, while ensuring that downstream steps still execute if the processing step is skipped. Which approach should the engineer take?
Hard87A team has 200 small ML models that need to be served via HTTPS endpoints. Each model is used infrequently, and the team wants to minimize hosting costs. Which SageMaker deployment approach is MOST cost-effective?
Medium88A machine learning team uses SageMaker Pipelines to automate retraining. They want to avoid re-running data processing steps if the data has not changed since the last successful pipeline run. Which built-in feature should they enable?
Medium89A company is using AWS Step Functions to orchestrate their ML retraining pipeline. They want to trigger retraining when new data arrives, but only if the model's performance has degraded below a threshold. Which THREE AWS services should they use together to achieve this? (Choose three.)
Medium90A company uses SageMaker Pipelines to automate their ML workflow. They need to add model versioning and approval workflow. Which THREE steps should they include in their pipeline to achieve this? (Choose THREE.)
Medium91A company wants to trigger a model retraining pipeline whenever new training data arrives in an S3 bucket. They also need to send a notification to a Slack channel when the retraining completes. Which TWO AWS services should they use to implement this event-driven workflow? (Select TWO.)
Easy92A company has 50 small PyTorch models that are used infrequently for inference. They want to minimize costs while maintaining the ability to serve all models from a single endpoint. Which SageMaker feature should they use?
Easy93A machine learning team needs to deploy a PyTorch model that has been compiled with SageMaker Neo to improve inference performance on edge devices. Which TWO statements about SageMaker Neo are correct? (Select TWO.)
Medium94A financial services company needs to deploy a machine learning model for real-time fraud detection. The model must be highly available across multiple Availability Zones and must support automatic scaling based on request volume. The company also needs to perform canary deployments to test new model versions with a small percentage of traffic before full rollout. Which SageMaker feature should they use?
HardOther domains
All MLA-C01 exam domains
Frequently asked questions
- What does the Deployment and Orchestration of ML Workflows domain cover on the MLA-C01 exam?
- Be able to wire a SageMaker Pipeline that trains, evaluates, and conditionally registers a model, then deploy it via a real-time endpoint using the right traffic-shifting and auto scaling configuration. The critical skill is matching each requirement to the correct SageMaker feature and parameter.
- How many questions are in this domain?
- This page lists all 94 Deployment and Orchestration of ML Workflows questions in the MLA-C01 question bank. The actual exam draws from this domain proportionally to its weighting in the official exam blueprint.
- What is the best way to practise this domain?
- Start with a short focused session (10 questions) to identify gaps, then work through explanations. Repeat with a longer session once the weak areas feel solid.
- Can I practise only Deployment and Orchestration of ML Workflows questions?
- Yes — the session launcher on this page filters questions to this domain only. Choose any session length for inline explanations and scoring.