Courseiva

MLA-C01 · domain

ML Solution Monitoring, Maintenance, and Security

This domain covers operating ML systems after deployment on AWS: detecting drift and data quality issues, monitoring model performance, maintaining lineage and versioning, and securing data and endpoints. Questions present operational scenarios and ask you to choose the correct SageMaker or AWS service, configuration setting, or remediation step.

94 questions26 easy47 medium21 hard

Focused practice

Practice ML Solution Monitoring, Maintenance, and Security questions

Scored sessions drawing only from this domain — pick a length below.

Start 20-question practice test →

What this domain covers

What to know about ML Solution Monitoring, Maintenance, and Security

You must be able to select the right SageMaker monitoring, lineage, and security configuration for a deployed model scenario. The most important thing is knowing which service detects drift versus which one records lineage, and enabling the correct encryption setting for training traffic.

SageMaker Model Monitor for data drift, model quality, bias, and feature attribution monitoring

SageMaker Lineage Tracking and Model Registry for dataset, hyperparameter, and model version lineage

SageMaker training job network isolation and inter-container traffic encryption settings

SageMaker JumpStart deployment options and cost tradeoffs such as serverless inference versus real-time endpoints

Watch out for

Common ML Solution Monitoring, Maintenance, and Security exam traps

  • ▸Assuming Model Monitor automatically retrains or fixes drift; it only detects and reports, so remediation must be configured separately.
  • ▸Confusing Model Registry versioning with Lineage Tracking; lineage records artifact relationships, while the registry manages approval and deployment status.
  • ▸Forgetting that inter-container encryption requires enabling the specific training job encryption setting, not just an encryption key on the volume.

Question index

All ML Solution Monitoring, Maintenance, and Security questions (94)

Click any question to see the full explanation, or start a practice session above.

1

A company wants to implement a retraining pipeline that automatically triggers when SageMaker Model Monitor detects data drift. The retraining job should use the latest approved pipeline version in SageMaker Pipelines. Which approach meets these requirements?

Medium
2

A data scientist notices that a SageMaker endpoint is returning HTTP 5XX errors under high load. The endpoint uses a single ml.m5.large instance. The team wants to reduce these errors without changing the instance type. What is the most cost-effective step?

Easy
3

A company uses SageMaker Model Monitor for data quality. They notice that monitoring jobs are failing intermittently with constraint violations. Upon review, they see that some features have different data types in production compared to the baseline (e.g., string instead of integer). Which type of drift is this?

Hard
4

A data scientist wants to track the lineage of models, datasets, and training jobs in SageMaker. Which SageMaker feature should they use to capture these relationships as artifacts and actions?

Easy
5

A data science team wants to automate the retraining of a model when data drift is detected. Which TWO AWS services should they use in combination to achieve this? (Choose TWO)

Easy
6

A company wants to automate remediation when a SageMaker endpoint's latency exceeds a threshold for more than 5 minutes. The team needs to be notified and a Lambda function should be invoked to scale up the endpoint. Which combination of services should be used?

Easy
7

A machine learning engineer wants to set up a retraining pipeline that triggers when model quality degrades. Which TWO components are essential for this automated retraining pipeline? (Select TWO)

Easy
8

A machine learning engineer wants to monitor a deployed model for data drift. Which SageMaker feature should they use to automatically detect drift in the input data distribution compared to the training data baseline?

Easy
9

A company is deploying a foundation model using SageMaker JumpStart. They want to minimize inference costs while maintaining low latency. Which TWO strategies should they consider? (Select TWO)

Hard
10

A fraud-detection team runs a SageMaker real-time endpoint. Compliance requires that every inference request be logged with its full request and response payloads, and that a security engineer be able to prove later which requests were captured. The team enables SageMaker Model Monitor data capture with a capture percentage of 100. Where are the captured records stored, and what must be configured so the records are encrypted with a customer-managed key rather than an AWS-managed key?

Medium
11

A healthcare company has deployed a SageMaker model that predicts patient risk scores. The security team requires that all access to the model's endpoint be authenticated and authorized, and that every invocation be traceable to a specific user or application for audit purposes. The team also wants to enforce least privilege so that only specific applications can invoke the endpoint. Which TWO actions should the machine learning engineer take to meet these requirements? (Choose two.)

Hard
12

A company has a SageMaker endpoint that serves a recommendation model. The security team wants to ensure that the model artifacts stored in Amazon S3 are encrypted at rest and that access to the S3 bucket is limited to the SageMaker execution role only. The team also wants to receive alerts if the bucket policy is changed. Which combination of actions should the machine learning engineer take?

Easy
13

A fraud detection team runs a SageMaker real-time endpoint that logs every request and response to an Amazon S3 bucket. Compliance requires that the model's prediction inputs and outputs be encrypted at rest with a customer-managed AWS KMS key, and that the endpoint be able to read the capture bucket only when necessary. The team enables data capture and specifies a KMS key on the endpoint configuration. Which additional configuration is required for the captured data written to S3 to be encrypted with that customer-managed key?

Medium
14

A machine learning team needs to ensure that all model training and inference jobs within SageMaker Studio run in a private network without internet access. The team also requires that inter-container traffic within the same training job be encrypted. Which configurations should they combine?

Hard
15

A data science team uses SageMaker to train and deploy models. They need to track model lineage, including datasets, training jobs, and model versions, to ensure reproducibility. Which THREE actions should they take? (Select THREE)

Medium
16

A data science team wants to track the lineage of models, including datasets, training jobs, and endpoints, for reproducibility and audit. They need a solution that captures relationships between artifacts automatically during training and deployment. Which service should they use?

Medium
17

An ML team has deployed a model to a SageMaker real-time endpoint and wants to set up automated monitoring for model quality. Which TWO elements are required to configure SageMaker Model Monitor for model quality? (Select TWO.)

Medium
18

An ML engineer monitors a SageMaker endpoint for data drift. They set up SageMaker Model Monitor to compare inference data against a baseline created from the training dataset. The monitoring schedule runs daily and reports violations. Which monitoring type should be configured to detect if the distribution of a numerical feature in real-time inference data differs significantly from the training distribution?

Medium
19

A company uses SageMaker Model Monitor's feature attribution drift monitoring with SHAP. They receive an alert that the average SHAP value for a particular feature has increased significantly compared to the baseline. The feature's input distribution has not changed. What does this likely indicate?

Hard
20

A financial institution uses SageMaker to train and deploy models. They need to track every experiment, model version, and deployment step for audit purposes. Which SageMaker feature should they use to capture the full lineage of artifacts, actions, and contexts?

Medium
21

A machine learning engineer must grant a data scientist the least-privilege permissions needed to invoke one specific SageMaker real-time endpoint from their own AWS account, and to view that endpoint's CloudWatch metrics without being able to modify the endpoint. The endpoint ARN is known. Which TWO IAM policy statements should the engineer include? (Choose two.)

Medium
22

A healthcare analytics team stores model artifacts and training datasets in Amazon S3 and uses SageMaker. An internal audit finds that some S3 buckets containing protected health information are missing encryption and that access is granted broadly. The team must remediate quickly and prevent future misconfiguration. Which combination of actions should the ML engineer take FIRST?

Medium
23

A retail company uses a SageMaker Model Monitor data quality monitor on a real-time endpoint. The monitor's baseline was generated from a training dataset in which the "promo_code" feature was often null. In production the feature is now populated for nearly every record, and the monitor reports violations even though model accuracy has not degraded. The team wants the monitor to stop flagging this expected change without disabling monitoring entirely. What should they do?

Hard
24

A financial services company must ensure that a SageMaker model deployed to a real-time endpoint only produces predictions consistent with a fairness constraint on a protected attribute, and that any violation is detected within minutes and triggers an alert to the compliance team. The model is already deployed and monitored for data quality. Which approach should the machine learning engineer implement?

Hard
25

A company wants to enable cross-account access to a SageMaker model endpoint. The model is in Account A, and Account B needs to invoke it. Which TWO steps are required? (Select TWO)

Hard
26

A retail company stores training datasets, model artifacts, and feature data in Amazon S3. An auditor requires that all objects be encrypted at rest with keys the company controls and that key usage be independently auditable. The team wants minimal operational overhead. Which approach should the ML engineer recommend?

Easy
27

A team receives alerts that their SageMaker endpoint latency has increased significantly. They check CloudWatch metrics and see Invocations rising, but ModelLatency remains stable. Which metric should they investigate to find the source of the increased latency?

Medium
28

A team uses SageMaker Clarify to monitor bias drift on a deployed model. They have defined a baseline with training data and set up a monitoring schedule. After one month, they receive a violation report indicating that the post-training metrics have deviated from the baseline. What does this violation indicate?

Medium
29

An organization uses SageMaker Studio and needs to restrict Studio's internet access while allowing users to install custom packages from a private PyPI mirror hosted in a VPC. Which networking configuration should they use?

Hard
30

An organization needs to ensure that all data used for inference on a SageMaker endpoint is encrypted at rest. The endpoint uses a SageMaker-provided container. Which configuration should be applied?

Easy
31

A company has a SageMaker real-time endpoint that serves predictions. They want to set up automated monitoring and remediation for when the number of 5XX errors exceeds a threshold. Which TWO steps should they take? (Choose TWO.)

Medium
32

A data scientist notices that a production model's accuracy has degraded over the past week. The training data distribution remains unchanged, but the relationship between features and the target has shifted. Which type of drift is occurring, and which monitoring approach should be used?

Medium
33

A machine learning team deploys a fraud detection model on a SageMaker endpoint. The model's predictions are used in real-time. The team wants to monitor for data drift by comparing incoming data distributions against a baseline created from the training data. Which SageMaker capability should they use?

Medium
34

A company plans to deploy a large foundation model using SageMaker JumpStart. They are concerned about costs because the model will be used intermittently. Which deployment option is MOST cost-effective for intermittent traffic?

Medium
35

A company wants to deploy a foundation model from SageMaker JumpStart with the lowest possible inference cost, given that latency requirements are flexible. They have a mix of traffic volumes. Which approach should they take?

Medium
36

A machine learning engineer is configuring a SageMaker Processing job that runs a custom container to compute bias metrics on a dataset containing personally identifiable information. The job reads input data from one S3 bucket and writes reports to another, and the security team requires that the container has no outbound internet access and that the input and output buckets are reached without traversing the public internet. Which configuration satisfies these requirements?

Hard
37

A company deploys a model for fraud detection. They want to monitor if the model's predictions become less accurate over time due to changes in the underlying data distribution, but they do not have immediate access to ground truth labels. Which type of drift should they monitor as a proxy?

Medium
38

A retail company has a SageMaker model that predicts customer churn. The model was trained on data that included a 'customer_zipcode' feature. After deployment, the data science team notices that the model's predictions for certain zip codes have become less accurate over time. They suspect that the relationship between zip code and churn has changed due to a recent relocation of a major employer. Which SageMaker monitoring capability should they use to detect this type of drift?

Medium
39

A team wants to secure SageMaker endpoints for a healthcare application. They must ensure data is encrypted at rest and in transit, and that the endpoint can only be accessed from within a VPC. Which THREE steps should they take? (Select THREE)

Medium
40

A team deploys a SageMaker real-time endpoint and configures it with an auto scaling policy targeting a variant. During a flash sale, traffic spikes and the team notices that the number of instances increases, but the average model latency still climbs above the target. The team wants the scaling behavior to react faster to sudden bursts without over-provisioning during steady periods. Which change should they make to the scaling policy?

Easy
41

A company deploys a real-time inference endpoint with auto-scaling using a target tracking policy based on average Invocations per instance. They notice that during a traffic spike, the endpoint scales out too late, causing increased latency. They want to scale proactively before the spike. Which strategy should they implement?

Hard
42

A machine learning team notices an increase in 5XXError count for a SageMaker endpoint. They want to set up automated remediation. Which THREE actions should they take? (Select THREE)

Medium
43

A machine learning engineer is monitoring a deployed model for data drift. The input features are a mix of categorical and numerical columns. The baseline is from the training data. Which SageMaker Model Monitor feature should they enable to detect changes in the distribution of each feature over time?

Medium
44

A machine learning team deploys a model for loan approval. They want to monitor data drift on the real-time endpoint using SageMaker Model Monitor. Which set of actions should they take to set up data quality monitoring?

Medium
45

A fraud-detection team runs a SageMaker real-time endpoint in a production account. Their security team requires that the endpoint be reachable only from within a specific Amazon VPC and that access to invoke the endpoint be governed by identity-based policies with least privilege. Which TWO configurations should the ML engineer implement to meet these requirements? (Choose two.)

Hard
46

An ML engineer needs to monitor the operational health of a SageMaker endpoint, specifically the time taken for the container to process an inference request and the overhead added by SageMaker. Which two CloudWatch metrics should they examine?

Easy
47

A machine learning team trains a model in SageMaker and wants to track every step — from dataset version to hyperparameters to final model artifact — for reproducibility and audit compliance. Which SageMaker feature should they use?

Medium
48

A team monitors a production endpoint and notices a sudden increase in 5XXError count. Which of the following is the most likely cause?

Medium
49

A company has deployed a machine learning model on Amazon SageMaker and wants to automatically detect when the distribution of input features deviates significantly from the training data distribution. Which SageMaker feature should they use?

Easy
50

A company deploys a model for fraud detection. They need to monitor for bias after deployment, specifically whether the model's false positive rate changes across demographic groups over time. Which SageMaker feature should they use?

Hard
51

A company wants to reduce costs for a real-time inference endpoint that experiences predictable traffic spikes during business hours and low traffic at night. Which auto-scaling policy is MOST cost-effective while maintaining performance?

Easy
52

A healthcare analytics team trains models in SageMaker and stores artifacts in an S3 bucket that contains protected health information. An auditor asks how the team can prove which training dataset and container image produced the model currently deployed to production, and wants the evidence retained even if someone deletes the training job. Which SageMaker capability should the team rely on to capture and retain this metadata automatically?

Medium
53

A team uses SageMaker ML Lineage Tracking to capture the metadata of their ML workflow. They want to query the lineage to see which model version was trained from a specific dataset. Which Lineage Tracking entity represents the dataset?

Medium
54

A hospital deploys a model to predict patient readmission risk. To comply with regulations, they must ensure that the model's predictions do not show bias against any demographic group over time. Which service should they use for ongoing monitoring?

Hard
55

A machine learning engineer wants to deploy a pre-trained foundation model for text summarization using SageMaker JumpStart. Which of the following is a primary cost consideration when deploying such a model?

Easy
56

A team wants to use SageMaker Clarify to monitor bias in their production model predictions. They have configured a bias drift monitor. What does SageMaker Clarify compare to detect bias drift?

Medium
57

A healthcare analytics team trains models in Amazon SageMaker and needs an immutable, queryable record of which dataset version and training job produced each registered model version, so an auditor can trace a deployed model back to its inputs months later. Which SageMaker capability should they rely on?

Easy
58

A company wants to track the lineage of their ML models for reproducibility and auditability. Which THREE services or features should they use together to achieve this? (Choose THREE.)

Medium
59

A company wants to reduce costs for a SageMaker real-time endpoint that has variable traffic. Which feature allows the endpoint to automatically adjust instance count based on demand?

Easy
60

A financial services company has a SageMaker real-time endpoint serving a fraud detection model. Compliance requires that all inference requests and responses be logged with the ability to detect anomalous input feature distributions over time. The team wants a managed solution that captures request/response payloads to Amazon S3 and automatically computes statistics and constraints against a baseline. Which combination of SageMaker features should they enable?

Medium
61

A machine learning engineer manages a SageMaker Model Monitor schedule for a real-time endpoint. The monitor's baseline was computed from a training dataset with a categorical feature named region. In production, a new category value appears that was never seen in training, and the monitor begins reporting violations. The engineer wants the monitor to flag only the appearance of unknown categories without treating normal distribution shifts in known categories as violations. Which approach should the engineer take?

Hard
62

A company uses SageMaker JumpStart to deploy a foundation model for a summarization task. They want to minimize costs while still meeting a latency requirement of under 2 seconds. Which option should they consider?

Medium
63

An ML team uses SageMaker to deploy a model for real-time inference. They want to monitor and improve cost efficiency. Which THREE actions should they take? (Select THREE.)

Hard
64

A retail company runs a SageMaker real-time endpoint serving a demand forecasting model. Security policy requires that all inference requests travel over the AWS private network and never traverse the public internet, and that the endpoint cannot be invoked from outside the company VPC. The endpoint already uses a customer-managed KMS key for volume encryption. Which TWO configurations should the engineer apply to meet these requirements? (Choose two.)

Hard
65

A machine learning engineer observes that model performance on a SageMaker endpoint has degraded over the past week. Ground truth labels are available with a 2-day delay. The engineer wants to automatically trigger a retraining pipeline when prediction quality drops below an acceptable threshold. Which approach is most appropriate?

Medium
66

A team wants to monitor the number of requests and latency of their SageMaker endpoint using a unified dashboard. Which AWS service should they use to create a custom dashboard with these metrics?

Easy
67

A company wants to automatically trigger a retraining pipeline when concept drift is detected in their deployed model. Which combination of services should they use?

Medium
68

A machine learning team needs to monitor a deployed model for both data drift and concept drift. Which TWO approaches should they implement? (Select TWO.)

Medium
69

A data scientist deploys a model and wants to monitor the endpoint's invocation latency. They notice that the CloudWatch metric 'ModelLatency' is high, but 'OverheadLatency' is low. Which statement correctly interprets these metrics?

Medium
70

A small team runs a SageMaker real-time endpoint in production. They want a low-effort way to know when the endpoint's invocations are failing so they can react quickly, and they want the alert delivered to their on-call channel. Which approach requires the least custom code?

Easy
71

A machine learning engineer notices that the latency of a SageMaker endpoint has increased over time. They need to identify which component (model inference vs. pre/post-processing) contributes most to the latency. Which CloudWatch metrics should they examine?

Medium
72

A company uses SageMaker endpoints for real-time inference. They want to automatically scale the number of instances based on the number of outstanding requests. Which auto-scaling policy type should they choose?

Medium
73

A machine learning engineer needs to give a data scientist read-only access to the model artifacts, training metrics, and monitoring reports stored in a single Amazon S3 bucket used by a SageMaker project, while ensuring the data scientist cannot delete or overwrite any object. Which approach follows least-privilege practice?

Easy
74

A company uses SageMaker Inference Recommender to select the optimal endpoint configuration. After running the recommender, they receive a recommendation for a specific instance type and initial instance count. What should they do next to optimize costs over time?

Medium
75

A company uses SageMaker Studio for collaborative ML development. The security team requires that all SageMaker Studio notebooks run within a VPC and cannot access the public internet. Which configuration should the administrator set?

Easy
76

An organization wants to schedule a retraining pipeline to run every Sunday night. Which AWS service should they use to trigger the pipeline on a schedule?

Easy
77

A team is monitoring a SageMaker endpoint and notices that the average latency (ModelLatency) is increasing over time, but the number of invocations is steady. They suspect that the model's inference code is becoming slower due to memory leaks. Which metric should they also examine to confirm this hypothesis?

Medium
78

An ML engineer wants to be notified when the average inference latency of a SageMaker endpoint exceeds 500 ms for 2 consecutive evaluation periods. Which AWS service combination should they use?

Easy
79

A company uses SageMaker Model Monitor for feature attribution drift monitoring with SHAP. Which THREE prerequisites must be in place before starting the monitoring schedule? (Select THREE)

Medium
80

A data science team uses Amazon SageMaker Model Monitor to detect data drift in production. They notice that the schema of incoming data (number of features) has changed compared to the training baseline. Which type of monitor is BEST suited to detect this issue?

Medium
81

A company is deploying a SageMaker real-time endpoint and needs to monitor inference latency. Which THREE metrics are available from SageMaker for this purpose? (Choose THREE.)

Medium
82

A machine learning engineer is deploying a model to a SageMaker real-time endpoint that must be accessible only from within a specific Amazon VPC and must not have a public IP address. The engineer also needs to ensure that all data in transit between the endpoint and the calling application is encrypted. Which configuration should the engineer use?

Medium
83

A machine learning engineer needs to monitor a SageMaker endpoint for data drift and receive alerts when drift is detected. They want to use a fully managed AWS service to schedule the monitoring jobs and send notifications. Which AWS service should they use to orchestrate the monitoring schedule and trigger alerts?

Easy
84

A company uses SageMaker Model Monitor to detect bias drift in their real-time inference endpoint. They have collected ground truth labels and want to monitor for bias across different demographic groups. Which type of monitoring should they configure?

Medium
85

A financial services company is deploying a fraud detection model on SageMaker. To comply with regulations, they must ensure that the model's predictions are not biased against protected groups. They plan to monitor bias drift post-deployment using SageMaker Clarify. Which data inputs are required to configure Clarify's bias drift monitoring?

Hard
86

A company wants to reduce costs for a SageMaker real-time endpoint that receives predictable traffic patterns: high during business hours and low at night. The model is a small PyTorch model. Which cost-saving strategy is most suitable?

Easy
87

A media company stores training data in an S3 bucket encrypted with an AWS KMS customer-managed key. A SageMaker training job runs inside a private VPC subnet with no internet access and must read that bucket. The job currently fails with an access-denied error from S3. Which change most directly resolves the failure while preserving the private-network requirement?

Hard
88

A team has deployed a real-time inference endpoint and wants to automatically scale based on CPU utilization. Which scaling policy type should they use with Application Auto Scaling for SageMaker endpoints?

Medium
89

A machine learning team needs to automatically retrain a model when concept drift is detected in the deployed endpoint's predictions. Which TWO steps should they take? (Choose TWO.)

Medium
90

A company wants to reduce costs for a production SageMaker endpoint that has predictable traffic patterns. They have purchased a Savings Plan. What additional step can they take to further optimize costs while maintaining performance?

Easy
91

A data scientist uses SageMaker Model Monitor to track feature attribution drift. Which technique does SageMaker Model Monitor use to compute feature attributions?

Medium
92

A company wants to track the lineage of their ML models, including the training dataset, hyperparameters, and training job used to produce each model version. Which AWS service should they use?

Easy
93

A company deploys a model with SageMaker and wants to monitor for concept drift. They have noticed that the relationship between input features and the target variable has changed, causing model accuracy to degrade. However, the input data distribution remains stable. Which type of drift is this, and what is the most appropriate response strategy?

Hard
94

An organization needs to ensure that all data transmitted between containers in a SageMaker training job is encrypted. In the training job configuration, which setting should they enable?

Hard

Frequently asked questions

What does the ML Solution Monitoring, Maintenance, and Security domain cover on the MLA-C01 exam?
You must be able to select the right SageMaker monitoring, lineage, and security configuration for a deployed model scenario. The most important thing is knowing which service detects drift versus which one records lineage, and enabling the correct encryption setting for training traffic.
How many questions are in this domain?
This page lists all 94 ML Solution Monitoring, Maintenance, and Security questions in the MLA-C01 question bank. The actual exam draws from this domain proportionally to its weighting in the official exam blueprint.
What is the best way to practise this domain?
Start with a short focused session (10 questions) to identify gaps, then work through explanations. Repeat with a longer session once the weak areas feel solid.
Can I practise only ML Solution Monitoring, Maintenance, and Security questions?
Yes — the session launcher on this page filters questions to this domain only. Choose any session length for inline explanations and scoring.
aws-ml-engineer-associate AWS-ML-ENGINEER-ASSOCIATE mla monitoring security Practice Questions