Courseiva

MLA-C01 · domain

scenario questions

Practise AWS Certified Machine Learning Engineer Associate MLA-C01 scenario questions practice questions — original exam-style scenarios with answer choices, explanations, and analysis of common mistakes.

665 questions193 easy305 medium167 hard

Focused practice

Practice scenario questions questions

Scored sessions drawing only from this domain — pick a length below.

Start 20-question practice test →

What this domain covers

What to know about scenario questions

scenario questions questions test whether you can apply the concept in context, not just recognise a definition.

How the topic appears in realistic exam-style scenarios.

Which detail in the question changes the correct answer.

How to eliminate plausible but wrong options.

How to connect the question back to the wider exam objective.

Watch out for

Common scenario questions exam traps

  • ▸Answering from memory before reading the full scenario.
  • ▸Missing a constraint such as cost, availability, security, scope or command context.
  • ▸Choosing a broad answer when the question asks for the most specific fix.
  • ▸Ignoring why the wrong options are tempting.

Question index

All scenario questions questions (665)

Click any question to see the full explanation, or start a practice session above.

1

A machine learning engineer is using Amazon SageMaker Feature Store to manage features for a fraud detection model. The engineer needs to ensure that the feature group can serve both batch and real-time predictions. The feature group is configured with an online store enabled. Which additional configuration is required to support batch predictions?

Medium
2

Refer to the exhibit. A data scientist used a SageMaker training job with a custom Scikit-learn script. The training job failed with the error shown. What is the most likely cause of this failure?

Hard
3

A company is building a CI/CD pipeline for ML models using AWS CodePipeline and SageMaker. The pipeline should include steps to automatically retrain, evaluate, and deploy models. Which THREE components are essential for this pipeline? (Choose three.)

Hard
4

A team needs to deploy a model that has compliance requirements to log all inference requests and responses for auditing. The model will be served using a real-time endpoint. How can they achieve this without custom code?

Hard
5

A financial institution is training a fraud detection model using SageMaker. The dataset is highly imbalanced, with only 0.1% fraudulent transactions. The team wants to use SageMaker Automatic Model Tuning to find the best hyperparameters. They notice that the tuning job spends most of its time on configurations that predict all transactions as non-fraudulent. Which hyperparameter should they tune to directly address this issue?

Hard
6

A machine learning engineer is using Amazon SageMaker Data Wrangler to prepare a dataset with a categorical feature that has over 5,000 distinct values (high cardinality). The engineer needs to transform this feature into a form suitable for a gradient boosting model while preserving as much information as possible. Which transform should be applied?

Hard
7

An MLOps team is designing a SageMaker Pipeline to automate model retraining. The pipeline must: (1) run training only if new training data is available, (2) register the model in SageMaker Model Registry only if evaluation metrics exceed a threshold, (3) deploy the approved model to a staging endpoint automatically. Which THREE steps should they include? (Choose THREE.)

Hard
8

A company wants to run inference on a large dataset stored in S3 using a pre-trained model. The inference can tolerate latency from minutes to hours, and they want a fully managed solution that autoscales to handle large volumes. Which SageMaker inference option is most suitable?

Medium
9

A data scientist wants to compare the performance of two model versions (V1 and V2) in production by splitting traffic between them. They want to gradually increase the percentage of traffic to the new version while monitoring metrics. Which SageMaker feature enables this?

Easy
10

A team is training a PyTorch model using SageMaker. They have a custom training script that requires specific Python packages not included in the SageMaker default PyTorch container. Which approach should they use?

Medium
11

A data scientist needs to deploy a single ML model that will serve real-time predictions with low latency (under 10 ms) for a high-traffic web application. The model fits in memory and requires GPU acceleration. Which SageMaker inference option is MOST suitable?

Easy
12

A machine learning team is preparing numerical features for a linear regression model. Feature 'A' ranges from 0 to 1000, feature 'B' ranges from 0 to 1, and feature 'C' ranges from -10000 to 10000. The team wants to ensure that feature scales do not affect the model's coefficients and that the features are bounded between 0 and 1. Which transformation should they apply?

Hard
13

A company wants to implement a retraining pipeline that automatically triggers when SageMaker Model Monitor detects data drift. The retraining job should use the latest approved pipeline version in SageMaker Pipelines. Which approach meets these requirements?

Medium
14

A company uses AWS Glue ETL jobs to transform data for machine learning. They have a dataset with a column 'income' that is heavily right-skewed. Which transformation should be applied to make the distribution more Gaussian-like?

Hard
15

A data scientist is using SageMaker Autopilot to automatically build a binary classification model on a balanced dataset. They want to understand the relationship between the input features and the model predictions. Which feature in SageMaker Autopilot should they use?

Medium
16

A company has a model that receives low traffic but needs to handle sudden spikes. Which deployment option is most cost-effective?

Easy
17

A data engineer needs to ingest streaming clickstream data from a website into an S3 data lake for ML training, with the ability to run real-time aggregations before storage. Which combination of AWS services meets these requirements?

Medium
18

A machine learning team is training a large natural language processing model on Amazon SageMaker using the SageMaker Hugging Face container. The training job runs on multiple instances and uses Managed Spot Training to reduce costs. However, the job frequently gets interrupted by Spot interruptions, causing long training times. What should the team do to mitigate this issue?

Hard
19

A machine learning engineer has trained a model using SageMaker and wants to deploy it to a real-time endpoint. The engineer needs to specify the model artifacts, the inference code, and the environment. Which SageMaker resource should the engineer create first?

Easy
20

A machine learning team is developing a deep learning model for image classification. They observe that the training loss decreases rapidly but the validation loss starts increasing after a few epochs. Which strategy should they implement to address this issue?

Hard
21

A data scientist notices that a SageMaker endpoint is returning HTTP 5XX errors under high load. The endpoint uses a single ml.m5.large instance. The team wants to reduce these errors without changing the instance type. What is the most cost-effective step?

Easy
22

A company uses SageMaker Model Monitor for data quality. They notice that monitoring jobs are failing intermittently with constraint violations. Upon review, they see that some features have different data types in production compared to the baseline (e.g., string instead of integer). Which type of drift is this?

Hard
23

A team is training a PyTorch model using SageMaker with a custom training script. They want to track hyperparameters and metrics across multiple experiments. Which service should they use?

Medium
24

A machine learning engineer is using SageMaker Automatic Model Tuning (AMT) to optimize a model. They want to ensure the tuning job explores the hyperparameter search space efficiently and stops poorly performing trials early. Which two strategies should they use? (Choose two.)

Medium
25

A machine learning engineer is using SageMaker Pipelines to automate the training and deployment of a model. The pipeline includes a processing step for feature engineering, a training step, and a model registration step. The engineer wants to ensure that the pipeline is reproducible and that the model artifacts are versioned. Which two actions should be taken? (Choose two.)

Medium
26

A data scientist wants to track the lineage of models, datasets, and training jobs in SageMaker. Which SageMaker feature should they use to capture these relationships as artifacts and actions?

Easy
27

A data scientist needs to label a large dataset of product images for a classification model. They want to reduce labeling costs by prioritizing uncertain samples. Which Amazon SageMaker Ground Truth feature should they use?

Easy
28

A data science team wants to automate the retraining of a model when data drift is detected. Which TWO AWS services should they use in combination to achieve this? (Choose TWO)

Easy
29

A team wants to deploy a new model using a canary deployment strategy on SageMaker. Which TWO configurations are necessary? (Choose two.)

Medium
30

A company uses Amazon SageMaker Data Wrangler for data preparation. The data science team wants to automatically detect potential bias in their dataset before training a model. Which feature of Data Wrangler should they use?

Medium
31

A company has 10 TB of log data in compressed JSON format stored in Amazon S3. The data needs to be processed and transformed into a structured format for machine learning. The processing requires complex transformations, including parsing nested JSON and joining with a reference table. The company wants to minimize infrastructure management. Which approach should the company use?

Easy
32

A data scientist needs to split a time-series dataset into training and testing sets while avoiding data leakage from future values. Which splitting technique should the data scientist use?

Easy
33

A company is adopting Amazon SageMaker Pipelines to automate their ML workflow. They want to choose three key benefits that SageMaker Pipelines provides over traditional manual scripts and ad-hoc steps. Which THREE benefits are correct?

Easy
34

A company is using SageMaker Model Registry to manage model versions. They want to automatically deploy the latest approved model to production after retraining. Which approach is best?

Medium
35

A data scientist needs to ensure that the same train/test split is used across multiple experiments for reproducibility in SageMaker. Which approach should they take?

Medium
36

A company stores its model training data in Amazon S3. To meet compliance requirements, all data in transit between the S3 bucket and SageMaker must be encrypted. What should the company enforce?

Easy
37

A company is using Amazon SageMaker Data Wrangler to prepare a dataset with over 200 features. The dataset includes a categorical feature with more than 10,000 unique values (high cardinality). The ML engineer wants to transform this feature into a numeric representation suitable for a linear model without increasing dimensionality too much. Which built-in transform in Data Wrangler should the engineer use?

Hard
38

A data engineer is preparing a large training dataset stored in Amazon S3 as many small Parquet files, and a SageMaker training job that reads directly from S3 is spending most of its time on the input channel rather than on model computation. The engineer needs to improve the input throughput without changing the model code or the training algorithm. Which action should the engineer take?

Medium
39

A company uses SageMaker Autopilot to build a binary classification model. The generated leaderboard shows an ensemble model as the best candidate. The team needs a model that can be deployed for real-time inference with latency < 10ms. What should they do?

Hard
40

A machine learning engineer is training a tabular regression model using the SageMaker built-in XGBoost algorithm. They want to reduce overfitting and improve generalization without changing the algorithm. Which SageMaker hyperparameter should they tune to control the fraction of features randomly sampled per tree?

Medium
41

A company deploys a real-time inference endpoint on SageMaker for a customer-facing application. Traffic patterns are unpredictable and sometimes spike. The endpoint must scale automatically to handle load while minimizing cost. Which approach should the company take?

Medium
42

A machine learning team is building a fraud detection model. They have a dataset with a categorical feature 'merchant_id' that has over 10,000 unique values. Which feature engineering technique should they apply to 'merchant_id' to reduce dimensionality while retaining predictive power?

Medium
43

A machine learning engineer is deploying a model to a SageMaker endpoint for real-time inference. The model must return predictions within 100 milliseconds for 95% of requests. The engineer wants to monitor the endpoint's latency and automatically roll back if latency exceeds the threshold. Which combination of SageMaker features should be used?

Medium
44

A company wants to automate remediation when a SageMaker endpoint's latency exceeds a threshold for more than 5 minutes. The team needs to be notified and a Lambda function should be invoked to scale up the endpoint. Which combination of services should be used?

Easy
45

A company is using SageMaker to train a model with a custom container. The training script requires a specific version of a Python library that is not included in the default SageMaker containers. How should they provide this library?

Hard
46

A data scientist is preparing a large dataset (50 GB) for training a TensorFlow model on SageMaker. The dataset consists of many small CSV files. Training is slow due to I/O bottlenecks. Which data preparation strategy most effectively accelerates training?

Hard
47

A data scientist is using SageMaker Autopilot for a regression problem. They want to see which data preprocessing steps Autopilot applied. Which TWO sources can they use to find this information?

Medium
48

A data scientist wants to version and manage trained models, require approval before deployment, and enable cross-account deployment. Which SageMaker feature provides these capabilities?

Easy
49

A company uses SageMaker Pipelines to orchestrate their ML workflow. They notice that if a pipeline step fails due to a transient error (e.g., a brief network issue), the entire pipeline fails and they must manually rerun from the beginning. They want to automatically retry failed steps a few times before failing. What should they do?

Hard
50

A data scientist is using SageMaker to train a deep learning model with the PyTorch estimator. They want to log custom scalar metrics such as validation accuracy and loss during training so they can monitor the job in SageMaker. Which approach should they use to emit these metrics from the training script?

Medium
51

A machine learning engineer is using a SageMaker training job with a custom training script. They need to save the trained model artifacts to Amazon S3 so that the model can be deployed later. Which parameter in the SageMaker estimator should they configure to specify the S3 location for model artifacts?

Medium
52

A machine learning engineer is training a model using SageMaker's built-in XGBoost algorithm. The training job fails with an error indicating insufficient memory. Which parameter should be adjusted to reduce memory usage?

Easy
53

A team deploys a model with SageMaker and notices that the model returns inconsistent results during inference. They suspect a mismatch in feature transformation between the training pipeline and the inference pipeline. Which SageMaker feature can help compare the feature distributions?

Medium
54

A financial services company is training a fraud detection model using SageMaker. The dataset is highly imbalanced, with only 0.2% fraudulent transactions. The team wants to optimize the model for recall at a fixed precision of 90%. They are using the SageMaker built-in XGBoost algorithm with binary:logistic objective. Which evaluation metric should they monitor during training and hyperparameter tuning?

Hard
55

A data engineer needs to prepare a large dataset for machine learning. The data is stored in an Amazon RDS MySQL database and needs to be transformed and moved to an S3 bucket in Parquet format for use with SageMaker. Which AWS service is most suitable for this extraction, transformation, and loading (ETL) task?

Easy
56

A company is using SageMaker to serve a model for real-time predictions. They want to test a new model version by routing a small percentage of live traffic to it while the rest goes to the current model. They also need to compare performance metrics. Which TWO actions should they take? (Select TWO.)

Medium
57

A machine learning engineer wants to set up a retraining pipeline that triggers when model quality degrades. Which TWO components are essential for this automated retraining pipeline? (Select TWO)

Easy
58

A company uses Amazon SageMaker Pipelines for automated retraining. The pipeline includes a processing step that runs a Python script. The script uses the boto3 library to call an AWS service, but the calls are being throttled. What is the MOST effective way to address this within the pipeline?

Medium
59

A machine learning engineer wants to monitor a deployed model for data drift. Which SageMaker feature should they use to automatically detect drift in the input data distribution compared to the training data baseline?

Easy
60

A machine learning engineer has trained a scikit-learn model and saved it as model.joblib in Amazon S3. The engineer wants SageMaker to host the model for real-time inference without writing a custom container or inference script, because the model uses only standard predict behavior. Which deployment approach should the engineer use?

Medium
61

A team uses SageMaker Pipelines for CI/CD. The training step fails due to insufficient memory. How to fix without rewriting code?

Hard
62

A company is deploying a foundation model using SageMaker JumpStart. They want to minimize inference costs while maintaining low latency. Which TWO strategies should they consider? (Select TWO)

Hard
63

A company operates multiple AWS accounts with SageMaker workloads. They need to implement governance and security controls for model monitoring and maintenance. Which THREE actions should they take to meet compliance requirements?

Hard
64

A fraud-detection team runs a SageMaker real-time endpoint. Compliance requires that every inference request be logged with its full request and response payloads, and that a security engineer be able to prove later which requests were captured. The team enables SageMaker Model Monitor data capture with a capture percentage of 100. Where are the captured records stored, and what must be configured so the records are encrypted with a customer-managed key rather than an AWS-managed key?

Medium
65

A company uses an Amazon SageMaker endpoint for real-time inference. The security team requires that all traffic between the endpoint and the client application be encrypted in transit. Which configuration ensures this?

Easy
66

A company is deploying a machine learning model using SageMaker hosting. They need to support multiple versions of the model for A/B testing. Which TWO actions are required to set up the A/B test? (Choose two.)

Medium
67

A company is using Amazon SageMaker Ground Truth to create a labeled dataset for object detection in images. The team wants to minimize labeling costs while maintaining high accuracy. Which feature should they use to achieve this?

Easy
68

A data scientist is building a model to predict customer churn based on historical data. The dataset has 10 features and 100,000 records, and the target is binary. Which algorithm is most appropriate for this binary classification problem?

Easy
69

A data scientist wants to use SageMaker Clarify to analyze bias during training of a binary classification model. Which TWO types of bias metrics can SageMaker Clarify compute? (Select TWO.)

Medium
70

A healthcare company is deploying a model for predicting patient outcomes. The model must be deployed across multiple AWS accounts to meet compliance requirements. Each account has its own Amazon SageMaker endpoint. The company wants to centralize monitoring of model performance without exposing data across accounts. Which solution should the company use?

Hard
71

A machine learning engineer is setting up automated retraining for a model using SageMaker Pipelines. The pipeline should trigger when a data drift alert is received from Model Monitor. Which event source should the engineer use to initiate the pipeline?

Hard
72

An ML engineer needs to orchestrate a multi-step workflow that includes data preprocessing on Spark, model training on SageMaker, and deployment to a production endpoint. They require tight integration with other AWS services and the ability to add custom logic. Which AWS service should they use alongside SageMaker?

Medium
73

A company is training a deep learning model for object detection using SageMaker. The training is very slow and the GPU memory is insufficient for the batch size. The team wants to scale across multiple GPUs efficiently. Which THREE actions should they take? (Choose THREE.)

Hard
74

A healthcare company has deployed a SageMaker model that predicts patient risk scores. The security team requires that all access to the model's endpoint be authenticated and authorized, and that every invocation be traceable to a specific user or application for audit purposes. The team also wants to enforce least privilege so that only specific applications can invoke the endpoint. Which TWO actions should the machine learning engineer take to meet these requirements? (Choose two.)

Hard
75

A team runs a SageMaker Pipeline that trains a model and registers it in the Model Registry. Compliance requires that the pipeline run automatically every time new labeled data lands in S3, and that each run record the exact S3 data prefix, the training image URI, and the git commit hash as lineage metadata. The engineer wants the least operational overhead. Which approach meets these requirements?

Hard
76

A data scientist is using SageMaker to train a custom PyTorch model for image classification. They want to use SageMaker Debugger to detect training issues. Which TWO built-in rules are most relevant for detecting common training problems? (Select TWO.)

Medium
77

A company has a SageMaker endpoint that serves a recommendation model. The security team wants to ensure that the model artifacts stored in Amazon S3 are encrypted at rest and that access to the S3 bucket is limited to the SageMaker execution role only. The team also wants to receive alerts if the bucket policy is changed. Which combination of actions should the machine learning engineer take?

Easy
78

A team is using SageMaker to run a large-scale distributed training job for a language model. They are using SageMaker's Pipe mode to stream data from S3 to reduce IO. They observe that the training throughput is lower than expected, and the CPU utilization is high while GPU utilization is low. The training script uses PyTorch's DataLoader with num_workers=0. The data preprocessing is minimal. Which change is most likely to improve GPU utilization?

Hard
79

A machine learning engineer is using Amazon SageMaker Debugger to monitor a training job for a deep neural network. They receive a rule alert indicating 'exploding gradients'. Which action should they take to address this issue?

Hard
80

A data scientist is training a deep learning model using SageMaker and wants to use distributed training across multiple GPUs to reduce training time. Which TWO actions should the scientist take to configure distributed training? (Select TWO.)

Medium
81

A machine learning engineer is configuring auto-scaling for a SageMaker real-time endpoint. The endpoint is expected to have steady traffic during business hours and low traffic at night. The engineer wants to minimize costs by scaling in during low traffic, but the model container has a long start-up time (about 5 minutes). Which scaling policy should the engineer use to prevent request drops during sudden traffic spikes?

Medium
82

A machine learning engineer needs to optimize a trained TensorFlow model for deployment on edge devices with limited compute. Which SageMaker feature should they use to compile the model for target hardware?

Easy
83

A team is preparing text data for a natural language processing (NLP) model. They have a corpus of customer reviews. Which THREE preprocessing steps are essential to reduce noise and improve model performance?

Hard
84

An ML engineer wants to use MLflow on SageMaker to track experiments and log metrics. They have set up MLflow on an EC2 instance. How can they best integrate MLflow tracking with SageMaker training jobs?

Medium
85

A machine learning engineer is using Amazon SageMaker Data Wrangler to create a data preparation pipeline. The pipeline includes multiple transforms such as handling missing values, scaling, and encoding. The engineer wants to export the prepared data directly to a feature group in Amazon SageMaker Feature Store for reuse in training and inference. Which export option should the engineer choose?

Medium
86

A fraud detection team runs a SageMaker real-time endpoint that logs every request and response to an Amazon S3 bucket. Compliance requires that the model's prediction inputs and outputs be encrypted at rest with a customer-managed AWS KMS key, and that the endpoint be able to read the capture bucket only when necessary. The team enables data capture and specifies a KMS key on the endpoint configuration. Which additional configuration is required for the captured data written to S3 to be encrypted with that customer-managed key?

Medium
87

A data engineer needs to assess the quality of a dataset containing customer information. The dataset has missing values, outliers, and duplicate records. Which TWO AWS services can be used to perform data quality assessment? (Select TWO.)

Medium
88

A machine learning team needs to deploy a model that was built using scikit-learn. They want to use SageMaker for hosting. Which approach should they take?

Easy
89

A startup wants to deploy a model that has variable traffic patterns, with some periods of no traffic and occasional spikes. They want to pay only for what they use and do not want to manage instances. Which SageMaker inference option should they choose?

Medium
90

A data scientist is preparing a dataset stored in Amazon S3 for a SageMaker training job. The dataset contains missing values in several columns. The scientist wants to impute missing values with the mean of each column. Which SageMaker built-in algorithm or processing method should be used to perform this imputation efficiently?

Easy
91

A data scientist is using SageMaker Experiments to track multiple training runs. They want to compare different hyperparameter configurations and visualize the impact on model accuracy. What should they use to track hyperparameters?

Medium
92

A machine learning engineer at a retail company is monitoring a production model that predicts inventory demand. The model's prediction accuracy has dropped significantly over the past week. The engineer checks the model's input data and notices a new product category was introduced with a different distribution. Which concept is most likely causing the performance degradation?

Easy
93

A data engineer needs to prepare a dataset for a fraud detection model. The dataset contains a highly skewed numerical feature with extreme outliers. The engineer decides to apply a logarithmic transformation to this feature before training. Which SageMaker Data Wrangler transform should be used to apply the logarithmic transformation?

Medium
94

A team is using Amazon SageMaker for feature engineering. They have a dataset with a column 'TransactionDate' in string format (e.g., '2023-01-15 10:30:00'). They need to create features: year, month, day, hour, and day_of_week. What is the most efficient way to do this in a SageMaker processing job?

Medium
95

A team is preparing text data for sentiment analysis. They have a large corpus of customer reviews. They want to convert the text into numerical features using a technique that captures word importance relative to the whole corpus. Which feature extraction method should they use?

Medium
96

A company is building a real-time fraud detection system. They need to store features with historical context for model training and also support low-latency lookups for inference. Which THREE configurations should they set up in Amazon SageMaker Feature Store? (Select THREE.)

Hard
97

A machine learning team needs to ensure that all model training and inference jobs within SageMaker Studio run in a private network without internet access. The team also requires that inter-container traffic within the same training job be encrypted. Which configurations should they combine?

Hard
98

A company uses Amazon SageMaker Ground Truth to label a dataset for object detection. To reduce labeling costs, they want to use active learning. Which configuration should they set up in Ground Truth?

Hard
99

A company is using SageMaker Autopilot to automatically build a regression model on a dataset. They want to understand which features are most important for the model's predictions. Which feature of Autopilot can provide this insight?

Medium
100

A company uses Amazon SageMaker Ground Truth to create a labeled dataset. They want to monitor the accuracy of human labelers during the labeling process. Which metric should they track?

Hard
101

A data science team deploys a PyTorch model on Amazon SageMaker for real-time inference. The model requires GPU for low latency. Which instance type is MOST cost-effective while meeting the GPU requirement?

Easy
102

A data engineer needs to perform feature selection on a dataset with 500 numeric features to train a regression model. The engineer wants to remove features that are redundant or have low predictive power. Which TWO techniques should the engineer consider? (Select TWO.)

Medium
103

A data scientist is working on a binary classification problem and wants to use AWS Glue for data preparation. The dataset has missing values in several numeric columns. Which imputation strategy is MOST appropriate for the scientist to apply in AWS Glue ETL?

Easy
104

A data science team uses SageMaker to train and deploy models. They need to track model lineage, including datasets, training jobs, and model versions, to ensure reproducibility. Which THREE actions should they take? (Select THREE)

Medium
105

A data engineer is using Amazon SageMaker Processing to run a data preprocessing script on a dataset with 500 million rows. The script runs out of memory on a single ml.r5.24xlarge instance. The engineer needs to modify the processing job to handle the dataset size. Which approach is most cost-effective and scalable?

Hard
106

A machine learning team needs to deploy a new model version for A/B testing, gradually shifting traffic from the old version to the new version over 24 hours. Which deployment strategy should they use?

Medium
107

A data science team wants to track the lineage of models, including datasets, training jobs, and endpoints, for reproducibility and audit. They need a solution that captures relationships between artifacts automatically during training and deployment. Which service should they use?

Medium
108

A data scientist wants to track hyperparameters, metrics, and artifacts for multiple training runs in SageMaker. They need to compare runs and identify the best performing model. Which SageMaker feature should they use?

Medium
109

A machine learning engineer is using Amazon SageMaker Processing to preprocess a large dataset. The processing job runs a custom Python script that uses the pandas library to read multiple CSV files from an S3 input prefix. The script must write the processed output to a different S3 prefix. Which configuration of the ProcessingInput and ProcessingOutput parameters is correct for this scenario?

Hard
110

During data preparation for a regression model, a data scientist notices that two features have a Pearson correlation coefficient of 0.95. The scientist is concerned about multicollinearity. Which action should be taken to address this issue?

Medium
111

A company needs to perform time-series forecasting on historical sales data. Which SageMaker built-in algorithm is BEST suited for this task?

Easy
112

A team uses SageMaker Pipelines to automate model retraining. After a successful pipeline run, they want to register the new model version in the SageMaker Model Registry so that it can be reviewed for approval. Which step type should they add to the pipeline?

Easy
113

A gaming company uses a SageMaker endpoint for real-time player churn prediction. The model is updated weekly. After a recent retraining, the team notices that the endpoint's predicted probabilities for churn have shifted dramatically: the average predicted probability dropped from 0.3 to 0.05. The team suspects concept drift (the relationship between features and target changed) rather than data drift. They have SageMaker Model Monitor set up for data drift and quality metrics, but not for bias or explainability. The team needs to confirm concept drift and take corrective action. Which approach should the team take FIRST?

Medium
114

An ML team has deployed a model to a SageMaker real-time endpoint and wants to set up automated monitoring for model quality. Which TWO elements are required to configure SageMaker Model Monitor for model quality? (Select TWO.)

Medium
115

A data engineer must prepare a 4 TB Parquet dataset stored in Amazon S3 for a SageMaker training job that runs on 8 ml.p4d.24xlarge instances. The engineer wants the fastest possible data throughput during training while minimizing per-epoch I/O overhead. The dataset is immutable for the duration of the training run. Which approach BEST meets these requirements?

Medium
116

A team uses SageMaker Model Monitor to track data quality. They notice that the monitor's constraint violations are increasing but the model performance remains good. What should they do?

Medium
117

An ML team uses AWS Step Functions to orchestrate a retraining pipeline triggered by EventBridge when new training data arrives. The pipeline includes a SageMaker training job and a model evaluation. If evaluation fails, the team wants to send an alert. How should they implement this?

Hard
118

An ML engineer monitors a SageMaker endpoint for data drift. They set up SageMaker Model Monitor to compare inference data against a baseline created from the training dataset. The monitoring schedule runs daily and reports violations. Which monitoring type should be configured to detect if the distribution of a numerical feature in real-time inference data differs significantly from the training distribution?

Medium
119

A company needs to deploy a large language model (LLM) on SageMaker with the Triton Inference Server to maximize GPU utilization and reduce latency. They have an NVIDIA A100 GPU. Which SageMaker inference option supports Triton?

Hard
120

A machine learning team at a retail company has deployed a product recommendation model using Amazon SageMaker. The model is updated weekly with new data. Recently, the team noticed that the model's accuracy on a holdout evaluation set has been declining over the past month. The data pipeline that feeds the training job has not changed. The team suspects data drift. They have SageMaker Model Monitor enabled on the inference endpoint and have set up Amazon CloudWatch metrics for feature distribution distances. Upon reviewing the CloudWatch dashboards, they see that the feature distribution distance metric for the most important feature 'product_category' has increased significantly. However, the team is unsure if this is the root cause. Which remediation step should the team take FIRST?

Easy
121

A data science team uses SageMaker Pipelines for automated training. They need to conditionally register a model only if evaluation metrics exceed a threshold. Which pipeline step type should they use after the evaluation step?

Medium
122

A machine learning engineer is preparing a dataset for a binary classification model. The dataset has 10,000 rows and 200 features, with 5% positive class. The engineer suspects class imbalance may affect model performance. Which TWO actions should the engineer take to mitigate imbalance? (Choose 2.)

Medium
123

An ML engineer is using Amazon SageMaker Automatic Model Tuning (AMT) to optimize hyperparameters for a gradient boosting model. The tuning job is taking a long time and has completed many training jobs. The engineer wants to stop training jobs that are unlikely to improve the objective metric. What should they configure?

Medium
124

A company trains a model daily using Amazon SageMaker and uses the model for real-time inference. They want to detect data drift between the training data and the inference data to decide when to retrain. Which AWS service should they use for this purpose?

Medium
125

A machine learning engineer wants to store, share, and manage features for multiple ML models across an organization. The features need to be accessible for both real-time inference (low-latency) and batch training. Which AWS service should the engineer use?

Easy
126

A data engineer needs to split a time-series dataset into training and validation sets for a forecasting model. Which split method should be used to avoid data leakage?

Easy
127

A data scientist has a 40 GB CSV dataset in Amazon S3 that will be used to train a SageMaker model. The training script reads the data with pandas, and the scientist wants to reduce both storage cost and training-time I/O without changing the logical schema. Which data preparation action should be taken?

Easy
128

A team is deploying a model that requires GPU acceleration for inference. They are using an Amazon SageMaker real-time endpoint. The model is a large language model (LLM) that does not fit on a single GPU. Which configuration should they use to minimize latency while fitting the model?

Hard
129

Which SageMaker feature allows you to automatically tune hyperparameters using Bayesian optimization?

Easy
130

A data engineer wants to transform a categorical feature with 1,000 possible values into numerical features for a linear model. Which feature engineering technique is most appropriate for this high-cardinality feature?

Easy
131

A company wants to ensure that only authorized users and services can invoke a SageMaker real-time endpoint. Which AWS service can be used to manage access control?

Easy
132

A data team is using Amazon SageMaker Data Wrangler to prepare a dataset. They need to detect potential bias in the data before training a model. Which feature of Data Wrangler should they use?

Medium
133

A team deploys a PyTorch model on Amazon SageMaker for real-time inference. They notice that inference latency is higher than expected. They suspect the serialization format used for input data is inefficient. Which approach would MOST likely reduce latency?

Medium
134

A machine learning engineer is using SageMaker Automatic Model Tuning to optimize hyperparameters for a regression model. The objective metric is RMSE. The training job is costly, and the engineer wants to find a good configuration quickly. Which tuning strategy should they use?

Hard
135

A machine learning team is using Amazon SageMaker to train a model. They notice that the training job is taking longer than expected and the logs show repeated warnings about 'loss not decreasing'. Which SageMaker feature should they use to diagnose and visualize the training process?

Medium
136

Refer to the exhibit. A team has configured data capture for a SageMaker endpoint. The endpoint is returning predictions but no captured data appears in the S3 bucket. What is the most likely cause?

Easy
137

An ML engineer is preparing a time-series dataset for a forecasting model that predicts daily sales for the next 30 days. The dataset contains 3 years of daily sales data. Which data splitting strategy should the engineer use to evaluate the model's performance on future data?

Hard
138

A company wants to secure access to a SageMaker real-time endpoint. Which TWO actions should be taken? (Select two.)

Medium
139

A company runs a batch inference job on 10 TB of image data stored in S3. Each image needs to be processed by a GPU-accelerated model. The job is not time-sensitive and cost is the primary concern. Which SageMaker option is MOST appropriate?

Medium
140

A company wants to use SageMaker to serve real-time predictions with a model that has a large memory footprint. They need to ensure the endpoint can handle traffic spikes. Which scaling policy should they use?

Easy
141

A team is training a PyTorch model using SageMaker and wants to use their own custom training container with a specific PyTorch version. Which approach should they use?

Medium
142

A data engineer is preparing a dataset for a k-means clustering algorithm. The features have different scales: age (18-100), income ($20k-$200k), and number of purchases (0-50). Without scaling, which feature will dominate the distance calculations?

Easy
143

A data scientist is using SageMaker to train a linear regression model on a dataset with a large number of features. They notice that the model's training time is long and want to speed it up by using a more efficient algorithm. They decide to use the SageMaker built-in Linear Learner algorithm. Which of the following is a key advantage of using the Linear Learner algorithm in SageMaker for this scenario?

Easy
144

Refer to the exhibit. A data scientist configured an automatic model tuning job for a classification model. The tuning job completed after 20 training jobs, but the best validation accuracy was only 0.65. What is the most effective way to potentially improve the result?

Medium
145

A social media company is processing a real-time stream of user activity data from Amazon Kinesis Data Streams to train a machine learning model for content recommendation. The raw data includes user ID, timestamp, content ID, interaction type (like, share, comment), and device type. The data scientists need to aggregate features per user over a sliding window of 7 days, including counts of interaction types, unique content IDs engaged, and a moving average of interaction timestamps. The aggregated data will be used to update a user embedding model. The streaming data volume is approximately 500 records per second, and the company uses an AWS Glue streaming ETL job for transformation. However, the Glue job is failing frequently with high latency and checkpoint errors. The team needs a more robust solution to prepare the streaming data features. Which approach should the team take?

Hard
146

A machine learning engineer is deploying a model using SageMaker and needs to ensure that the endpoint can automatically scale based on traffic patterns. Which TWO actions should the engineer take? (Choose two.)

Medium
147

An ML engineer needs to create a feature store that supports both low-latency online inference and large-scale offline training. The features are updated hourly from a streaming source. Which Amazon SageMaker Feature Store configuration should the engineer use?

Easy
148

A company wants to forecast monthly sales that show clear seasonality. Which algorithm is most suitable?

Hard
149

A company uses SageMaker Model Monitor's feature attribution drift monitoring with SHAP. They receive an alert that the average SHAP value for a particular feature has increased significantly compared to the baseline. The feature's input distribution has not changed. What does this likely indicate?

Hard
150

A machine learning team is building a feature store using Amazon SageMaker Feature Store. They need to store features that support both real-time inference (low latency) and historical training. Which configuration should they choose?

Hard
151

A data engineer is preparing a dataset in Amazon SageMaker Data Wrangler for a binary classification model. The dataset contains missing values in several numeric columns, and the engineer wants a reusable, reproducible transformation that can be applied identically to the training data and to future inference data. The engineer plans to export the transformation and integrate it into a SageMaker Pipeline. Which TWO actions should the engineer take to ensure the imputation is consistent between training and inference? (Choose two.)

Medium
152

A company has a trained machine learning model that needs to be deployed as a real-time inference endpoint on Amazon SageMaker. The endpoint must automatically scale based on incoming traffic. Which SageMaker feature should be used?

Easy
153

A company wants to detect anomalies in login events from a large user base, focusing on unusual patterns that may indicate compromised accounts. Which SageMaker built-in algorithm is most suitable for this task?

Easy
154

A data scientist has a 200 GB Parquet dataset in Amazon S3 that will be used to train a SageMaker model. The training script reads the data with the SageMaker training toolkit's File mode, and the job currently spends a long time downloading before training begins. The team wants to reduce startup time without changing the training algorithm. Which change should the data scientist make?

Easy
155

A company wants to use a pre-trained NLP model from SageMaker JumpStart for sentiment analysis. Which step is required to make predictions?

Hard
156

A company wants to version and track ML models, with an approval workflow for promoting models from staging to production. Which SageMaker feature should they use?

Easy
157

A financial services company is deploying a model for loan approval. They must ensure that the model's predictions do not show bias against protected groups. They plan to monitor for bias drift after deployment. Which SageMaker feature should they use?

Medium
158

A machine learning engineer runs a training job and notices the loss is NaN after a few steps. Which SageMaker Debugger rule can help identify this issue?

Medium
159

A machine learning engineer is training a deep learning model on SageMaker using the PyTorch estimator. The training job fails with an error indicating that the GPU memory is exhausted. The engineer wants to reduce memory usage without changing the model architecture. Which SageMaker feature should the engineer use?

Hard
160

A financial institution uses SageMaker to train and deploy models. They need to track every experiment, model version, and deployment step for audit purposes. Which SageMaker feature should they use to capture the full lineage of artifacts, actions, and contexts?

Medium
161

A team is deploying a real-time inference endpoint in SageMaker. The model requires access to an S3 bucket containing customer data, which is encrypted with SSE-KMS. The team needs to ensure that the endpoint can decrypt the data. Which IAM role configuration is necessary?

Hard
162

A data science team is using SageMaker Experiments to track hyperparameters and metrics for a model training project. They need to compare multiple trials and identify the best model. Which THREE actions are part of a typical workflow? (Select THREE.)

Hard
163

A data scientist wants to train a binary classification model using Amazon SageMaker with a built-in algorithm that performs well on tabular data. Which algorithm should they choose?

Easy
164

A company is deploying a large NLP model on SageMaker for real-time inference. They want to reduce inference latency and cost by optimizing the model for the target hardware. The model is trained in PyTorch. Which SageMaker feature should they use to compile the model for best performance on the chosen instance?

Medium
165

A machine learning engineer must grant a data scientist the least-privilege permissions needed to invoke one specific SageMaker real-time endpoint from their own AWS account, and to view that endpoint's CloudWatch metrics without being able to modify the endpoint. The endpoint ARN is known. Which TWO IAM policy statements should the engineer include? (Choose two.)

Medium
166

A healthcare analytics team stores model artifacts and training datasets in Amazon S3 and uses SageMaker. An internal audit finds that some S3 buckets containing protected health information are missing encryption and that access is granted broadly. The team must remediate quickly and prevent future misconfiguration. Which combination of actions should the ML engineer take FIRST?

Medium
167

Refer to the exhibit. A data scientist reviews the CloudWatch Logs from an Amazon SageMaker real-time endpoint. What is the MOST likely root cause of the NaN output?

Easy
168

A machine learning engineer runs a SageMaker Processing job that must load a 200 GB dataset from S3, compute statistics, and write a small summary to S3. The job repeatedly fails with an out-of-disk-space error on the processing instance. Which change is MOST likely to resolve the failure?

Hard
169

A machine learning engineer uses Amazon SageMaker Data Wrangler to preprocess a dataset. After applying a transform, the engineer wants to export the data to a feature group in Amazon SageMaker Feature Store for reuse in training and inference. Which export option should they choose?

Medium
170

A data scientist is preparing text data for a sentiment analysis model using Amazon SageMaker. Which two data preprocessing techniques are commonly used when working with text data for natural language processing? (Choose two.)

Medium
171

A company is training a deep learning model using SageMaker and wants to reduce the time spent on data loading from Amazon S3 during training. The training dataset consists of many small files. Which approach is MOST effective to accelerate data loading?

Hard
172

A data scientist wants to track feature definitions, share them across teams, and serve features for both training and real-time inference. Which AWS service provides these capabilities?

Easy
173

Which TWO of the following are best practices for deploying machine learning models on SageMaker? (Select TWO.)

Medium
174

A machine learning engineer is performing feature selection for a regression model with 200 features. The dataset has 10,000 samples. The engineer wants to remove irrelevant features while keeping those that have a strong non-linear relationship with the target. Which feature selection method is best suited for this requirement?

Hard
175

Refer to the exhibit. A data scientist ran a training job using a custom algorithm container. The job failed with the error shown. What is the most likely cause?

Easy
176

A team is using AWS Step Functions to orchestrate a machine learning workflow that includes data preprocessing, training, and model evaluation. The team wants to run the workflow whenever new data arrives in an S3 bucket. Which approach should they use to trigger the Step Functions workflow?

Medium
177

A data engineer is building a data pipeline for a machine learning model that requires both structured and unstructured data. The structured data (customer demographics) is in Amazon RDS, and the unstructured data (customer support chat logs) is in Amazon S3 as JSON files. The engineer needs to combine these datasets into a single training dataset stored in S3 in Parquet format. They must also perform feature engineering such as text vectorization on the chat logs. The pipeline should be serverless and cost-effective. Which approach should they use?

Medium
178

A machine learning engineer is preparing a training job on SageMaker with a custom Docker container. Which TWO actions are required to use the container with SageMaker? (Choose TWO.)

Medium
179

A data scientist is building a binary classification model on a highly imbalanced dataset where the positive class represents only 1% of the data. The scientist needs to train the model using Amazon SageMaker's built-in XGBoost algorithm. Which strategy should be used to address the class imbalance?

Medium
180

An ML team uses AWS Step Functions to orchestrate a multi-step inference pipeline: data preprocessing, model inference, and postprocessing. The pipeline runs on demand for single records. The team notices that the pipeline occasionally fails due to timeouts in the preprocessing step. They want to implement retries with exponential backoff and a maximum retry count of 3 for that step. How should they configure this?

Medium
181

A data scientist is preparing a dataset for a linear regression model. The features have different scales: one feature ranges from 0 to 1000, another from 0 to 1, and a third from -5 to 5. The scientist wants to ensure that all features contribute equally to the model. Which TWO scaling techniques should the scientist consider? (Select TWO.)

Medium
182

An engineer runs: aws sagemaker describe-endpoint --endpoint-name my-endpoint and receives the exhibit output. The engineer wants to update the endpoint to use a new model version stored in ECR with tag ':2'. Which step is necessary to perform the update?

Medium
183

A team wants to evaluate a binary classification model for credit risk. They need to understand the trade-off between false positives and false negatives. Which TWO metrics should they use? (Select TWO.)

Medium
184

A data engineer needs to catalog metadata from multiple data sources across the organization for use in ML workflows. Which AWS Glue component should be used to store and manage this metadata?

Easy
185

A data scientist needs to split a dataset into training, validation, and test sets. The dataset has a categorical target variable with imbalanced class distribution. Which splitting technique ensures that each subset has a similar proportion of each class?

Medium
186

A data scientist is preparing a dataset for a binary classification model. The dataset has 10,000 samples, but the positive class represents only 2% of the data. The data scientist needs to train a model that will be evaluated on a hold-out test set that preserves the original class distribution. Which data preparation strategy is MOST appropriate?

Medium
187

A company uses SageMaker Pipelines to automate model retraining. The pipeline runs daily but sometimes fails due to data quality issues. What is the best design to handle this?

Medium
188

A data engineer is preparing a dataset for a SageMaker training job. The dataset contains a timestamp column and is stored in Amazon S3 as CSV files. The engineer needs to ensure that the training job reads the data efficiently and that the data is partitioned by date to improve query performance in Amazon Athena. Which action should the engineer take?

Medium
189

A company wants to use SageMaker Autopilot to automatically build a binary classification model. Which output does Autopilot provide to help understand model decisions?

Medium
190

A machine learning engineer is using AWS Glue ETL to transform a large dataset stored in Amazon S3. The transformation involves joining two tables on a high-cardinality column and aggregating results. The job is running slowly and the engineer needs to improve performance. Which optimization technique should the engineer apply?

Medium
191

A data scientist runs this pipeline but the Train step fails with "ResourceLimitExceeded". What is the most likely cause?

Medium
192

A data scientist has trained a model that achieves 95% accuracy on the training set but only 70% on the test set. Which of the following is the most likely cause?

Medium
193

A company deploys a model using SageMaker and enables data capture for monitoring. After a week, they notice that the captured data is not being written to the specified S3 bucket. The endpoint is running and invocations are successful. What is the most likely cause?

Hard
194

Refer to the exhibit. An IAM policy is attached to a user to allow invoking a SageMaker endpoint. A developer tries to call the endpoint from a laptop with IP 203.0.113.5 and receives an access denied error. What is the most likely reason?

Hard
195

A data scientist is working with a dataset containing a categorical feature 'country' with 200 unique values. They plan to use a linear regression model. Which encoding method is most suitable to avoid the dummy variable trap while maintaining interpretability?

Medium
196

An ML engineer is setting up monitoring for a SageMaker endpoint. Which THREE metrics should be monitored to detect performance issues? (Select THREE.)

Medium
197

A SageMaker endpoint is logging an error when processing inference requests that require database access. What is the most likely cause?

Medium
198

A data engineer is using AWS Glue to prepare a dataset for ML. The engineer wants to split the dataset into training and testing sets while preserving the distribution of the target variable. Which TWO methods achieve this goal? (Select TWO)

Easy
199

A data scientist is using SageMaker Data Wrangler to prepare a dataset for a binary classification model. The dataset contains a mix of numerical and categorical features. The scientist wants to perform feature engineering to improve model performance. Which TWO actions are appropriate for handling categorical features in Data Wrangler? (Choose two.)

Medium
200

Which SageMaker built-in algorithm should be used for forecasting time series data with seasonal patterns?

Easy
201

A data scientist is using Amazon SageMaker Data Wrangler to prepare a dataset for classification. They want to detect potential bias in the data before training. Which SageMaker service should they use in conjunction with Data Wrangler to detect bias?

Medium
202

A retail company uses a SageMaker Model Monitor data quality monitor on a real-time endpoint. The monitor's baseline was generated from a training dataset in which the "promo_code" feature was often null. In production the feature is now populated for nearly every record, and the monitor reports violations even though model accuracy has not degraded. The team wants the monitor to stop flagging this expected change without disabling monitoring entirely. What should they do?

Hard
203

A machine learning engineer trains a binary classifier and obtains an accuracy of 95% on the test set. The dataset is imbalanced with 95% positive class. What is the most important metric to evaluate the model's performance?

Easy
204

A financial services company must ensure that a SageMaker model deployed to a real-time endpoint only produces predictions consistent with a fairness constraint on a protected attribute, and that any violation is detected within minutes and triggers an alert to the compliance team. The model is already deployed and monitored for data quality. Which approach should the machine learning engineer implement?

Hard
205

A data science team detects that a deployed model's prediction accuracy is degrading over time due to concept drift. They need to implement a retraining strategy. Which THREE actions are recommended best practices for handling concept drift?

Medium
206

A company's ML pipeline runs in multiple AWS accounts (dev, test, prod). They want to enforce that only approved models from a central Model Registry can be deployed to the production account. Which combination of services is MOST appropriate to implement this governance?

Hard
207

A data scientist needs to ingest streaming customer clickstream data from a website into an S3 data lake for ML training. The data must be delivered within 1 minute of ingestion, and JSON records must be converted to Parquet. Which AWS service combination should be used?

Medium
208

A machine learning engineer wants to reduce training costs by using excess EC2 capacity. Which instance purchasing option should they choose for SageMaker training jobs?

Easy
209

A data scientist is training a large transformer model using SageMaker's model parallelism library. The training job is failing with an out-of-memory (OOM) error. Which two actions can help resolve the OOM error? (Choose two.)

Hard
210

A data engineer is designing a feature engineering pipeline using Amazon SageMaker Feature Store. The team needs to support both real-time inference (millisecond latency) and batch training jobs that require access to historical feature values at specific points in time. Which configuration should the engineer choose?

Medium
211

A data scientist wants to train a model on SageMaker using a custom PyTorch script, then register the best model in the SageMaker Model Registry. The training job is part of a SageMaker Pipeline. Which pipeline step should be used to register the model?

Medium
212

A company wants to enable cross-account access to a SageMaker model endpoint. The model is in Account A, and Account B needs to invoke it. Which TWO steps are required? (Select TWO)

Hard
213

A team is building a recommendation system and wants to store and serve features for online and offline models. The features include user statistics (updated daily) and movie metadata (static). The team needs low-latency inference for real-time recommendations and wants to reuse features across multiple models. Which AWS service should the team use to store, manage, and serve these features?

Medium
214

A machine learning engineer needs to split a time-series dataset for a forecasting model. The data spans 3 years of daily sales. Which splitting strategy should they use to avoid look-ahead bias?

Medium
215

A team is building a time-series forecasting model for daily sales data. They want to evaluate model performance using cross-validation while respecting the temporal order of the data. Which data splitting strategy should they use?

Hard
216

A team is using Amazon SageMaker Data Wrangler to prepare a large dataset. They need to detect potential bias in the data before training. Which capability of Data Wrangler should they use?

Hard
217

Refer to the exhibit. A data engineer investigates why a SageMaker endpoint is returning errors. The endpoint configuration has been updated to point to a new model version. What is the MOST likely cause of the error?

Medium
218

A company is using SageMaker Pipelines to orchestrate their ML workflow. They have a Condition step that checks if a model's accuracy exceeds 0.9. If true, they want to register the model in the model registry; otherwise, they want to run a retraining step. Which step type should they use for the decision?

Medium
219

A company wants to deploy 50 small models (each ~100 MB) for real-time inference. They need to minimize hosting costs while maintaining low latency. Which SageMaker hosting option is most cost-effective?

Medium
220

An ML team uses Amazon SageMaker Data Wrangler to prepare a dataset for a binary classification model. They suspect the dataset might contain bias against a certain demographic group. They want to detect and visualize potential bias before training the model. Which feature of SageMaker should they use?

Medium
221

A retail company stores training datasets, model artifacts, and feature data in Amazon S3. An auditor requires that all objects be encrypted at rest with keys the company controls and that key usage be independently auditable. The team wants minimal operational overhead. Which approach should the ML engineer recommend?

Easy
222

A team deploys a machine learning model using an Amazon SageMaker endpoint. They need to monitor for data drift and model quality issues. Which AWS services or features should they use? (Choose THREE.)

Medium
223

A company uses Amazon SageMaker Data Wrangler to prepare data for ML. The dataset contains a timestamp column and sensor readings from IoT devices. The data scientist needs to create features such as moving averages and rolling statistics over time windows. Which Data Wrangler transformation type should be selected?

Hard
224

A machine learning engineer is preparing a training dataset in Amazon SageMaker for a binary classification model. The dataset is stored as a single CSV file in Amazon S3 and contains 12 categorical features with high cardinality (thousands of unique values each). The engineer wants to avoid the curse of dimensionality and reduce training time while preserving predictive power. Which preprocessing approach should be used with the SageMaker built-in XGBoost algorithm?

Medium
225

A team is building a SageMaker Pipeline that trains a model and then registers it in the SageMaker Model Registry. They want the pipeline to automatically deploy the model to a real-time endpoint only after a human approves the model package. Which two actions should the team take to implement this approval-gated deployment? (Choose two.)

Hard
226

A data scientist is building a text classification model using a pre-trained BERT model from the Hugging Face library on SageMaker. The scientist wants to fine-tune the model on a custom dataset. Which TWO steps are necessary to set up the fine-tuning job? (Select TWO.)

Medium
227

A data science team has trained a PyTorch model for real-time inference and needs to deploy it on AWS with GPU acceleration while minimizing cold-start latency. Which SageMaker inference option should they choose?

Easy
228

A team wants to use a custom PyTorch training script in SageMaker. They need to install additional Python packages not included in the base PyTorch container. Which approach should they take?

Medium
229

Which TWO SageMaker Pipelines steps are essential for automating a complete ML workflow from data processing to model deployment? (Choose 2.)

Medium
230

A data scientist is working with a dataset that contains missing values in several numeric features. The data scientist wants to impute the missing values with the median of each feature. Which Amazon SageMaker Data Wrangler transformation should be used?

Easy
231

A team is deploying a TensorFlow model on a SageMaker real-time endpoint with automatic scaling. They set the scaling policy to target an average CPU utilization of 50%. However, during traffic spikes, the endpoint experiences high latency and 503 errors. The instance type is ml.c5.large. What should the team do to resolve this while minimizing cost?

Hard
232

A company deploys a machine learning model as a SageMaker real-time endpoint. They need to implement a mechanism to automatically roll back to the previous model version if performance degrades after a deployment. Which approach should they use?

Hard
233

After deploying a model to a SageMaker endpoint, the operations team notices high inference latency. They suspect it is due to insufficient instance capacity. Which first step should they take to diagnose the issue?

Medium
234

A machine learning engineer is using SageMaker to train a model with the built-in LightGBM algorithm. The engineer wants to use early stopping to prevent overfitting. The training job is configured with a validation dataset. Which hyperparameter should be set to enable early stopping?

Hard
235

A company wants to automate its machine learning pipeline using AWS CodePipeline and Amazon SageMaker. The pipeline should train a model, evaluate it, and if the evaluation passes, register the model in the SageMaker Model Registry. Which service should the company use to orchestrate the training and evaluation steps?

Easy
236

A team receives alerts that their SageMaker endpoint latency has increased significantly. They check CloudWatch metrics and see Invocations rising, but ModelLatency remains stable. Which metric should they investigate to find the source of the increased latency?

Medium
237

A data science team wants to host 50 different models for a recommendation engine. Each model is small (under 100 MB) and traffic patterns are unpredictable. They need to minimize cost and operational overhead. Which approach should they take?

Hard
238

A company uses SageMaker Clarify to detect bias in their training data. They find that the model has a high disparate impact for a protected attribute. What should they do to mitigate this bias during training?

Medium
239

A machine learning engineer is deploying a model to a SageMaker endpoint and wants to ensure that the model's predictions can be explained. The engineer needs to understand which features contributed most to each prediction. Which SageMaker feature should be used?

Hard
240

A media company runs a real-time recommendation model on a SageMaker endpoint. Traffic triples every evening between 18:00 and 22:00 and drops to near zero overnight, and the team wants to cut costs without underserving evening users. They want the endpoint to scale out automatically as invocations rise and scale back in when traffic falls. Which solution should they implement?

Medium
241

A team uses SageMaker Clarify to monitor bias drift on a deployed model. They have defined a baseline with training data and set up a monitoring schedule. After one month, they receive a violation report indicating that the post-training metrics have deviated from the baseline. What does this violation indicate?

Medium
242

A team is using Amazon SageMaker Processing for data preprocessing. They have a Parquet dataset in Amazon S3. Which configuration will provide the most efficient reading of the dataset during processing?

Medium
243

A machine learning engineer wants to automatically track hyperparameters, metrics, and artifacts for multiple training runs. Which SageMaker feature should they use?

Easy
244

A financial services company deploys a fraud detection model on a SageMaker real-time endpoint. The inference logic includes a pre-processing step that requires access to a DynamoDB table for user metadata. The model container is a custom Docker image. How should the team grant the endpoint access to DynamoDB?

Hard
245

A data scientist is performing text preprocessing for a sentiment analysis model. The dataset contains many stop words and rare words. Which combination of preprocessing steps will reduce dimensionality and improve model performance?

Medium
246

An organization uses SageMaker Studio and needs to restrict Studio's internet access while allowing users to install custom packages from a private PyPI mirror hosted in a VPC. Which networking configuration should they use?

Hard
247

An organization needs to ensure that all data used for inference on a SageMaker endpoint is encrypted at rest. The endpoint uses a SageMaker-provided container. Which configuration should be applied?

Easy
248

A data engineer is setting up a Glue ETL job to process a large dataset stored in Amazon S3. The job needs to read data in Parquet format, apply a filter, and write the results back to S3 in Parquet. The engineer wants to minimize the cost and runtime. Which optimization technique is MOST effective?

Easy
249

A company is running a SageMaker endpoint serving multiple models. They need to monitor for data drift and model quality. Which THREE actions are necessary? (Choose three.)

Hard
250

A financial services company trains multiple models on SageMaker and needs to track hyperparameters, metrics, and artifacts for each experiment. Which SageMaker feature should they use to organize and compare experiments?

Medium
251

A company operates an e-commerce platform that uses a machine learning model to recommend products to users. The model is deployed on an Amazon SageMaker endpoint with automatic scaling enabled based on average CPU utilization. The model was trained on historical data and is updated weekly. Recently, the platform experienced a flash sale event that caused a sudden spike in traffic. During the event, the endpoint's latency increased dramatically, and many requests timed out. After the event, the team reviews the CloudWatch metrics and notices that the CPU utilization never exceeded 70%, and the scaling policy was triggered but instances took several minutes to become available. The team wants to prevent similar issues in future flash sales. Which course of action would be MOST effective?

Hard
252

An ML team wants to deploy a model that was trained using XGBoost in SageMaker. They want to use the built-in XGBoost algorithm container for inference. Which inference option requires the least custom code?

Easy
253

A company wants to serve a large ensemble of models using NVIDIA Triton Inference Server on SageMaker for high throughput GPU inference. Which SageMaker inference option supports this?

Hard
254

A data engineer is using Amazon SageMaker Data Wrangler to create a data preparation flow for a dataset with 500 columns, many of which are highly correlated. The goal is to reduce dimensionality while preserving interpretability. Which built-in transform in Data Wrangler should be applied?

Hard
255

A retailer runs a nightly batch scoring job that processes 40 GB of transaction data and writes predictions to S3. Occasionally a single partition is corrupt, causing the entire job to fail after several hours. The team wants the job to skip the corrupt partition, log which partition failed, and still complete processing of the remaining data with minimal changes to their existing SageMaker Processing job. Which change should they make?

Medium
256

A machine learning engineer is using SageMaker Automatic Model Tuning (AMT) to optimize hyperparameters for a random forest model. The engineer notices that the tuning job is taking too long and many hyperparameter combinations are being evaluated but not improving the objective metric. Which action should the engineer take to make the tuning more efficient?

Hard
257

A team uses SageMaker Pipelines to train and register a model. They want to conditionally run a hyperparameter tuning step only if the data quality check passes. Which pipeline step type should they use to branch the execution?

Hard
258

Which SageMaker built-in algorithm is specifically designed for time series forecasting?

Easy
259

A company has a SageMaker real-time endpoint that serves predictions. They want to set up automated monitoring and remediation for when the number of 5XX errors exceeds a threshold. Which TWO steps should they take? (Choose TWO.)

Medium
260

A machine learning team is deploying a model to a SageMaker endpoint and needs to implement A/B testing between two model versions. They want to split traffic 80/20 and monitor performance metrics for each variant. Which two actions should they take? (Choose two.)

Hard
261

A company uses AWS Glue ETL jobs to clean and transform data from S3 before training. The data contains a column with 40% missing values. The column is normally distributed. Which imputation strategy should the data engineer use?

Medium
262

A data scientist is preparing a CSV dataset in Amazon S3 for a SageMaker training job. Several rows contain missing values in numeric feature columns, and the chosen algorithm cannot handle NaNs. The scientist wants a repeatable, code-based transformation that runs inside a SageMaker Processing job before training. Which step is the MOST appropriate?

Easy
263

A company needs to deploy a new model version to a SageMaker real-time endpoint. They want to route 5% of traffic to the new version initially to monitor for errors before full rollout. Which deployment strategy should they use?

Medium
264

A machine learning engineer observes that a SageMaker training job fails with the error shown in the exhibit. What is the most likely cause of the failure?

Medium
265

A machine learning engineer needs to run a one-time scoring job over 500 GB of data stored in Amazon S3 using a trained model, and the results must be written back to S3. There is no requirement for a persistent HTTPS endpoint. Which SageMaker feature should the engineer use?

Easy
266

A company requires that all SageMaker notebook instances be created within a private VPC without internet access. Which configuration step is mandatory?

Easy
267

A company is fine-tuning a large language model using LoRA on SageMaker. They want to reduce GPU memory usage during training. Which configuration change would help?

Hard
268

A machine learning engineer is building a pipeline to preprocess text data for a sentiment analysis model. The data consists of customer reviews. The engineer wants to convert the text into numerical features while preserving the semantic meaning of words. Which technique should be used?

Medium
269

Refer to the exhibit. A user launches a SageMaker notebook instance with this lifecycle configuration. What happens?

Easy
270

A data scientist trains a binary classification model using SageMaker and obtains an AUC of 0.95 on the test set. However, the precision-recall curve shows low precision for high recall thresholds. The business requires a model that performs well on the minority class. Which metric should the team primarily optimize during hyperparameter tuning?

Hard
271

A data engineer is preparing a dataset for a time series forecasting model. The dataset contains a timestamp column and a target variable. The engineer wants to create additional features such as lag values and rolling averages. Which SageMaker Data Wrangler transform should be used to generate these time series features?

Medium
272

A data science team deployed a model on Amazon SageMaker and enabled Model Monitor to detect data drift. After a week, they receive alerts indicating that the distribution of a key feature has shifted significantly. However, the model's accuracy on the recent production data remains high. Which action should the team take next?

Medium
273

A data science team deploys a regression model using Amazon SageMaker. After one week, the model's prediction accuracy drops significantly. The team needs to detect this degradation automatically and trigger retraining. Which AWS service should they use to monitor the model's performance over time and set up alerts?

Easy
274

A data scientist notices that a production model's accuracy has degraded over the past week. The training data distribution remains unchanged, but the relationship between features and the target has shifted. Which type of drift is occurring, and which monitoring approach should be used?

Medium
275

A company needs to detect bias in a pre-trained model before deployment. They want to compute metrics like disparate impact and equal opportunity difference. Which AWS service should they use?

Hard
276

A data engineer needs to prepare a large dataset (10 TB) stored in Amazon S3 for a training job on SageMaker. The data is in CSV format, but the training algorithm expects Parquet for performance. The engineer must transform the data with minimal cost and without writing custom code. Which service should be used?

Medium
277

A data scientist needs to create a feature group in Amazon SageMaker Feature Store for real-time recommendations. Which TWO configurations are required? (Select TWO.)

Medium
278

A data scientist is training a binary classifier on a highly imbalanced dataset (1:100 class ratio). The dataset contains 500,000 rows and 30 features. The data is stored in S3 in Parquet format. The data scientist wants to use SageMaker's built-in XGBoost algorithm. Which data preparation technique should the data scientist apply to best address the class imbalance without causing data leakage?

Hard
279

A machine learning team deploys a fraud detection model on a SageMaker endpoint. The model's predictions are used in real-time. The team wants to monitor for data drift by comparing incoming data distributions against a baseline created from the training data. Which SageMaker capability should they use?

Medium
280

An ML team uses SageMaker Pipelines to automate model retraining. They want to skip redundant training steps when input data has not changed. Which feature should they enable?

Hard
281

A company plans to deploy a large foundation model using SageMaker JumpStart. They are concerned about costs because the model will be used intermittently. Which deployment option is MOST cost-effective for intermittent traffic?

Medium
282

Refer to the exhibit. The training job failed. What is the MOST likely cause?

Hard
283

A team is deploying a model using SageMaker Pipelines. They have defined a pipeline with steps: preprocessing, training, evaluation, and conditional registration. The evaluation step produces a JSON file with metrics. If accuracy > 0.9, the model is registered; else, the pipeline fails. Which TWO statements about this pipeline are correct? (Choose TWO.)

Hard
284

A machine learning engineer is deploying a real-time inference endpoint on Amazon SageMaker AI for a fraud detection model. The model must serve predictions with consistent latency under 50 ms and the team expects traffic to fluctuate unpredictably, with occasional bursts. The engineer wants to automatically adjust the number of instances based on actual workload while minimizing cost during idle periods. Which SageMaker AI feature should the engineer configure?

Medium
285

A data scientist is training a deep learning model on SageMaker and notices that the training loss oscillates and does not converge. They want to debug this issue. Which SageMaker feature can they use to monitor and analyze the training process?

Easy
286

A team has a large number of models that need to be deployed for batch inference weekly. They want to minimize cost and management overhead. Which approach is MOST efficient?

Medium
287

A company wants to deploy a foundation model from SageMaker JumpStart with the lowest possible inference cost, given that latency requirements are flexible. They have a mix of traffic volumes. Which approach should they take?

Medium
288

A company wants to deploy a machine learning model using infrastructure as code to ensure reproducibility. They need to define the SageMaker Studio domain, user profiles, and the endpoint configuration. Which tool should they use?

Medium
289

An ML engineer creates a SageMaker inference pipeline with two containers: a preprocessor and a predictor. The preprocessor is a lightweight Python script that transforms input data. How should the engineer structure the endpoints to ensure both containers run sequentially?

Medium
290

Which SageMaker feature compiles a trained model into an optimized binary for a specific hardware target (e.g., Intel, ARM, NVIDIA, or edge devices) to improve inference performance?

Easy
291

A team is training a large language model using PyTorch on SageMaker. They need to reduce training time. The model has 10 billion parameters. Which distributed training strategy should they use?

Medium
292

A machine learning engineer is configuring a SageMaker Processing job that runs a custom container to compute bias metrics on a dataset containing personally identifiable information. The job reads input data from one S3 bucket and writes reports to another, and the security team requires that the container has no outbound internet access and that the input and output buckets are reached without traversing the public internet. Which configuration satisfies these requirements?

Hard
293

A company deploys a model for fraud detection. They want to monitor if the model's predictions become less accurate over time due to changes in the underlying data distribution, but they do not have immediate access to ground truth labels. Which type of drift should they monitor as a proxy?

Medium
294

A company is training a deep learning model with SageMaker and wants to reduce training time by using pipe mode instead of file mode for a large dataset stored as TFRecord files in Amazon S3. After switching the estimator's input mode to Pipe, the training job fails immediately with a dataset format error. The data scientist confirms the files are valid TFRecords and that the same script works with File mode. What is the most likely cause?

Hard
295

A company needs to serve real-time predictions from a large ensemble of three deep learning models, each requiring different inference environments (PyTorch, TensorFlow, MXNet). Which SageMaker endpoint type supports running multiple inference containers together?

Medium
296

A company wants to build a customer service chatbot that answers questions about their internal policy documents. The documents are updated monthly, and the team cannot afford to retrain a model each time. Which approach is MOST appropriate?

Medium
297

An organization stores raw data in Amazon S3 as CSV files. They need to perform serverless data transformation and convert the data to Parquet format for efficient ML training. Which AWS service is most appropriate?

Easy
298

A team is building a fraud detection model using SageMaker and wants to detect anomalies in user login events. Which SageMaker built-in algorithm is specifically designed for anomaly detection in event-based data?

Medium
299

A retail company has a SageMaker model that predicts customer churn. The model was trained on data that included a 'customer_zipcode' feature. After deployment, the data science team notices that the model's predictions for certain zip codes have become less accurate over time. They suspect that the relationship between zip code and churn has changed due to a recent relocation of a major employer. Which SageMaker monitoring capability should they use to detect this type of drift?

Medium
300

A data scientist is preparing text data for sentiment analysis. They need to convert the text into numerical features while reducing the impact of common words. Which feature extraction method should they use?

Medium
301

A financial services company operates a real-time inference endpoint for a fraud detection model on Amazon SageMaker. The model was trained on historical transaction data from 2023. Over the past month, the model's precision has dropped from 92% to 78%, while recall remains high at 95%. The data science team suspects data drift and has already enabled SageMaker Model Monitor with data capture and a baseline from the training data. The latest monitoring report indicates no statistically significant drift in any of the input features. The team also verified that the inference code and model artifact have not changed. Despite the stable feature distributions, the model is misclassifying an increasing number of legitimate transactions as fraudulent (false positives). The business is concerned about the impact on customer experience. What is the best course of action?

Hard
302

Refer to the exhibit. The data scientist wants to update the endpoint to use a new model version without downtime. Which approach should they use?

Easy
303

A team wants to secure SageMaker endpoints for a healthcare application. They must ensure data is encrypted at rest and in transit, and that the endpoint can only be accessed from within a VPC. Which THREE steps should they take? (Select THREE)

Medium
304

A data scientist is using Amazon SageMaker Data Wrangler to create a data preparation flow. After completing the flow, the scientist wants to export the processed data to a feature group in Amazon SageMaker Feature Store for reuse in multiple training jobs. Which export option should the scientist choose in Data Wrangler?

Medium
305

A data scientist needs to store training data in Amazon S3 and wants to optimize read performance for iterative training jobs. Which S3 feature should they use?

Easy
306

A fraud-detection team trains a model in SageMaker and wants to shift 10 percent of live prediction traffic to a newly retrained model to compare accuracy before a full cutover. The endpoint already serves the current model on one production variant. They need the endpoint to route a controlled fraction of requests to the new model without changing the client application. What should they do?

Medium
307

A team deploys a SageMaker real-time endpoint and configures it with an auto scaling policy targeting a variant. During a flash sale, traffic spikes and the team notices that the number of instances increases, but the average model latency still climbs above the target. The team wants the scaling behavior to react faster to sudden bursts without over-provisioning during steady periods. Which change should they make to the scaling policy?

Easy
308

A company ingests streaming transaction data from multiple sources using Amazon Kinesis Data Streams. The data must be transformed (e.g., JSON parsing, data type conversions) and then stored in Amazon S3 for ML training. The transformation logic may change over time. Which approach provides the greatest flexibility and ease of maintenance?

Medium
309

A company wants to use SageMaker to fine-tune a foundation model for a text generation task using RLHF (Reinforcement Learning from Human Feedback). Which THREE components are required in the RLHF pipeline?

Hard
310

A data scientist is using Amazon SageMaker Data Wrangler to prepare a dataset. They need to identify potential bias in the data before training. Which SageMaker feature should they use?

Easy
311

A machine learning engineer has trained a model in SageMaker and wants to deploy it to a real-time endpoint for low-latency inference. The model artifacts are stored in Amazon S3, and the engineer needs to create the endpoint with the least operational effort. Which sequence of actions should the engineer take?

Easy
312

A data scientist needs to run a hyperparameter tuning job for a PyTorch model using SageMaker. They want to use Hyperband for efficient resource allocation. Which tuning strategy should they select in the HyperparameterTuner?

Medium
313

A data science team deploys a real-time inference endpoint on Amazon SageMaker. They want to monitor for data drift in the input features over time. Which AWS service should they use to capture and analyze the input data distribution?

Easy
314

A data science team needs to deploy a PyTorch model for real-time inference with low latency. The model requires GPU acceleration. Which SageMaker endpoint configuration should they use?

Easy
315

A data scientist discovers that a dataset for binary classification contains 95% negative samples and 5% positive samples. Which technique is MOST appropriate to address the class imbalance?

Easy
316

A company wants to serve 200 different PyTorch models. Each model is small (under 1 GB) and only a fraction are used at any time. To minimize cost and management overhead, which SageMaker inference option should be used?

Medium
317

A machine learning engineer is evaluating a binary classification model for detecting fraudulent transactions. The dataset is highly imbalanced, and the cost of false negatives (missing a fraud) is very high. Which two evaluation metrics should the engineer consider? (Choose two.)

Hard
318

An ML team is using Amazon SageMaker Feature Store to serve features for both real-time inference and batch training. They need to ensure that training data uses feature values as they were at the time of each event. Which type of query should they use?

Medium
319

A company deploys a real-time inference endpoint with auto-scaling using a target tracking policy based on average Invocations per instance. They notice that during a traffic spike, the endpoint scales out too late, causing increased latency. They want to scale proactively before the spike. Which strategy should they implement?

Hard
320

A company wants to deploy a PyTorch model on SageMaker for real-time inference. Which two steps are required? (Select TWO.)

Medium
321

A company is building a sentiment analysis model for customer reviews. The text data contains many typos, abbreviations, and informal language. Which text preprocessing step would be most beneficial to support the ML model?

Medium
322

A company is using SageMaker Debugger to monitor a training job for a deep learning model. They want to detect when gradients become extremely large, which may cause training instability. Which built-in rule should they use?

Medium
323

A machine learning engineer is using Amazon SageMaker Experiments to track multiple training runs. They want to compare the performance of different hyperparameter configurations visually. Which SageMaker tool provides an interactive interface to compare experiments?

Easy
324

A team is fine-tuning a foundation model using reinforcement learning from human feedback (RLHF) on SageMaker. They have a dataset of human preferences. Which SageMaker capability is most suitable for the reward model training step?

Hard
325

A data scientist needs to annotate a large dataset of images for an object detection model. The team wants to minimize manual labeling effort and cost. Which Amazon SageMaker feature should they use?

Easy
326

A company is setting up a data pipeline to ingest streaming clickstream data from their website for real-time analytics and machine learning. The data must be reliably ingested, transformed, and stored in Amazon S3 for batch processing. Which combination of AWS services should be used?

Easy
327

A team is fine-tuning a Hugging Face transformer model on SageMaker. They need to use a custom training script with the Hugging Face Estimator. Which SageMaker feature does this represent?

Medium
328

A machine learning team notices an increase in 5XXError count for a SageMaker endpoint. They want to set up automated remediation. Which THREE actions should they take? (Select THREE)

Medium
329

A data scientist needs to version and manage multiple models for a team of five. The team frequently experiments with different algorithms and hyperparameters. They need a centralized registry to store, deploy, and compare model versions. Which AWS service should the data scientist use?

Easy
330

A data scientist suspects that a deep learning model is overfitting. They enable SageMaker Debugger and want to detect overfitting automatically. Which built-in rule should they use?

Medium
331

A machine learning engineer is monitoring a deployed model for data drift. The input features are a mix of categorical and numerical columns. The baseline is from the training data. Which SageMaker Model Monitor feature should they enable to detect changes in the distribution of each feature over time?

Medium
332

A machine learning team deploys a model for loan approval. They want to monitor data drift on the real-time endpoint using SageMaker Model Monitor. Which set of actions should they take to set up data quality monitoring?

Medium
333

A company is using SageMaker endpoints for inference. To reduce costs, they want to use Automatic Scaling. However, they observe that scaling up takes several minutes, causing latency spikes during traffic bursts. What should they do to mitigate this?

Medium
334

Which SageMaker built-in algorithm is best suited for detecting anomalous login attempts based on IP addresses and user behavior?

Easy
335

A company needs to deploy a model that processes large payloads (up to 1 GB) asynchronously. The results should be written to S3, and the team needs SNS notifications upon completion. Which SageMaker inference option is MOST suitable?

Easy
336

A data engineer needs to ingest streaming data from IoT devices into Amazon S3 for machine learning. The data arrives continuously and must be available for querying within minutes. Which service should be used to collect and deliver the streaming data to S3?

Easy
337

A company uses Amazon SageMaker Feature Store to manage features for real-time inference. They need to store customer transaction data that is updated frequently and must be available for low-latency lookups. Which type of Feature Store should be used?

Medium
338

A company runs an online retail business and wants to build a product recommendation system. They have a dataset of customer purchases stored in Amazon S3 as CSV files. The dataset includes columns: 'customer_id', 'product_id', 'purchase_date', 'quantity', 'price', and 'category'. The data science team plans to use Amazon SageMaker to train a factorization machines model. During data exploration, they discover that the 'category' column has 1,200 unique values, and many categories appear only a few times. The 'product_id' column has 50,000 unique values. They want to include both features in the model. The team is concerned about the high cardinality of these features. Which approach should they take to prepare these features for the factorization machines model?

Medium
339

A company deploys a large NLP model on a SageMaker real-time endpoint using an ml.p3.2xlarge instance. To reduce inference cost without sacrificing throughput, they want to compile the model for their target hardware. Which service should they use?

Hard
340

A company ingests daily log data into an S3 bucket. They need to update the existing ML training dataset with new data without reprocessing the entire history. Which two strategies should they adopt? (Choose two.)

Easy
341

A healthcare company uses Amazon SageMaker to deploy a real-time inference endpoint for a diagnostic model. The endpoint is configured with a single ml.p3.2xlarge instance. The model processes patient data and returns a risk score. Recently, the endpoint has been experiencing intermittent 504 errors along with increased latency. The team uses Amazon CloudWatch to monitor the endpoint's InvocationsPerInstance and ModelLatency metrics. They observe that InvocationsPerInstance is well below the throttling threshold, but ModelLatency shows periodic spikes lasting 5-10 seconds. The endpoint's CPU utilization remains below 60%, but memory utilization occasionally spikes to 90% during those spikes. The team has checked the inference code and found no obvious memory leaks or performance bottlenecks in the custom logic. The model itself is a deep neural network hosted using Apache MXNet. The team suspects that the issue might be related to resource contention or an external dependency. What should the team do FIRST to diagnose and resolve the issue?

Hard
342

A machine learning engineer is preparing a dataset for training a SageMaker built-in Linear Learner model for binary classification. The dataset contains a highly imbalanced target with only 2% positive examples. They want to improve the model's ability to detect positives without collecting more data. Which SageMaker Linear Learner hyperparameter should they adjust to assign more weight to the positive class?

Hard
343

A machine learning model is deployed on SageMaker and its predictions are used in a production application. The model's accuracy has degraded over time. What is the most likely cause?

Medium
344

A data scientist is training a model on text data from customer reviews. The dataset contains a mix of English and Spanish reviews. The scientist wants to convert the text into numerical features for a classification model. Which approach is MOST appropriate for this multilingual dataset?

Medium
345

A fraud-detection team runs a SageMaker real-time endpoint in a production account. Their security team requires that the endpoint be reachable only from within a specific Amazon VPC and that access to invoke the endpoint be governed by identity-based policies with least privilege. Which TWO configurations should the ML engineer implement to meet these requirements? (Choose two.)

Hard
346

An ML engineer needs to monitor the operational health of a SageMaker endpoint, specifically the time taken for the container to process an inference request and the overhead added by SageMaker. Which two CloudWatch metrics should they examine?

Easy
347

Refer to the exhibit. A data scientist is trying to use AWS Glue to read data from the S3 bucket `ml-data-bucket`. The Glue job fails with an access denied error. What is the most likely cause?

Easy
348

A team is evaluating classification models for a medical diagnosis application. The cost of a false negative is much higher than the cost of a false positive. Which metric should be optimized during model selection?

Medium
349

A machine learning team trains a model in SageMaker and wants to track every step — from dataset version to hyperparameters to final model artifact — for reproducibility and audit compliance. Which SageMaker feature should they use?

Medium
350

A data scientist trained a model using SageMaker and wants to automate the retraining process when new data becomes available. Which AWS service is best suited to trigger a SageMaker training job based on an S3 event?

Easy
351

A data scientist is using SageMaker Experiments to track multiple training runs. They want to compare the F1 scores across runs. Which component should they use to log the F1 score?

Medium
352

A company is fine-tuning a large language model using reinforcement learning from human feedback (RLHF). Which THREE components are typically required?

Hard
353

Refer to the exhibit. A SageMaker training job logs show training AUC increasing but validation AUC plateauing at 0.880. What is the most likely issue?

Hard
354

A company deploys a deep learning model to a real-time SageMaker endpoint. After deployment, users report high inference latency. Which action is the MOST effective first step to reduce latency?

Easy
355

A company is training a binary classifier in SageMaker and observes that the training loss decreases but validation loss increases after a few epochs. What is the most likely issue?

Easy
356

Which feature scaling method is most robust to outliers in the data?

Medium
357

A company has a dataset of 2 billion records stored as text files in Amazon S3. The data is partitioned by year and month. The data science team wants to read only the last 6 months of data for model training using SageMaker. To minimize data scanned and reduce costs, which approach should the team use?

Easy
358

A team monitors a production endpoint and notices a sudden increase in 5XXError count. Which of the following is the most likely cause?

Medium
359

A data scientist is performing feature selection for a linear regression model and wants to remove features that are highly correlated with each other to reduce multicollinearity. Which technique is BEST suited for this purpose?

Easy
360

A team is using SageMaker Pipelines to train a model. The pipeline has multiple steps: data processing, training, evaluation, and registration. They use a Condition step to evaluate the model's accuracy and if it exceeds a threshold, register the model. They run the pipeline and the training step succeeds, but the pipeline fails at the Condition step with an error: 'Unable to evaluate condition: the property 'Accuracy' does not exist.' The evaluation step output is a JSON file with key 'accuracy'. What is the most likely cause?

Hard
361

A company is building a real-time inference pipeline for an ML model. The raw data arrives in JSON format via Amazon Kinesis Data Streams. Before invoking the SageMaker endpoint, the data must be preprocessed to match the training data format. Which THREE steps should be included in the preprocessing function? (Select THREE)

Hard
362

A team uses SageMaker Clarify to monitor bias drift in production. They schedule weekly analysis. After a month, Clarify reports a significant increase in a bias metric. What should the team do first?

Medium
363

A financial services firm is training a fraud detection model using SageMaker. The dataset is highly imbalanced (0.1% fraudulent transactions). The model currently achieves 99.9% accuracy but only catches 5% of fraud cases. Which metric should the team prioritize to evaluate model performance?

Hard
364

A data engineer is ingesting streaming clickstream data from a website into Amazon S3 for ML training. The data arrives at a rate of 10,000 events per second, and the team needs near-real-time availability with minimal transformation. Which AWS service should the engineer use to ingest the data into S3 with the LEAST operational overhead?

Hard
365

A company has deployed a machine learning model on Amazon SageMaker and wants to automatically detect when the distribution of input features deviates significantly from the training data distribution. Which SageMaker feature should they use?

Easy
366

A company wants to use SageMaker Clarify to analyze bias in their training data and model predictions. Which TWO types of bias can Clarify detect? (Choose TWO.)

Easy
367

A data science team is using Amazon SageMaker to train and deploy a binary classification model. They want to continuously monitor the model for data drift in production. Which combination of AWS services and SageMaker features should they use to implement automated drift detection with minimal operational overhead?

Medium
368

A data scientist is using Amazon SageMaker Data Wrangler to prepare a dataset. The dataset contains a column with date strings in the format 'YYYY-MM-DD'. The data scientist wants to extract the year, month, and day as separate features. Which Data Wrangler transform should be used?

Medium
369

A machine learning engineer is using Amazon SageMaker Data Wrangler to prepare a dataset for a regression model. The dataset contains a categorical feature with high cardinality (over 10,000 unique values). The engineer wants to encode this feature efficiently without creating thousands of binary columns, which would explode the feature space. Which encoding technique should the engineer use in Data Wrangler?

Hard
370

A company deploys a model for fraud detection. They need to monitor for bias after deployment, specifically whether the model's false positive rate changes across demographic groups over time. Which SageMaker feature should they use?

Hard
371

A data scientist is using SageMaker Automatic Model Tuning to optimize hyperparameters for an XGBoost model. They want to maximize AUC. Which search strategy is MOST appropriate for efficient exploration?

Medium
372

A machine learning engineer is preparing data for a SageMaker training job and needs to split a large dataset into training, validation, and test sets while avoiding data leakage from the same entity appearing in multiple splits. The dataset contains multiple rows per customer, and the target is customer churn. Which TWO strategies are appropriate? (Choose two.)

Hard
373

A company is building a fraud detection model on credit card transactions. The dataset contains a column 'merchant_id' with 50,000 unique values, many with low frequency. The team wants to avoid overfitting while preserving predictive signal. Which feature engineering approach is most appropriate?

Hard
374

A company is building a recommender system using implicit feedback (clicks) and explicit feedback (ratings). They plan to use Amazon SageMaker to train a model. The data includes user ID, item ID, timestamp, and rating (if any). Which TWO data preparation steps should the team perform? (Choose TWO.)

Medium
375

A machine learning engineer is preparing a dataset for training a model on Amazon SageMaker. The dataset contains numerical features with varying scales, and the engineer wants to ensure that all features contribute equally during training. Which data preparation step should the engineer take?

Easy
376

Which TWO actions are recommended best practices for securing an Amazon SageMaker notebook instance? (Select TWO.)

Easy
377

A company is training a large NLP model on SageMaker and wants to reduce costs by using Spot Instances. Which TWO configurations should they implement to handle Spot interruptions gracefully?

Medium
378

A company wants to deploy a new model using a canary deployment strategy on SageMaker. Which two actions should they take? (Select TWO.)

Medium
379

A team has a SageMaker Pipeline that trains a model and registers it in the Model Registry. They want to automate the deployment of the approved model to a staging environment. Which event-driven approach should they use?

Medium
380

A company is using SageMaker Automatic Model Tuning to optimize a regression model. They want to minimize the root mean squared error (RMSE). The tuner has completed 20 jobs, and the RMSE has plateaued. Which action should the data scientist take to potentially improve the results?

Medium
381

A data scientist wants to fine-tune a Llama 2 7B model using SageMaker for a text summarization task. The dataset is 10 GB. The budget is limited, so cost efficiency is important. Which THREE steps should the data scientist take? (Choose THREE.)

Medium
382

A machine learning engineer is using SageMaker to train a model and wants to automatically stop a training job when the validation loss has not improved for 10 consecutive epochs, while still saving the best model artifacts. The engineer is using the SageMaker training toolkit in a custom container. Which combination of actions should the engineer take?

Medium
383

A company wants to reduce costs for a real-time inference endpoint that experiences predictable traffic spikes during business hours and low traffic at night. Which auto-scaling policy is MOST cost-effective while maintaining performance?

Easy
384

A healthcare analytics team trains models in SageMaker and stores artifacts in an S3 bucket that contains protected health information. An auditor asks how the team can prove which training dataset and container image produced the model currently deployed to production, and wants the evidence retained even if someone deletes the training job. Which SageMaker capability should the team rely on to capture and retain this metadata automatically?

Medium
385

A data scientist is using SageMaker Experiments to track multiple training runs for a PyTorch model. They want to compare metrics across runs and identify the best hyperparameters. Which TWO capabilities should they use? (Choose TWO.)

Medium
386

A data scientist is preparing a dataset for a multi-class classification problem. The dataset contains a categorical feature with 50,000 unique values (high cardinality). The scientist wants to reduce dimensionality while preserving predictive information. Which TWO approaches are appropriate? (Choose 2)

Hard
387

A team wants to automatically retrain a model when new labeled data arrives. Which SageMaker feature can orchestrate this workflow?

Easy
388

A team uses SageMaker ML Lineage Tracking to capture the metadata of their ML workflow. They want to query the lineage to see which model version was trained from a specific dataset. Which Lineage Tracking entity represents the dataset?

Medium
389

A data scientist wants to normalize a feature to have a range between 0 and 1 for a neural network. Which scaling technique should be applied?

Easy
390

A machine learning engineer is preparing a dataset for a binary classification model. The dataset has 10,000 samples with a 1:100 class imbalance. The engineer needs to balance the classes before training. Which technique would create a balanced dataset without discarding majority class samples and without generating synthetic data?

Medium
391

A data scientist is building a time-series forecasting model for daily sales data. The data spans two years. To evaluate the model's performance, the data scientist needs to simulate a realistic rolling forecast scenario. Which data splitting strategy should be used?

Hard
392

A machine learning engineer needs to standardize features to have zero mean and unit variance before training a support vector machine. Which scaling method should they apply?

Easy
393

A company has a SageMaker endpoint that uses a trained model to classify images. The endpoint is experiencing high latency and the team suspects it is due to the model size. Which action can the team take to reduce latency without significantly impacting accuracy?

Easy
394

A hospital deploys a model to predict patient readmission risk. To comply with regulations, they must ensure that the model's predictions do not show bias against any demographic group over time. Which service should they use for ongoing monitoring?

Hard
395

A data scientist is using SageMaker Data Wrangler to prepare a large dataset. The data contains duplicate rows, which could bias the model. Which built-in step in Data Wrangler can automatically detect and remove duplicates?

Medium
396

A financial services company is developing a fraud detection model using Amazon SageMaker. They have a dataset with 10 million transactions, each with 300 features. The dataset is highly imbalanced (0.1% fraud). They have performed feature engineering and now need to split the data for training, validation, and test sets. The data is stored in CSV files in Amazon S3. They plan to use SageMaker's built-in XGBoost algorithm. To ensure proper evaluation and avoid data leakage, which data splitting strategy should they use?

Hard
397

A company runs a regression model to predict house prices. They have 50 features including 'zip_code' (high cardinality), 'square_footage', and 'year_built'. They want to select the most important features to reduce overfitting. Which feature selection method is computationally efficient for high-dimensional data and can handle multicollinearity?

Hard
398

A team uses SageMaker Ground Truth to create labeled datasets. They need to ensure labeling jobs are cost-effective. Which TWO measures should they take? (Select TWO.)

Easy
399

A machine learning team is building a model to predict customer churn. They have historical data that includes customer activity logs, each with a timestamp. The team wants to ensure that the training data does not contain any data leakage from the future. Which approach should they take when preparing the training and validation datasets?

Hard
400

A company stores its raw IoT sensor data in Amazon S3. The data is in CSV format and contains timestamps, sensor IDs, and readings. A data engineer needs to catalog this data for discoverability and querying by other team members. Which AWS service should they use to create a searchable metadata catalog?

Easy
401

A company is using SageMaker to train a neural network for image classification. The training job is taking too long. The team wants to reduce training time without sacrificing model accuracy. Which approach should they recommend?

Medium
402

A model deployed on SageMaker is returning inaccurate predictions for certain customer segments. The team suspects data drift. Which SageMaker feature should they use to continuously monitor input data distribution?

Hard
403

A team wants to apply a custom container for inference on SageMaker. The container needs to implement a web server that responds to API requests. Which protocol and port must the container listen on to be compatible with SageMaker hosting?

Easy
404

A startup wants to deploy a containerized ML application that includes both a model inference server and a preprocessing component in the same endpoint. Which SageMaker endpoint type supports running multiple containers?

Medium
405

A company uses Amazon SageMaker to deploy a real-time inference endpoint. They notice increased latency in predictions during peak hours. Which should they investigate first to address the issue?

Easy
406

A company uses SageMaker to train a model. They want to ensure that training data is encrypted at rest and in transit, and that only authorized users can access the training artifacts. Which three steps should they take? (Choose three.)

Hard
407

A retail company is building a machine learning model to predict customer churn. The data engineering team has extracted customer transaction data from Amazon Aurora and stored it as CSV files in Amazon S3. The data includes customer IDs, transaction amounts, timestamps, and product categories. A data scientist discovers that the dataset contains several missing values in the 'transaction_amount' column for about 15% of the records. The data scientist also notices that the 'customer_id' column has some duplicate entries. The team wants to prepare the data for training a churn model using Amazon SageMaker. The data is approximately 50 GB in size. What should the data scientist do to handle the missing values and duplicates efficiently while preparing the data for training?

Easy
408

A data scientist is preparing a large dataset for training a machine learning model. The dataset contains missing values in several columns. Which approach is the MOST efficient for handling missing values in a large dataset using AWS services?

Easy
409

A data engineer needs to ingest streaming clickstream data from a website into an S3 data lake for ML training. The data arrives continuously and must be written to S3 in near real-time. Which AWS service is best suited for this task?

Easy
410

A machine learning team deploys a custom container image for an Amazon SageMaker training job. The container needs to access an S3 bucket that contains sensitive data. The team wants to follow the principle of least privilege. How should the team grant access?

Medium
411

A company wants to update an existing SageMaker real-time endpoint to serve a new model version. They need to route a small percentage of traffic to the new version initially and monitor for errors before switching fully. Which deployment pattern supports this?

Easy
412

During model training on Amazon SageMaker, the training job fails with a 'ResourceLimitExceeded' error. What is the most likely cause?

Medium
413

A team wants to track and compare multiple machine learning experiments, including hyperparameters, metrics, and artifacts. They are using Amazon SageMaker. Which AWS service or feature should they use to achieve this?

Easy
414

An organization wants to ensure that only approved model versions can be deployed to production. They use the SageMaker Model Registry to track model versions. How can they enforce that only approved models are deployed?

Medium
415

A machine learning engineer is training a SageMaker job with the TensorFlow estimator and wants to automatically capture model training metadata such as loss curves and accuracy for later comparison, without writing any custom code. Which SageMaker feature should they enable?

Medium
416

A company uses Amazon SageMaker to deploy a model for real-time inference. They want to perform A/B testing between two model versions. Which TWO actions should the company take to set up A/B testing? (Choose TWO.)

Medium
417

A machine learning engineer trains a binary classifier in SageMaker and the model outputs class probabilities. The business requires that the model achieve at least 90% recall on the positive class, while keeping precision above 70%. The engineer uses the default threshold of 0.5 when deploying. Which approach should the engineer take to meet these requirements?

Medium
418

A company wants to maintain multiple versions of a trained model in a central repository and track metadata such as training metrics, hyperparameters, and approval status. Which SageMaker feature should they use?

Easy
419

Refer to the exhibit. A data scientist configured SageMaker Debugger to monitor training for overfitting. However, the rule never triggers even though the model appears to be overfitting. What is the most likely reason?

Hard
420

A financial services company uses Amazon SageMaker to deploy a fraud detection model for real-time inference. The model is deployed on an ml.m5.large instance with a SageMaker real-time endpoint. The endpoint has an auto scaling policy configured using a custom scaling policy based on average CPU utilization, with scale out threshold at 70% and scale in threshold at 30%. During a flash sale event, the traffic to the endpoint spikes tenfold within minutes. The endpoint fails to handle the load, resulting in increased latency and timeouts. The data science team needs to improve the scalability of the endpoint to handle sudden traffic spikes. Which solution should the team implement?

Hard
421

A company is building a real-time fraud detection system using Amazon Kinesis Data Streams. The data must be joined with a reference table (e.g., customer profile) that is stored in Amazon DynamoDB and updated frequently. The enriched data will be used for ML predictions. Which THREE AWS services should the company use to build this streaming pipeline? (Select THREE.)

Medium
422

A company is using Amazon SageMaker to train a model on sensitive customer data. The security team requires that all data be encrypted in transit and at rest, and that the training job does not have internet access. Which configuration should the team use to meet these requirements?

Easy
423

A machine learning engineer wants to deploy a pre-trained foundation model for text summarization using SageMaker JumpStart. Which of the following is a primary cost consideration when deploying such a model?

Easy
424

A data engineer is building a feature store using Amazon SageMaker Feature Store. The team needs to store features that are updated frequently and require low-latency retrieval for real-time inference. Which type of store should the engineer use?

Easy
425

A healthcare company deploys a model to predict patient readmission risk. The model was trained on historical data and is now showing signs of concept drift. The team needs to implement a monitoring solution that can detect drift and automatically retrain the model when drift is detected. Which THREE steps should the team take to build this solution? (Choose THREE.)

Hard
426

A team wants to use SageMaker Clarify to monitor bias in their production model predictions. They have configured a bias drift monitor. What does SageMaker Clarify compare to detect bias drift?

Medium
427

A data scientist is using SageMaker built-in XGBoost algorithm for a regression problem. Which metric is most appropriate as the objective metric for hyperparameter tuning?

Easy
428

An ML engineer needs to split a dataset into training, validation, and test sets. The dataset has a time-based column that should not be leaked. Which split method is most appropriate?

Easy
429

A data engineer is designing an ETL pipeline using AWS Glue to transform raw data from S3 into a curated set for ML training. The data contains personally identifiable information (PII) that must be masked before being used by data scientists. Which TWO actions should the engineer take? (Choose TWO.)

Medium
430

A data scientist is training a binary classification model using Amazon SageMaker. The dataset has a severe class imbalance (95% negative, 5% positive). The model achieves 99% accuracy but fails to identify positive cases correctly. Which action should the data scientist take to improve the model's ability to detect positive cases?

Medium
431

A company deploys a model using SageMaker real-time endpoint with auto scaling. They observe that during a traffic spike, the endpoint quickly scales up to 10 instances, but after the spike, it takes a long time to scale down, leading to high costs. The scaling policy is based on a simple average CPU utilization threshold. Which adjustment would optimize the scaling down behavior?

Hard
432

An MLOps engineer is setting up a SageMaker endpoint for a model that performs inference on large images. The model is containerized and expects input in a specific format. The team wants to preprocess the images (resize and normalize) before passing them to the model. What is the most efficient way to implement this?

Medium
433

A data engineer is using AWS Glue to run an ETL job that joins two large datasets and writes the output to S3 for ML training. The job is failing due to out-of-memory errors. Which THREE actions can help resolve this issue? (Select THREE.)

Hard
434

A healthcare analytics team trains models in Amazon SageMaker and needs an immutable, queryable record of which dataset version and training job produced each registered model version, so an auditor can trace a deployed model back to its inputs months later. Which SageMaker capability should they rely on?

Easy
435

A machine learning engineer is preparing a dataset for training a SageMaker model. The dataset contains missing values in several numerical features. The engineer wants to handle these missing values during the training pipeline. Which two methods are valid ways to handle missing values in SageMaker? (Choose two.)

Medium
436

A company is using SageMaker to train a model for image classification. The training dataset contains 100,000 labeled images. The team wants to use a pre-trained model to reduce training time. Which SageMaker feature should they use?

Medium
437

A company uses SageMaker Neo to compile a trained model for deployment on edge devices. What is the primary benefit of using Neo?

Easy
438

A company wants to track the lineage of their ML models for reproducibility and auditability. Which THREE services or features should they use together to achieve this? (Choose THREE.)

Medium
439

A data science team needs to deploy a trained PyTorch model for real-time inference with sub-100ms latency. The model fits on a single GPU. Which SageMaker inference option is MOST cost-effective while meeting the latency requirement?

Easy
440

A company wants to reduce costs for a SageMaker real-time endpoint that has variable traffic. Which feature allows the endpoint to automatically adjust instance count based on demand?

Easy
441

A financial services company has a SageMaker real-time endpoint serving a fraud detection model. Compliance requires that all inference requests and responses be logged with the ability to detect anomalous input feature distributions over time. The team wants a managed solution that captures request/response payloads to Amazon S3 and automatically computes statistics and constraints against a baseline. Which combination of SageMaker features should they enable?

Medium
442

An ML engineer is fine-tuning a foundation model using RLHF on SageMaker. Which THREE components are essential for this workflow? (Select THREE.)

Hard
443

A data scientist is training a deep learning model on Amazon SageMaker and notices that the training loss decreases but the validation loss starts increasing after a certain number of epochs. The model is likely overfitting. Which SageMaker feature can they use to detect and diagnose this issue during training?

Medium
444

A company is building a fraud detection model on an imbalanced dataset (99% legitimate, 1% fraudulent). To improve recall on the minority class, they want to resample data. Which combination of techniques should they use?

Medium
445

A company wants to deploy a scikit-learn model to a SageMaker AI real-time endpoint. The model must be loaded from a custom Python module that contains preprocessing logic not present in the built-in scikit-learn container. The team wants to minimize operational overhead and does not need to change system-level libraries. Which approach should the engineer take?

Medium
446

An ML engineer has a real-time SageMaker endpoint serving a fraud-detection model. The team wants to release a new model version to a small percentage of live traffic first, monitor CloudWatch metrics for accuracy regressions, and roll back quickly if performance degrades. They also want the production and candidate variants to share the same endpoint so latency comparisons are apples-to-apples. Which SageMaker deployment strategy should they use?

Medium
447

A team wants to fine-tune a pre-trained Hugging Face transformer model for text classification using SageMaker. They have a custom training script. Which SageMaker estimator should they use?

Easy
448

A financial services company uses SageMaker to train a fraud detection model. They have imbalanced data with 1% fraud. They trained a Gradient Boosting model using SMOTE for oversampling and achieved 99% accuracy on the test set, but the fraud recall is only 10%. The data scientist is concerned about the model's performance. Which change is most likely to improve fraud recall without sacrificing too much precision?

Medium
449

A machine learning team has a model that needs to serve predictions with very low latency (under 10 ms) for a real-time web application. The model is a small ensemble of three neural networks that fits in memory. Which SageMaker inference option is MOST appropriate?

Medium
450

A company deployed a machine learning model on an Amazon SageMaker real-time endpoint. Over several weeks, they notice that inference latency has been gradually increasing, especially during peak business hours. The model and instance type have remained unchanged. What is the most likely cause of the increased latency?

Easy
451

A data science team deploys a regression model to Amazon SageMaker for real-time inference. After one month, the model's prediction errors increase significantly, but data distributions remain unchanged. Which monitoring approach is MOST suitable for detecting this issue?

Easy
452

A team needs to deploy a new model version to production while minimizing risk. They want to route 5% of live traffic to the new model and 95% to the current model, and then gradually increase the new model's traffic. Which SageMaker deployment pattern should they use?

Medium
453

A machine learning engineer is setting up an Amazon SageMaker notebook instance. The instance needs to access a private S3 bucket that contains training data. The notebook instance is in a VPC. Which combination of steps will grant access to the S3 bucket? (Choose TWO.)

Easy
454

An ML platform team must orchestrate a workflow that trains a model, evaluates it against a baseline, and only registers the model if evaluation passes. If evaluation fails, the workflow must notify the data science channel and stop without registering. The team wants the orchestration logic to be expressed as code, versioned in git, and integrated with SageMaker training jobs and Model Registry. Which approach best fits these requirements?

Hard
455

A company uses Amazon SageMaker Ground Truth to label images for object detection. They want to minimize labeling costs while maintaining high accuracy. Which feature should they enable?

Medium
456

A trained model needs to be deployed for real-time inference with low latency. Which AWS service is best suited for this?

Medium
457

A machine learning engineer is preparing a training dataset stored in Amazon S3 for a SageMaker training job. The data is in CSV format, and the engineer wants to ensure that the training job can access the data efficiently and securely. The S3 bucket is in the same AWS Region as the SageMaker training job. The engineer needs to provide the training job with the necessary permissions to read the data. Which of the following is the MOST secure and appropriate way to grant the training job access to the S3 bucket?

Medium
458

A data science team is using Amazon SageMaker Pipelines to orchestrate a multi-step workflow that includes data preprocessing, training, and model evaluation. They want to reuse the preprocessed data across multiple pipeline executions without re-running the preprocessing step if the source data hasn't changed. What should they configure?

Hard
459

A data scientist is training a large model on SageMaker and wants to reduce training time by using multiple GPUs. The model is small enough to fit on a single GPU but training is slow. Which SageMaker feature should be used?

Medium
460

A data scientist is using SageMaker built-in XGBoost algorithm for a binary classification task. Which objective metric is MOST appropriate for SageMaker Automatic Model Tuning to maximize?

Easy
461

A machine learning engineer is training a deep learning model on SageMaker and notices that the training loss decreases rapidly in the first few epochs but then plateaus. The validation loss starts increasing after 10 epochs. Which action should the engineer take to improve generalization?

Medium
462

A machine learning engineer manages a SageMaker Model Monitor schedule for a real-time endpoint. The monitor's baseline was computed from a training dataset with a categorical feature named region. In production, a new category value appears that was never seen in training, and the monitor begins reporting violations. The engineer wants the monitor to flag only the appearance of unknown categories without treating normal distribution shifts in known categories as violations. Which approach should the engineer take?

Hard
463

An ML engineer needs to compile a trained TensorFlow model to run efficiently on a target edge device with an ARM CPU. Which AWS service should they use?

Easy
464

A data scientist wants to track the training and validation accuracy of a SageMaker training job over time. They need to visualize these metrics in Amazon CloudWatch. Which action should they take?

Easy
465

A company wants to monitor its Amazon SageMaker real-time endpoint for data quality issues. Which TWO actions should the company take?

Easy
466

A company uses SageMaker JumpStart to deploy a foundation model for a summarization task. They want to minimize costs while still meeting a latency requirement of under 2 seconds. Which option should they consider?

Medium
467

A financial services company ingests transaction data from multiple sources into an S3 data lake. They want to use AWS Glue to catalog this data and make it queryable by Amazon Athena. The data schema changes frequently as new sources are added. Which AWS Glue feature should they enable to automatically detect and update the schema?

Medium
468

An ML team uses SageMaker to deploy a model for real-time inference. They want to monitor and improve cost efficiency. Which THREE actions should they take? (Select THREE.)

Hard
469

A data science team deploys a machine learning model to a SageMaker endpoint for real-time inference. They need to monitor the model for feature distribution drift over time to ensure the model's predictions remain accurate. Which AWS service should they use?

Easy
470

A machine learning engineer needs to split a dataset for binary classification where the positive class represents only 2% of the data. Which data splitting strategy ensures that both training and test sets maintain the same class proportion as the original dataset?

Easy
471

A data scientist is training a binary classification model using Amazon SageMaker. The dataset is highly imbalanced (95% negative class, 5% positive class). The model is evaluated on a held-out test set, and the F1 score is 0.12. The data scientist wants to improve the F1 score. Which two actions should the data scientist take? (Choose two.)

Medium
472

A company runs a real-time fraud detection model on a SageMaker endpoint. The model is updated weekly, and each update must be validated against live traffic without affecting existing predictions. The team wants to compare the new model's performance against the current model using a small percentage of incoming requests, while ensuring that the current model continues to serve the majority of traffic. Which SageMaker deployment strategy should they use?

Medium
473

A hospital's ML team deploys a diagnostic model to a SageMaker real-time endpoint. Compliance requires that every inference request and response be recorded for later auditing, and the records must be retrievable months later. The team needs a low-effort way to capture this data. Which approach should they use?

Easy
474

A company is using SageMaker Pipelines to automate a multi-step ML workflow. The pipeline includes data preprocessing, training, and model evaluation. The team wants to ensure that if the evaluation step fails, the pipeline stops and sends an alert to the operations team. Which SageMaker Pipelines feature should they use?

Easy
475

Which THREE steps should be taken to optimize a large-scale distributed training job on SageMaker? (Choose 3.)

Hard
476

A data science team uses SageMaker Pipelines to orchestrate their ML workflow. They noticed that even when source data hasn't changed, the pipeline re-runs all steps, wasting compute time. What should they enable to avoid redundant runs?

Hard
477

A company is building a time series forecasting model using SageMaker DeepAR. The raw data is a CSV with columns: timestamp, item_id, and value. What is the correct data format required for DeepAR training?

Medium
478

A data scientist is preparing a dataset for a machine learning model that predicts customer churn. The dataset contains a column 'CustomerID' that is a unique identifier. What should the data scientist do with this column before training the model?

Easy
479

A team uses SageMaker Pipelines to automate retraining. They want to skip the training step if the data has not changed since the last run. Which feature should they enable?

Medium
480

A data scientist is using SageMaker Experiments to track multiple training runs. They want to compare runs based on the objective metric and visualize performance. Which THREE steps should they perform? (Choose THREE.)

Hard
481

A company wants to deploy a machine learning model that was trained on-premises using TensorFlow. The model is a TensorFlow SavedModel. The company uses AWS and wants to minimize operational overhead. Which deployment option meets these requirements?

Easy
482

A machine learning engineer deploys a new model version to a SageMaker endpoint with production variants. They want to gradually shift traffic from the old model to the new model, monitoring for errors, and automatically roll back if the error rate exceeds 5%. Which deployment pattern should they use?

Hard
483

A data scientist wants to quickly build a binary classification model without writing any code. Which SageMaker feature is MOST suitable?

Easy
484

A data scientist needs to ingest streaming clickstream data from a website into an S3 data lake for ML training. The data must be processed in near real-time and partitioned by hour. Which AWS service combination should be used?

Medium
485

A machine learning engineer needs to split a dataset into training, validation, and test sets for a SageMaker training job. The dataset is stored in Amazon S3 as a single CSV file. The engineer wants to ensure that the splits are reproducible and that the test set is never used during training or hyperparameter tuning. Which approach should the engineer use?

Easy
486

A machine learning team stores training data in an Amazon S3 bucket and wants to catalog it so that Amazon Athena and Amazon SageMaker Feature Store can discover the schema. The data is partitioned by year, month, and day in Hive-style prefixes, and new partitions are added daily. A data engineer must ensure new partitions are automatically discoverable without manual intervention. Which solution meets these requirements?

Medium
487

A company has 200 small PyTorch models that are each used infrequently but need to be available for real-time inference. To minimize costs, they want to host all models on a single endpoint. Which SageMaker feature should they use?

Medium
488

An organization wants to automate ML retraining using an event-driven architecture. Which THREE services should they combine? (Select THREE.)

Medium
489

A machine learning engineer is building a real-time fraud detection pipeline using Amazon Kinesis Data Streams. The data must be prepared (e.g., feature engineering, normalization) before being fed into a SageMaker endpoint. Which TWO steps should the engineer implement to ensure low-latency data preparation?

Medium
490

A company wants to deploy a trained XGBoost model for batch inference on a large dataset stored in S3. The inference job should be cost-effective and does not require real-time responses. Which SageMaker inference option should they use?

Easy
491

A machine learning engineer needs to deploy a new version of a model gradually, initially sending 5% of traffic to the new version and 95% to the current version, while monitoring for errors. Which deployment pattern should they use?

Easy
492

A media company uses SageMaker endpoints to serve a model that predicts video engagement. They have two production variants: Variant A (ml.c5.large) for regular traffic and Variant B (ml.c5.xlarge) for burst traffic. They use weighted routing (90% to A, 10% to B). Recently, during peak hours, Variant A's latency increase causes many requests to time out. The metrics show that both variants are under similar CPU load, but the number of concurrent requests to Variant A is very high. The team wants to ensure that burst traffic is handled properly without manual intervention. What should they do?

Medium
493

A retail company runs a SageMaker real-time endpoint serving a demand forecasting model. Security policy requires that all inference requests travel over the AWS private network and never traverse the public internet, and that the endpoint cannot be invoked from outside the company VPC. The endpoint already uses a customer-managed KMS key for volume encryption. Which TWO configurations should the engineer apply to meet these requirements? (Choose two.)

Hard
494

A company deploys a SageMaker model using AWS KMS for encryption at rest. They have a compliance requirement to rotate the KMS key every year without causing downtime for the inference endpoint. Which approach should they take?

Hard
495

Which TWO data storage options are commonly used by Amazon SageMaker Feature Store for offline and online storage?

Easy
496

A data scientist is trying to create a SageMaker endpoint configuration with 6 instances of ml.c5.large for a production variant. The creation fails with the error shown in the exhibit. Which action should the data scientist take to resolve this issue?

Hard
497

A data scientist is training a model with SageMaker and needs to reduce the cost of a long-running training job that can tolerate interruptions. The job uses a custom training script and reads data from Amazon S3. The data scientist wants the job to resume from the last saved state if the underlying compute is reclaimed. (Choose two.)

Hard
498

A machine learning engineer observes that model performance on a SageMaker endpoint has degraded over the past week. Ground truth labels are available with a 2-day delay. The engineer wants to automatically trigger a retraining pipeline when prediction quality drops below an acceptable threshold. Which approach is most appropriate?

Medium
499

A team wants to monitor the number of requests and latency of their SageMaker endpoint using a unified dashboard. Which AWS service should they use to create a custom dashboard with these metrics?

Easy
500

A team uses SageMaker real-time endpoints for inference. They want to deploy a new model version and compare its performance with the current version under live traffic without affecting user experience. Which method should they use?

Hard
501

A company is deploying a large language model (LLM) to a SageMaker endpoint. They want to minimize inference latency and cost by using GPU acceleration and model parallelism. The model is too large to fit on a single GPU. Which SageMaker feature should they use?

Hard
502

A data engineer stores raw ML training data in Amazon S3 and needs to catalog the schema, track partition changes, and make the data queryable by Amazon Athena without running ETL. Which AWS service should the engineer use?

Easy
503

A data engineer needs to convert a JSON dataset to Parquet format for efficient querying with Amazon Athena. The JSON files are in an S3 bucket. Which service can perform this conversion with minimal coding?

Easy
504

A company wants to automatically trigger a retraining pipeline when concept drift is detected in their deployed model. Which combination of services should they use?

Medium
505

A company uses SageMaker endpoints with auto-scaling. The endpoint is experiencing high latency during peak hours. The metrics show CPU utilization is low but memory is high. What is the most likely cause?

Medium
506

Which TWO actions are recommended best practices when preparing training data for a machine learning model in AWS? (Choose two.)

Easy
507

A team is using SageMaker Pipelines to automate retraining and deployment. They want to trigger the pipeline automatically when new training data is available in an S3 bucket. Which approach should they use?

Medium
508

A machine learning team runs a SageMaker AI Pipeline that trains a model and registers it in the SageMaker AI Model Registry. A separate deployment process must promote the model to production only after a human reviewer approves the model version. The team wants to automate the promotion so that approval in the Model Registry triggers the deployment without manual intervention. Which combination of steps should the engineer implement?

Hard
509

Refer to the exhibit. A user is unable to invoke a SageMaker endpoint. The IAM policy shown is attached to the user. Which permission is missing to allow invocation?

Easy
510

A company uses Amazon SageMaker Feature Store to store features for a real-time recommendation model. The feature data is updated continuously, and the model must use the most recent feature values for each user at inference time. Which type of Feature Store should the company use for serving features to the model?

Hard
511

A machine learning team needs to monitor a deployed model for both data drift and concept drift. Which TWO approaches should they implement? (Select TWO.)

Medium
512

A machine learning engineer is deploying a model to a SageMaker endpoint that must handle occasional large payloads up to 1 GB. The inference time can take up to 10 minutes. The team wants to minimize cost and avoid idle compute. Which deployment option is most appropriate?

Medium
513

An ML engineer runs the CLI command shown in the exhibit. However, the training job fails immediately with an error: 'Unable to assume role'. What is the most likely cause?

Easy
514

A company is using Amazon SageMaker to deploy a model for real-time inference. The model requires access to a private S3 bucket that contains reference data. The company wants to ensure that the endpoint can access the S3 bucket without using a public internet connection. Which TWO actions should they take? (Select TWO.)

Easy
515

A data engineer is using Amazon SageMaker Ground Truth to create a labeled dataset for an object detection task. The dataset contains millions of images, and the labeling budget is limited. Which approach can reduce labeling costs while maintaining high model accuracy?

Medium
516

A company wants to deploy a PyTorch model on SageMaker using the NVIDIA Triton Inference Server for GPU acceleration. They have an existing Triton configuration. Which approach should they take?

Medium
517

A team uses AWS Auto Scaling for a SageMaker real-time endpoint. They notice that when scaling in, the latest instance is always terminated first, causing disruption to recent requests. How can they configure the scaling policy to terminate the oldest instance first?

Medium
518

A team is using Amazon SageMaker to train a neural network. They want to minimize training time while effectively exploring the hyperparameter space. Which approach should they use?

Medium
519

A company needs to update a model in production without any downtime. They currently have a single real-time endpoint serving traffic. Which approach allows them to deploy a new model version and switch traffic gradually while being able to roll back quickly?

Hard
520

An ML engineer has trained a model and stored the model artifacts in an Amazon S3 bucket in the same AWS Region as the planned SageMaker AI endpoint. During endpoint creation, the engineer must specify the S3 location of the model artifacts. Which permission must the SageMaker AI execution role have for the endpoint to load the model successfully?

Easy
521

An ML team wants to use Amazon SageMaker Ground Truth to create a labeled dataset for a multi-class image classification task. They have a large set of unlabeled images and want to minimize labeling costs while maintaining high accuracy. Which Ground Truth feature should they enable?

Easy
522

A data scientist deploys a model and wants to monitor the endpoint's invocation latency. They notice that the CloudWatch metric 'ModelLatency' is high, but 'OverheadLatency' is low. Which statement correctly interprets these metrics?

Medium
523

A data engineer is preparing a large dataset of 10 TB for ML training on Amazon SageMaker. The data is stored in Amazon S3 as CSV files. To reduce training time and cost, the engineer wants to use a columnar format that is optimized for analytical queries. Which format should the engineer convert the data to?

Easy
524

A small team runs a SageMaker real-time endpoint in production. They want a low-effort way to know when the endpoint's invocations are failing so they can react quickly, and they want the alert delivered to their on-call channel. Which approach requires the least custom code?

Easy
525

A data scientist is using SageMaker to train a linear regression model. After training, they evaluate the model on the test set and get an R² of 0.95. However, when they deploy the model to a SageMaker endpoint and run predictions on new data, the predictions are far off. What is the most likely cause?

Easy
526

A machine learning engineer notices that the latency of a SageMaker endpoint has increased over time. They need to identify which component (model inference vs. pre/post-processing) contributes most to the latency. Which CloudWatch metrics should they examine?

Medium
527

A healthcare startup has deployed a machine learning model on Amazon SageMaker that predicts patient readmission risks. The model uses sensitive health data stored in an S3 bucket encrypted with AWS KMS. The SageMaker endpoint is configured with an IAM role that has the following policy attached: { "Version": "2012-10-17", "Statement": [ { "Effect": "Allow", "Action": "s3:*", "Resource": "arn:aws:s3:::healthcare-data/*", "Condition": { "Bool": { "aws:SecureTransport": "true" } } }, { "Effect": "Allow", "Action": "kms:Decrypt", "Resource": "*" } ] }. During a security audit, the team discovers that the IAM role's KMS permission is too permissive because it allows decryption of any KMS key in the account. The team needs to modify the policy to follow the principle of least privilege while still allowing the SageMaker endpoint to read the encrypted data. Which modification should the team make?

Hard
528

A data scientist is training a model using SageMaker and wants to use spot instances to reduce costs. The training job is checkpointed every 5 minutes. However, the job gets interrupted frequently and never completes. What is the MOST likely cause?

Hard
529

A data engineer is using Amazon SageMaker Data Wrangler to prepare a dataset for a machine learning model. The engineer needs to ensure that the data preparation flow is reproducible and can be integrated into an MLOps pipeline. Which TWO actions should the engineer take to achieve this? (Choose two.)

Medium
530

A data scientist is using Amazon SageMaker Ground Truth to create a labeled dataset for an object detection model. The dataset contains 1 million images, and the team wants to reduce labeling cost by labeling only the most informative samples. Which feature of Ground Truth should they use?

Medium
531

A company has deployed a SageMaker real-time endpoint for a model that predicts customer churn. The endpoint uses a single ml.m5.large instance. After deployment, the team notices that during peak hours, the endpoint returns 5xx errors for about 20% of requests. The endpoint has not been configured with any scaling policy. The team needs to resolve this issue with minimal cost increase. Which solution should the team implement?

Easy
532

A data engineer is building an AWS Glue ETL job that reads raw JSON clickstream events from Amazon S3, flattens nested structures, and writes Parquet to a curated S3 prefix for SageMaker training. The job must run daily on only the newly arrived files and must keep the Glue Data Catalog table current so Athena and SageMaker can query it. Which TWO configurations should the engineer apply to meet these requirements? (Choose two.)

Hard
533

A machine learning engineer is deploying a pre-trained NLP model on Amazon SageMaker for real-time inference. The model expects input sequences of variable length, and performance is critical. The engineer wants to minimize latency while handling the variable-length inputs efficiently. Which approach should the engineer choose?

Hard
534

A machine learning engineer is using SageMaker Autopilot for AutoML. Which TWO outputs does Autopilot produce?

Medium
535

A data science team is using AWS Step Functions to orchestrate a machine learning workflow that includes a SageMaker training job followed by a model deployment. They want to ensure that if the training job fails, the workflow retries up to three times with exponential backoff before sending a notification to an Amazon SNS topic. Which Step Functions feature should they use to implement this?

Medium
536

A company uses SageMaker Pipelines to automate their ML workflow. They notice that the pipeline reruns all steps even when the input data has not changed. Which feature should they enable to avoid unnecessary recomputation?

Easy
537

A company uses SageMaker endpoints for real-time inference. They want to automatically scale the number of instances based on the number of outstanding requests. Which auto-scaling policy type should they choose?

Medium
538

A team is fine-tuning a Hugging Face BERT model for text classification using SageMaker. They want to use the Hugging Face estimator for convenience. Which parameter must be set to use a custom training script?

Medium
539

A machine learning engineer is training a TensorFlow model using SageMaker with distributed training. They need to implement data parallelism across multiple GPUs. Which SageMaker feature should they use to distribute the training?

Medium
540

Refer to the exhibit. A data scientist reviews the output of a SageMaker training job. The model has 95% training accuracy and 92% validation accuracy. Which statement is true?

Easy
541

An e-commerce company uses a SageMaker endpoint to serve a product recommendation model. The model is retrained every month using batch transforms. The ML team has set up a retraining pipeline using SageMaker Processing jobs and Step Functions. Recently, the Step Functions workflow has been failing at the retraining step with an error: 'AccessDeniedException: User: arn:aws:sts::123456789012:assumed-role/RetrainingRole/abc123 is not authorized to perform: s3:GetObject on resource: arn:aws:s3:::training-data/processed/latest.parquet'. The team confirms that the S3 bucket exists and the object is present. The retraining role has the following policy: { "Version": "2012-10-17", "Statement": [ { "Effect": "Allow", "Action": [ "s3:GetObject", "s3:PutObject" ], "Resource": "arn:aws:s3:::training-data/*" } ] }. The team also verifies that the bucket policy does not explicitly deny access. What is the MOST likely cause of the AccessDenied error?

Easy
542

A company uses SageMaker to deploy a model and wants to perform A/B testing by splitting traffic between two model variants. Which TWO actions should they take? (Select TWO.)

Medium
543

A data scientist is using SageMaker Autopilot to automatically build a model. Which TWO aspects does Autopilot handle? (Choose TWO.)

Easy
544

A team is fine-tuning a large language model (LLM) using SageMaker and wants to reduce memory footprint during training. Which technique should they use?

Medium
545

Which TWO options are recommended best practices for monitoring model performance in production on SageMaker? (Choose 2.)

Medium
546

A machine learning engineer needs to give a data scientist read-only access to the model artifacts, training metrics, and monitoring reports stored in a single Amazon S3 bucket used by a SageMaker project, while ensuring the data scientist cannot delete or overwrite any object. Which approach follows least-privilege practice?

Easy
547

A data scientist is using SageMaker Automatic Model Tuning to find the best hyperparameters for a model. They want to reduce the total tuning time for a given number of training jobs. Which tuning strategy should they choose?

Easy
548

A data science team has trained a model using SageMaker and wants to deploy it to a production endpoint with automatic scaling based on request volume. Which SageMaker feature should they use to configure scaling?

Easy
549

A team uses MLflow on SageMaker for experiment tracking. They want to automate the retraining of a model when new training data arrives in an S3 bucket. Which combination of services should they use?

Medium
550

A company uses SageMaker Inference Recommender to select the optimal endpoint configuration. After running the recommender, they receive a recommendation for a specific instance type and initial instance count. What should they do next to optimize costs over time?

Medium
551

A data scientist creates a feature group as shown in the exhibit. When ingesting data with an 'age' column of integer values, the ingestion fails. What is the most likely cause?

Hard
552

A machine learning engineer is building a time-series forecasting model to predict daily sales for the next 30 days. The dataset spans two years of daily sales data. To evaluate model performance, the engineer needs to simulate a realistic forecasting scenario where the model is trained on past data and tested on future data without leakage. Which data splitting strategy should they use?

Hard
553

A company uses SageMaker Studio for collaborative ML development. The security team requires that all SageMaker Studio notebooks run within a VPC and cannot access the public internet. Which configuration should the administrator set?

Easy
554

A company needs to secure a SageMaker notebook instance that contains sensitive data. Which THREE of the following are effective security measures? (Select THREE.)

Hard
555

A company collects sensor data from IoT devices. The data arrives with missing timestamps due to network issues. For anomaly detection, the engineer needs to create features that capture rolling statistics over fixed windows. Which data preprocessing step is essential before feature generation?

Medium
556

A data scientist is using SageMaker Automatic Model Tuning with Hyperband. They want to stop poorly performing trials early to save resources. Which strategy does Hyperband use?

Medium
557

An organization wants to schedule a retraining pipeline to run every Sunday night. Which AWS service should they use to trigger the pipeline on a schedule?

Easy
558

A data engineer is using AWS Glue to prepare a dataset for machine learning. The dataset has several columns with outliers. The engineer wants to detect and handle outliers in a scalable manner. Which TWO approaches should the engineer consider? (Select TWO.)

Easy
559

A machine learning team at a bank is training a binary classification model using SageMaker's built-in XGBoost algorithm on a dataset with 20 million rows and 300 features. They need to reduce training time while maintaining model accuracy. The data is stored in Amazon S3 as CSV files. Which approach should they take to speed up training?

Medium
560

A team is monitoring a SageMaker endpoint and notices that the average latency (ModelLatency) is increasing over time, but the number of invocations is steady. They suspect that the model's inference code is becoming slower due to memory leaks. Which metric should they also examine to confirm this hypothesis?

Medium
561

A company is using Amazon SageMaker Data Wrangler to prepare a dataset for training. They have created a data flow with multiple transforms. Which TWO actions can they take to operationalize the data preparation pipeline for production? (Choose 2)

Medium
562

A machine learning engineer is using SageMaker Debugger to monitor a training job and wants to detect issues early. The engineer wants to receive alerts when the training job is likely to fail due to vanishing gradients and when the loss is not decreasing. Which two actions should the engineer take to achieve this? (Choose two.)

Medium
563

A team needs to split a time-series dataset for a forecasting model. They want to avoid data leakage and evaluate model performance on future unseen data. Which data splitting strategy should they use?

Medium
564

A machine learning engineer needs to deploy a TensorFlow model that requires a custom inference environment with specific system libraries. The model will be used in a real-time application with variable traffic. They want to minimize cold start latency. Which SageMaker hosting option should they choose?

Medium
565

An ML engineer wants to be notified when the average inference latency of a SageMaker endpoint exceeds 500 ms for 2 consecutive evaluation periods. Which AWS service combination should they use?

Easy
566

A team uses SageMaker Neo to compile a model for deployment on a target device. After compilation, they deploy the compiled model to a SageMaker endpoint using the Neo-optimized container. The endpoint fails to start with error "RuntimeError: Unable to load model". What could be the issue?

Hard
567

A company wants to use SageMaker Autopilot for a regression problem. They require an explainability report that shows feature importance globally. Which Autopilot feature should they enable?

Medium
568

A company uses SageMaker to train a model. The training job is failing with an error "ResourceLimitExceeded". What is the most likely cause?

Easy
569

A machine learning engineer needs to reduce costs when training a large model on SageMaker. They are willing to accept potential interruptions and have checkpointing enabled. Which instance purchasing option should they use?

Easy
570

A company wants to deploy a model using a serverless inference endpoint that can automatically scale to zero when not in use and has a configurable maximum concurrency. Which SageMaker inference option meets these requirements?

Easy
571

A company wants to automate the deployment of a SageMaker model into production whenever a new model version is approved in the Model Registry. Which service can be used to trigger the deployment pipeline?

Easy
572

A machine learning engineer is evaluating a binary classification model that predicts customer churn. The model achieves 95% accuracy, but the engineer suspects class imbalance is causing a misleading metric. Which THREE evaluation steps should the engineer perform to properly assess the model? (Choose THREE.)

Hard
573

A company uses SageMaker Model Monitor for feature attribution drift monitoring with SHAP. Which THREE prerequisites must be in place before starting the monitoring schedule? (Select THREE)

Medium
574

A data science team uses Amazon SageMaker Model Monitor to detect data drift in production. They notice that the schema of incoming data (number of features) has changed compared to the training baseline. Which type of monitor is BEST suited to detect this issue?

Medium
575

A data scientist must join a 50 GB transactional table with a small 5 MB lookup table in AWS Glue before writing Parquet output for SageMaker training. The join currently shuffles the large table across the cluster and the job runs slowly. Which optimization should the data scientist apply?

Medium
576

An ML team is deploying a model using SageMaker. The model requires GPU inference and must be available in multiple AWS regions for low latency. The team has created a multi-model endpoint with GPU instances. After deployment, they notice high latency spikes when a new model is loaded. What is the most likely cause?

Medium
577

A financial services company must ensure that all data used by Amazon SageMaker training jobs is encrypted at rest. The company wants to use a customer-managed key (CMK) for the encryption. Which steps are necessary to achieve this? (Choose TWO.)

Hard
578

A team wants to automate the retraining and deployment of an ML model whenever new labeled data arrives in S3. The workflow includes data preprocessing, training, evaluation, and conditional deployment. Which AWS service is best suited for orchestrating this end-to-end pipeline?

Easy
579

A data scientist is training an object detection model using SageMaker built-in Object Detection algorithm. They want to visualize the bounding boxes on validation images after training. Which approach should they use?

Medium
580

A data scientist wants to version control trained models and manage approvals for deployment. Which SageMaker feature should they use?

Easy
581

A company has a batch transform job in Amazon SageMaker that processes large datasets every night. Recently, the job has been failing sporadically with an out-of-memory error. The data size has not increased. What is the MOST likely cause?

Medium
582

A company is deploying a SageMaker real-time endpoint and needs to monitor inference latency. Which THREE metrics are available from SageMaker for this purpose? (Choose THREE.)

Medium
583

A machine learning engineer wants to automatically trigger a retraining pipeline whenever new training data arrives in an S3 bucket. The pipeline uses SageMaker Pipelines. Which AWS service should be used to detect the S3 event and start the pipeline?

Easy
584

A team built a SageMaker Pipeline that includes a training step and a model evaluation step. They want to automatically register a model in SageMaker Model Registry only if the evaluation metric (accuracy) exceeds 0.9. Which pipeline step should be used to implement this conditional logic?

Medium
585

An ML engineer is designing a SageMaker Pipeline for model training and registration. They need to ensure that the pipeline can be re-run with different datasets without manual intervention, and that the steps are only re-executed if inputs have changed. Which THREE features should they configure? (Select THREE.)

Hard
586

Refer to the exhibit. A team configured a SageMaker Model Monitor schedule for data quality. The baseline was created from a training dataset. After running for a day, the monitoring results show frequent violations. What is the most likely cause?

Easy
587

An ML engineer is fine-tuning a large language model using LoRA on SageMaker. The training is converging slowly, and GPU utilization is low. The engineer suspects the bottleneck is data loading. Which action should the engineer take to improve GPU utilization?

Hard
588

A data engineer is optimizing Amazon Athena queries on large datasets stored in S3 for machine learning data preparation. Which THREE practices improve query performance?

Hard
589

Refer to the exhibit. A user has the above IAM policy attached but cannot access files in SageMaker Studio. What additional permission is most likely needed?

Easy
590

A machine learning engineer needs to deploy a TensorFlow model to Amazon SageMaker and wants to use the built-in TensorFlow Serving container. What should the engineer provide in the model archive?

Easy
591

A company wants to deploy its trained model to edge devices such as cameras and IoT devices. The model must run efficiently with low latency and minimal memory footprint. Which THREE actions should the company take to prepare the model for edge deployment? (Choose THREE.)

Easy
592

A machine learning engineer is deploying a model to a SageMaker real-time endpoint that must be accessible only from within a specific Amazon VPC and must not have a public IP address. The engineer also needs to ensure that all data in transit between the endpoint and the calling application is encrypted. Which configuration should the engineer use?

Medium
593

A machine learning engineer needs to monitor a SageMaker endpoint for data drift and receive alerts when drift is detected. They want to use a fully managed AWS service to schedule the monitoring jobs and send notifications. Which AWS service should they use to orchestrate the monitoring schedule and trigger alerts?

Easy
594

A machine learning practitioner is building a binary classifier with severe class imbalance (1:1000). They want to use SMOTE for oversampling. What is a potential drawback of applying SMOTE on the entire dataset before splitting into training and test sets?

Hard
595

A team is training a large language model and needs to split the model layers across multiple GPUs due to memory constraints. Which distributed training strategy should they use?

Medium
596

A data engineer stores a training dataset in Amazon S3 across a prefix hierarchy organized by year, month, and day, and analysts query it with Amazon Athena. The engineer wants Athena to scan only the relevant date range instead of the entire dataset. Which action should the engineer take?

Easy
597

A team is migrating their ML infrastructure to AWS and wants to use infrastructure as code to manage SageMaker Studio domains, user profiles, and associated resources. Which services can they use for this purpose? (Select THREE.)

Medium
598

A team is training a deep learning model on Amazon SageMaker using a custom Docker container. Which three practices should they follow to optimize training performance? (Choose three.)

Medium
599

A company uses SageMaker Model Monitor to detect bias drift in their real-time inference endpoint. They have collected ground truth labels and want to monitor for bias across different demographic groups. Which type of monitoring should they configure?

Medium
600

A machine learning engineer is preparing a dataset for binary classification. The target variable has a severe class imbalance (95% negative, 5% positive). Which technique can help address this imbalance during data preparation?

Easy
601

An ML team is preparing time-series data for a demand forecasting model. They want to evaluate model performance over time without leaking future information into past training windows. Which data splitting strategy is MOST appropriate?

Medium
602

A team needs to deploy a PyTorch model that uses custom CUDA kernels. They want to use NVIDIA Triton Inference Server on SageMaker for high-performance serving. Which SageMaker configuration is required to use Triton?

Hard
603

A data scientist is preparing a dataset for training a binary classification model. The dataset has 100,000 rows and 50 features. The target variable is imbalanced, with only 5% positive cases. Which technique should the data scientist apply to address the class imbalance BEFORE training?

Easy
604

During deployment of a Hugging Face model, the endpoint logs show this error. Which step was likely missed?

Hard
605

A company is training a deep learning model on Amazon SageMaker. The training job started but has been stuck in 'InProgress' state for an unusually long time with low CPU utilization. The data scientist suspects a bottleneck. What should be the first troubleshooting step?

Medium
606

A machine learning engineer needs to prepare a dataset containing customer transactions for training a fraud detection model. The dataset includes features such as transaction amount, timestamp, merchant category, and customer ID. The engineer wants to create a feature representing the average transaction amount per customer over the last 7 days. Which approach should be used in Amazon SageMaker Data Wrangler?

Medium
607

A financial services company is deploying a fraud detection model on SageMaker. To comply with regulations, they must ensure that the model's predictions are not biased against protected groups. They plan to monitor bias drift post-deployment using SageMaker Clarify. Which data inputs are required to configure Clarify's bias drift monitoring?

Hard
608

A company deploys a model on Amazon SageMaker for real-time inference. The inference latency is too high. The model is a large deep learning model. The company wants to reduce latency without significantly impacting accuracy. Which approach should the company consider?

Medium
609

A data scientist is building a text classification model using Amazon SageMaker. The dataset is stored as a CSV file in Amazon S3. The scientist wants to use the SageMaker built-in BlazingText algorithm. Which of the following steps are required to prepare the data for training? (Choose TWO.)

Hard
610

A company wants to reduce costs for a SageMaker real-time endpoint that receives predictable traffic patterns: high during business hours and low at night. The model is a small PyTorch model. Which cost-saving strategy is most suitable?

Easy
611

A machine learning engineer needs to ingest streaming data from thousands of IoT devices into Amazon S3 for batch training. The data should be available in S3 within minutes of arrival. Which combination of services should the engineer use?

Medium
612

Which technique is commonly used to handle missing values in a categorical feature?

Easy
613

An MLOps engineer is building an automated retraining pipeline for a fraud detection model. The model must be retrained weekly, and the new model should only be promoted to production if it meets predefined performance thresholds compared to the current model. Which combination of SageMaker capabilities should the engineer use?

Hard
614

A machine learning engineer is using SageMaker Pipelines to orchestrate a training workflow. The pipeline includes a processing step that outputs a dataset, which is then used by a training step. The engineer notices that the processing step runs every time the pipeline executes, even when the input data has not changed. The engineer wants to avoid re-running the processing step if the input data and code are unchanged, while ensuring that downstream steps still execute if the processing step is skipped. Which approach should the engineer take?

Hard
615

A media company stores training data in an S3 bucket encrypted with an AWS KMS customer-managed key. A SageMaker training job runs inside a private VPC subnet with no internet access and must read that bucket. The job currently fails with an access-denied error from S3. Which change most directly resolves the failure while preserving the private-network requirement?

Hard
616

A team has deployed a real-time inference endpoint and wants to automatically scale based on CPU utilization. Which scaling policy type should they use with Application Auto Scaling for SageMaker endpoints?

Medium
617

A company uses Amazon SageMaker to train and deploy a machine learning model. After deployment, they notice that the model's accuracy drops significantly over time due to changes in the underlying data distribution. Which monitoring solution should they implement to detect this issue automatically?

Medium
618

A team uses SageMaker for training. They need to monitor training progress and view metrics like loss and accuracy. Which SageMaker feature should they use?

Easy
619

A data scientist wants to evaluate the performance of a binary classification model. The dataset is highly imbalanced with only 5% positive class. Which metric should be used to evaluate the model?

Easy
620

A team has 200 small ML models that need to be served via HTTPS endpoints. Each model is used infrequently, and the team wants to minimize hosting costs. Which SageMaker deployment approach is MOST cost-effective?

Medium
621

A machine learning team uses SageMaker Pipelines to automate retraining. They want to avoid re-running data processing steps if the data has not changed since the last successful pipeline run. Which built-in feature should they enable?

Medium
622

A model deployed on a SageMaker endpoint is returning predictions. The team wants to log all predictions to an S3 bucket for auditing. What is the most efficient way to achieve this?

Medium
623

A data scientist observes that a linear regression model has many irrelevant features. They want to perform feature selection to improve generalization. Which method combines feature selection with model training using a penalty that can shrink coefficients to zero?

Hard
624

A team is training a large deep learning model on SageMaker using a single ml.p3.16xlarge instance. Training is taking too long. They want to reduce time by distributing across multiple GPUs but are constrained by model size that does not fit in a single GPU memory. Which distributed training strategy should they use?

Medium
625

A machine learning team is building a model to predict customer churn. The dataset includes a feature 'customer_tenure' with values ranging from 1 to 100 months, and 'monthly_spend' ranging from $10 to $5000. The model will use gradient boosting. Which feature scaling approach is most appropriate?

Hard
626

A healthcare company is building a model to predict patient readmission rates. The dataset contains a mix of numeric features (age, blood pressure, lab test results) and categorical features (gender, diagnosis code, hospital department). The dataset has 2 million rows. The data is stored in an Amazon S3 bucket, and they use AWS Glue to catalog and preprocess the data. The data scientist notices that the 'diagnosis_code' column has 10,000 unique codes, and 20% of the rows have missing values for 'blood_pressure'. They plan to use a SageMaker built-in XGBoost model. For optimal model performance, which preprocessing steps should they apply using AWS Glue ETL?

Medium
627

A company deploys a model on SageMaker that serves predictions to a web application. The model's performance degrades over time due to data drift. The company wants to set up continuous monitoring. Which TWO actions should the company take to monitor and retrain the model effectively? (Choose TWO.)

Medium
628

A company is using AWS Step Functions to orchestrate their ML retraining pipeline. They want to trigger retraining when new data arrives, but only if the model's performance has degraded below a threshold. Which THREE AWS services should they use together to achieve this? (Choose three.)

Medium
629

A data scientist is building a text classification model using Amazon SageMaker. The dataset is large and includes imbalanced classes. Which three techniques can help improve model performance? (Choose three.)

Medium
630

Which TWO tools are specifically designed for debugging and analyzing training jobs in SageMaker?

Hard
631

A data scientist is using Amazon SageMaker Debugger to monitor training metrics. They want to stop training automatically if the model is overfitting. Which action should they take?

Hard
632

Which SageMaker feature provides AutoML capabilities, including automatic data preprocessing, model selection, and hyperparameter tuning?

Easy
633

A company uses SageMaker Pipelines to automate their ML workflow. They need to add model versioning and approval workflow. Which THREE steps should they include in their pipeline to achieve this? (Choose THREE.)

Medium
634

Which THREE components are required to set up automated model retraining in response to performance degradation using Amazon SageMaker? (Select THREE.)

Medium
635

A company wants to trigger a model retraining pipeline whenever new training data arrives in an S3 bucket. They also need to send a notification to a Slack channel when the retraining completes. Which TWO AWS services should they use to implement this event-driven workflow? (Select TWO.)

Easy
636

A data scientist is building a regression model to predict house prices. The dataset contains a feature 'neighborhood' with 500 distinct values, and most neighborhoods have fewer than 10 samples. Which approach is MOST appropriate for handling this high-cardinality categorical feature?

Medium
637

A company has 50 small PyTorch models that are used infrequently for inference. They want to minimize costs while maintaining the ability to serve all models from a single endpoint. Which SageMaker feature should they use?

Easy
638

A machine learning engineer needs to select features for a regression model. The dataset contains 50 numeric features, and the target variable is continuous. The engineer wants to reduce dimensionality by selecting features that have the strongest linear relationship with the target. Which feature selection method is MOST appropriate?

Medium
639

Which SageMaker built-in algorithm is designed for time series forecasting?

Easy
640

A data scientist is evaluating a binary classification model. They have the confusion matrix and want to assess the model's performance comprehensively. Which THREE metrics should they consider? (Select THREE.)

Medium
641

A machine learning engineer is training a model using SageMaker and wants to set up monitoring to detect if gradients become too large, which could destabilize training. Which SageMaker Debugger built-in rule should they enable?

Medium
642

A company is using Amazon SageMaker to train a large deep learning model. The training job is taking a very long time. The data scientist suspects that the GPU utilization is low due to inefficient data loading. Which action should the data scientist take to diagnose and address this issue?

Medium
643

Refer to the exhibit. A team observes that their SageMaker endpoint scales out quickly when load increases, but scales in very slowly when load decreases, causing over-provisioning. What is the most likely cause?

Medium
644

A machine learning team needs to automatically retrain a model when concept drift is detected in the deployed endpoint's predictions. Which TWO steps should they take? (Choose TWO.)

Medium
645

A machine learning engineer is building an ML pipeline using Amazon SageMaker. The engineer needs to prepare the data, detect bias in the dataset, and then create features for training. Which TWO AWS services or features should the engineer use? (Choose TWO.)

Medium
646

A machine learning engineer is developing a text classification model using Amazon SageMaker. The dataset consists of 1 million customer reviews, with labels indicating sentiment (positive, negative, neutral). The engineer uses a pre-trained BERT model from the Hugging Face Model Hub and fine-tunes it on the dataset using SageMaker's Hugging Face estimator with a ml.p3.2xlarge instance. After 2 hours of training, the training job fails with a 'ResourceExhaustedError: CUDA out of memory' error. The error occurs during the forward pass of the first epoch. The engineer confirms that the batch size is set to 32, the maximum sequence length is 512 tokens, and the dataset is stored in a S3 bucket in the same AWS region. The engineer needs to complete fine-tuning without increasing instance costs. Which course of action should the engineer take?

Medium
647

A company wants to reduce costs for a production SageMaker endpoint that has predictable traffic patterns. They have purchased a Savings Plan. What additional step can they take to further optimize costs while maintaining performance?

Easy
648

A data scientist uses SageMaker Model Monitor to track feature attribution drift. Which technique does SageMaker Model Monitor use to compute feature attributions?

Medium
649

A team is using AWS Glue to process streaming data from Amazon Kinesis. The streaming data contains both structured and semi-structured fields. The team needs to flatten the semi-structured fields into columns for downstream ML training. Which Glue feature is BEST suited?

Hard
650

A machine learning team needs to deploy a PyTorch model that has been compiled with SageMaker Neo to improve inference performance on edge devices. Which TWO statements about SageMaker Neo are correct? (Select TWO.)

Medium
651

A company uses SageMaker endpoints with auto-scaling based on CPU utilization. During a flash sale, latency increases despite low CPU. What should be done?

Hard
652

A data scientist needs to evaluate a binary classification model. The dataset is highly imbalanced (5% positive class). Which metric is MOST appropriate for assessing model performance?

Medium
653

A company wants to track the lineage of their ML models, including the training dataset, hyperparameters, and training job used to produce each model version. Which AWS service should they use?

Easy
654

A data scientist wants to use SageMaker Autopilot to automatically build a regression model. The dataset contains 200 features and 50,000 rows. Which output does SageMaker Autopilot provide?

Easy
655

A company has a SageMaker endpoint running a model that provides real-time recommendations. Recently, the model's accuracy has degraded due to data drift. The team wants to automatically retrain the model when a drift metric exceeds a threshold and deploy the new model without downtime. Which architecture should the team implement?

Hard
656

A data scientist is preparing a dataset for binary classification. The dataset has a target variable with 90% of samples belonging to class 0 and 10% to class 1. Which data splitting strategy should the scientist use to ensure that the training and test sets maintain the same class proportion as the original dataset?

Easy
657

A company deploys a model with SageMaker and wants to monitor for concept drift. They have noticed that the relationship between input features and the target variable has changed, causing model accuracy to degrade. However, the input data distribution remains stable. Which type of drift is this, and what is the most appropriate response strategy?

Hard
658

Refer to the exhibit. A data scientist tries to deploy a model from an S3 bucket encrypted with SSE-KMS. What should the administrator do to resolve this?

Medium
659

A machine learning engineer is using Lasso regression for feature selection. After training, many coefficients become zero. The engineer notices that some features with high mutual information with the target also have zero coefficients. What is the most likely reason?

Hard
660

A data scientist is using SageMaker to train an XGBoost model for a regression problem. After training, they evaluate the model on a test set and get an RMSE of 10 and an R² of 0.85. Which additional metric would give the MOST insight into the model's average prediction error magnitude?

Medium
661

An ML team is building a time-series forecasting model for daily sales. They need to split the data into training and validation sets without data leakage, and the validation set should be the most recent 30 days. Which splitting strategy should they use?

Hard
662

A financial services company needs to deploy a machine learning model for real-time fraud detection. The model must be highly available across multiple Availability Zones and must support automatic scaling based on request volume. The company also needs to perform canary deployments to test new model versions with a small percentage of traffic before full rollout. Which SageMaker feature should they use?

Hard
663

A machine learning engineer is preparing a training script that must run on multiple GPU instances with SageMaker. The script currently reads the entire training dataset from local disk into memory, which fails on larger datasets. The engineer wants the script to stream training data from the SageMaker training channel path without loading everything into memory. Which approach should the engineer take?

Medium
664

An organization needs to ensure that all data transmitted between containers in a SageMaker training job is encrypted. In the training job configuration, which setting should they enable?

Hard
665

A data scientist wants to fine-tune a large language model for a question-answering task. They want to reduce memory usage during training by using a low-rank approximation of the weight updates. Which technique should they use?

Medium

Frequently asked questions

What does the scenario questions domain cover on the MLA-C01 exam?
scenario questions questions test whether you can apply the concept in context, not just recognise a definition.
How many questions are in this domain?
This page lists all 665 scenario questions questions in the MLA-C01 question bank. The actual exam draws from this domain proportionally to its weighting in the official exam blueprint.
What is the best way to practise this domain?
Start with a short focused session (10 questions) to identify gaps, then work through explanations. Repeat with a longer session once the weak areas feel solid.
Can I practise only scenario questions questions?
Yes — the session launcher on this page filters questions to this domain only. Choose any session length for inline explanations and scoring.