Be able to configure and troubleshoot SageMaker training and hosting: pick the right endpoint type, wire IAM roles to S3 correctly, and read AWS CLI errors. The single most important thing is verifying permissions and data paths before blaming the model code.
Start practicing
Machine Learning Implementation and Operations — choose a session length
Free · No account required
Domain overview
This domain covers deploying, operating, and troubleshooting ML workloads on AWS. Expect questions on SageMaker hosting options, training job configuration, IAM roles and S3 permissions, container images, hyperparameter tuning jobs, and debugging access-denied or input-data failures using the AWS CLI and console.
Exam objectives
Choosing SageMaker hosting options such as real-time, serverless, async, and batch transform for latency and payload needs
Diagnosing IAM role and S3 bucket policy errors that cause training jobs to fail with access denied
Tracking and querying hyperparameter tuning job metadata, configurations, and results with SageMaker and related services
Creating and debugging SageMaker training jobs via the AWS CLI, including input channels and container configuration
Assuming the SageMaker execution role automatically has S3 access; you must grant explicit permissions to the bucket and prefix.
Confusing real-time endpoints with Serverless Inference or Batch Transform when cold start latency or payload size matters.
Forgetting that custom containers must expose the correct inference handler and be pushed to Amazon ECR before hosting.
Click any question to see the full explanation and answer options, or start a focused practice session above.
A company is using Amazon SageMaker to train a model with a custom algorithm. The training script reads data from an S3 bucket using boto3. The training job fails with an 'AccessDenied' error when trying to access the S3 bucket. The IAM role attached to the SageMaker notebook instance has full S3 access. What is the most likely cause?
2An IAM policy is attached to a SageMaker execution role. A data scientist tries to create a training job using a custom algorithm stored in an ECR repository. The training job fails with an 'AccessDenied' error when pulling the Docker image from ECR. What is the missing permission?
3A DevOps engineer created a SageMaker notebook instance using the Terraform configuration shown. The notebook instance is in a VPC with a public subnet. However, the notebook instance cannot access the internet. What is the most likely cause?
4A company is using Amazon SageMaker to train a XGBoost model on a large dataset. The training job is taking a long time. The data scientist wants to reduce training time without sacrificing model accuracy. The dataset is 100 GB in CSV format stored in S3. What is the most effective approach?
5A data scientist is deploying a model using Amazon SageMaker for real-time inference. The model is memory-intensive and requires a GPU. Which instance type should be selected for the endpoint?
6A data scientist is deploying a model on Amazon SageMaker for real-time inference. The model is a PyTorch model that requires custom inference code. The data scientist needs to handle variable-length inputs and optimize inference latency. Which TWO steps should the data scientist take? (Choose TWO.)
7A media company uses Amazon SageMaker to train a deep learning model for video classification. The training job uses a single ml.p3.2xlarge instance and processes 50 GB of labeled video data stored in Amazon S3. The training completes successfully in 12 hours. However, the data scientists report that the model’s accuracy is lower than expected. They suspect the training data contains labeling errors. To improve model accuracy without incurring significant additional cost, they want to identify and remove mislabeled training examples before retraining. They have a small budget of $50 and need to complete the analysis within 2 hours. Which approach should the data scientists take?
8An ML engineer is troubleshooting why an automated CI/CD pipeline cannot deploy an updated model to an existing SageMaker endpoint. The pipeline uses the IAM role that has the attached policy shown in the exhibit. What is the MOST likely cause of the failure?
9Drag and drop the steps to train a model using Amazon SageMaker built-in algorithm in the correct order.
10A data scientist is training a linear regression model on a dataset with 100 features. The model shows high variance on the test set. Which action is MOST likely to reduce overfitting?
11A machine learning team is building a fraud detection system using Amazon SageMaker. The training data is highly imbalanced (99% legitimate, 1% fraudulent). They need to maximize the recall of the fraud class while keeping precision above 90%. Which approach should they take?
12A data scientist wants to use Amazon SageMaker to train a deep learning model on a large dataset stored in S3. The training job is expected to take several hours. Which storage option should be used to minimize data loading time and cost?
13An ML engineer is deploying a model to a SageMaker endpoint for real-time inference. The model requires a custom inference script that preprocesses input data and postprocesses predictions. Which SageMaker feature should be used to implement this custom logic?
14A data science team is deploying a machine learning model to production using Amazon SageMaker. The model requires real-time inference with low latency. Which SageMaker feature should they use to deploy the model?
15During training of a deep learning model on a GPU instance in SageMaker, the training job fails with an insufficient memory error. Which step should be taken first to resolve this issue?
16A company uses SageMaker to train a model each night. The training data is stored in an S3 bucket with SSE-S3 encryption. The training job fails with an access denied error. Which configuration is needed?
17A data scientist is deploying a model to a SageMaker endpoint and needs to optimize for cost while maintaining low latency. Which TWO actions should the data scientist take?
18Refer to the exhibit. A data scientist runs the AWS CLI command to create a SageMaker training job. The training job fails because the input data is not accessible. Which step should the data scientist take to fix the issue?
19Refer to the exhibit. A SageMaker training job uses an IAM role with this policy. The training job writes output to s3://my-bucket/output/. Which statement about the policy is true?
20A data scientist needs to create a SageMaker notebook instance with access to a private S3 bucket. The bucket uses SSE-KMS encryption. Which additional configuration is required?
21A company uses SageMaker to run training jobs on a schedule. The training data is stored in an S3 bucket that receives new data every hour. Which TWO approaches can the company use to trigger a training job when new data arrives?
22A data scientist needs to deploy a PyTorch model for real-time inference. Which AWS service is best suited for this task?
23A deployed SageMaker endpoint is returning high latency. The model is a scikit-learn Random Forest. Which action is most likely to reduce latency?
24A company wants to serve predictions from a model using a REST API with low latency. Which SageMaker deployment option is most appropriate?
25A company wants to monitor a deployed model for data drift. Which AWS service should they use?
26Which THREE factors should be considered when choosing an instance type for a SageMaker training job?
27Which TWO services can be used to orchestrate a machine learning pipeline?
28Refer to the exhibit. A SageMaker training job is launched with the CLI command shown. The job fails with an error 'S3 data distribution type not supported for File mode'. What is the most likely fix?
29A machine learning team is deploying a model using Amazon SageMaker. The model inference code runs on GPUs and requires a custom container. The team wants to minimize cold start latency. Which SageMaker hosting option should they use?
30A data scientist is training a deep learning model on a large dataset using SageMaker. The training job is taking too long. Upon reviewing the CloudWatch logs, the scientist notices that the GPU utilization is below 10% most of the time. Which change is MOST likely to improve GPU utilization and reduce training time?
31A company is using Amazon SageMaker to train a model and wants to track hyperparameter tuning jobs. Which AWS service is BEST suited to store and query metadata such as tuning job configurations and results?
32A company is using Amazon SageMaker to train a model and wants to automatically retrain the model every week using new data. Which AWS service should be used to orchestrate the retraining pipeline?
33A machine learning team is using SageMaker to train a model. The training data is stored in an S3 bucket encrypted with AWS KMS. The training job fails with an 'AccessDenied' error. Which IAM permission is MOST likely missing from the SageMaker execution role?
34A company is using Amazon SageMaker Ground Truth to create a labeled dataset for object detection. The labeling job is taking longer than expected. The team notices that many workers are spending a lot of time on images with no objects. Which labeling strategy should they use to reduce costs and time?
35Which TWO factors should be considered when choosing between Amazon SageMaker's real-time endpoints and serverless inference? (Select TWO.)
36Which THREE measures can help reduce inference latency for a deep learning model deployed on SageMaker real-time endpoints? (Select THREE.)
37Which TWO actions are best practices for securing a SageMaker notebook instance? (Select TWO.)
38A data scientist is training a neural network on a GPU instance in Amazon SageMaker. The training job fails with an 'OutOfMemoryError'. Which action should the data scientist take to resolve this issue?
39A machine learning engineer is deploying a model to an Amazon SageMaker endpoint. The model is a PyTorch model that requires a custom inference script. The engineer notices that the endpoint is returning 500 errors after deployment. Which step should the engineer take to debug the issue?
40A data scientist is using AWS Glue to prepare training data. The job reads from an S3 bucket, performs transformations, and writes to another S3 bucket. The job is failing due to insufficient memory. Which solution should the data scientist use to fix this?
41A company is building a fraud detection model. The dataset is highly imbalanced (99% legitimate, 1% fraud). The data scientist trains a model using Amazon SageMaker's built-in XGBoost algorithm. The model achieves 99% accuracy but only catches 10% of fraud cases. Which technique should the data scientist apply to improve recall for the minority class?
42A data scientist needs to version control datasets used for machine learning experiments. Which AWS service should the data scientist use?
43A company is using Amazon SageMaker to train a model on a large dataset stored in S3. The training job is taking a long time due to slow data loading. Which action can the data scientist take to reduce data loading time?
44A data scientist is using Amazon SageMaker to train a model. The training data is stored in an S3 bucket encrypted with AWS KMS. Which TWO actions are necessary to allow SageMaker to access the data?
45A company is deploying a machine learning model using Amazon SageMaker. The model needs to be updated frequently. Which THREE practices should the company implement for model versioning and deployment?
46A data scientist is reviewing the training logs from a SageMaker training job. The model's loss decreases steadily and accuracy increases. However, when the model is evaluated on a holdout test set, the accuracy is only 0.65. Which issue does this behavior suggest?
47A data scientist is using Amazon SageMaker to train a custom image classification model using a PyTorch script. The training job runs successfully but the model accuracy is lower than expected. The scientist wants to debug the training process by inspecting gradients and layer outputs. Which SageMaker feature should be used to capture this internal state during training?
48A company uses Amazon SageMaker to train a model using a custom Docker container. The training job fails with an error: "Unable to write to /opt/ml/output/data". The data scientist checks the container and finds that the /opt/ml directory is not writable. What is the MOST likely cause?
49An ML team wants to perform batch inference on a large dataset stored in Amazon S3 using a pre-trained model. The team needs to process the data in parallel across multiple instances to reduce processing time. Which approach should they use?
50Refer to the exhibit. An ML engineer attaches this IAM policy to a user. The user wants to invoke the SageMaker endpoint my-endpoint from an EC2 instance with public IP 52.1.1.1. What will happen?
51A financial services company is deploying a machine learning model for credit risk assessment. The model must have an inference latency under 200ms and must be able to handle up to 1000 transactions per second (TPS). The company wants to minimize costs. The model is a gradient boosting model implemented in XGBoost. Which SageMaker deployment option should the team choose?
52A company wants to use Amazon SageMaker to host a model that was trained using a custom algorithm. The model artifact is stored in Amazon S3. The company wants to ensure that the endpoint can automatically scale based on the number of incoming requests. Which configuration should the company use?
53A company uses Amazon SageMaker to train a model. The training job fails with an 'OutOfMemory' error. The training data is stored in S3 and the instance type is ml.m5.xlarge. What is the most efficient way to resolve this issue?
54A team uses AWS Glue ETL jobs to preprocess data for SageMaker training. The job runs successfully but the output data is empty. What is the most likely cause?
55A company uses SageMaker to host a model for real-time predictions. The model is updated weekly. To minimize downtime during model updates, what should the company do?
56A company is using SageMaker to host a model that makes predictions on streaming data from Amazon Kinesis. The model must provide predictions with sub-second latency. Which approach should the company use?
57A team is using SageMaker to train a model. They want to track hyperparameters, metrics, and model artifacts. Which SageMaker feature should they use?
58A company is deploying a SageMaker model for real-time inference. The endpoint must be highly available and cost-effective. Which TWO actions should the company take? (Select TWO.)
59A company uses SageMaker to train a model. The training job is taking too long and the data scientist wants to speed it up. Which THREE strategies should the data scientist consider? (Select THREE.)
60A SageMaker endpoint has a CloudWatch alarm configured as shown in the exhibit. The alarm fires when the p99 latency exceeds 500 ms for two consecutive minutes. Which action should the data scientist take to reduce latency?
61A company is using Amazon SageMaker Ground Truth to create labeled datasets for a text classification task. The labeling job uses a private workforce of 10 annotators. After labeling 10,000 items, the quality of labels is inconsistent. Which approach will MOST effectively improve labeling consistency?
62A data scientist trains a model using Amazon SageMaker's built-in XGBoost algorithm. The model overfits on the training data. Which hyperparameter adjustment is MOST likely to reduce overfitting?
63A company is using Amazon SageMaker to train a deep learning model for image classification. The training job is using a single p3.2xlarge instance and takes 10 hours. The data scientist wants to reduce training time using distributed training. Which SageMaker feature should be used?
64A company is deploying a machine learning model using Amazon SageMaker. The model needs to be updated frequently with new data. Which TWO approaches can be used to update the model without downtime? (Choose TWO.)
65A company is building a machine learning pipeline on AWS. The pipeline includes data ingestion, preprocessing, training, and deployment. Which THREE AWS services can be used to orchestrate the pipeline? (Choose THREE.)
66A data scientist wants to deploy a PyTorch model for real-time inference with low latency. Which AWS service should they use?
67A company's ML model training on Amazon SageMaker is taking longer than expected. The training job uses a single ml.p3.2xlarge instance. Which change is most likely to reduce training time?
68A team is using Amazon SageMaker Autopilot to automatically build models. The dataset has 50 features and 1 million rows. After training, Autopilot generates multiple candidates. The team wants to deploy the model with the highest accuracy. What is the best practice to select and deploy the model?
69An ML engineer needs to store and version training datasets and model artifacts. Which AWS service should they use?
70A team is training a large language model using PyTorch on multiple GPUs. The training is taking too long due to inefficient data loading. Which AWS service can help accelerate data loading by caching data close to the GPU instances?
71A company's ML pipeline uses AWS Step Functions to orchestrate data preprocessing, training, and evaluation. The training step occasionally fails due to a transient error. What is the most robust way to handle this without manual intervention?
72A data scientist needs to perform hyperparameter optimization for a gradient boosting model. Which built-in Amazon SageMaker feature should they use?
73A company's ML model is deployed on a SageMaker endpoint. The model's predictions are used in a customer-facing application that requires low latency. Over time, the model's performance degrades due to data drift. What is the most suitable approach to detect this drift automatically?
74An ML team is using SageMaker Processing jobs to run feature engineering scripts. The scripts require a specific Python package not included in the default SageMaker image. How should the team provide this package?
75Which TWO options are valid ways to reduce inference latency for a model deployed on a SageMaker real-time endpoint? (Select TWO.)
76Which THREE steps should be taken to secure a SageMaker notebook instance that accesses sensitive data? (Select THREE.)
77Which TWO AWS services can be used to deploy a trained model for serverless inference? (Select TWO.)
78A company is running a real-time inference endpoint on Amazon SageMaker. The endpoint is using an ml.c5.xlarge instance. Over the past month, the CPU utilization has been consistently below 10%, and the latency is well within requirements. The company wants to reduce costs. What should they do?
79A data scientist is using Amazon SageMaker to train a model. The training job is taking longer than expected. The data scientist notices that the GPU utilization is low. Which action would most likely improve GPU utilization?
80A company is using Amazon SageMaker to deploy a model. The model is a large ensemble that requires 8 GB of memory. The company wants to minimize endpoint cost. Which instance type should they choose?
81A company is using Amazon SageMaker to train a model. The training data includes sensitive personally identifiable information (PII). The company needs to ensure that the training data is protected and that the trained model does not inadvertently expose PII. Which TWO actions should the company take? (Choose TWO.)
82A data scientist is deploying a model on Amazon SageMaker. The model requires inference on images, and the data scientist wants to use a GPU instance for low latency. However, the data scientist is unsure about the instance type to choose for the endpoint. Which TWO factors should the data scientist consider when selecting the instance type? (Choose TWO.)
83A data scientist is using Amazon SageMaker to train a model on a large dataset (10 TB) stored in S3 in Parquet format. The training job uses an ml.p3.16xlarge instance with multiple GPUs. The data scientist notices that the GPU utilization is low (around 30%) and the training is slow. The dataset consists of hundreds of thousands of small Parquet files. The data scientist suspects that the I/O is bottlenecked. What should the data scientist do to improve GPU utilization and training speed?
84A data scientist is using SageMaker to train a model. The training job is failing with a 'ResourceLimitExceeded' error. Which action should be taken to resolve this issue?
85A machine learning team is using AWS Glue to prepare data for training. They notice that the ETL job takes a long time to process large datasets. Which change is most likely to improve performance?
86A data scientist is using SageMaker Debugger to monitor a training job. The training loss is not decreasing as expected. Which Debugger feature can help identify the issue?
87A company is using Amazon Rekognition to detect objects in images stored in S3. They want to reduce costs by processing images only when they are uploaded. Which AWS service should be used to trigger Rekognition automatically?
88A machine learning engineer needs to deploy a model that requires custom inference code with dependencies. Which SageMaker deployment option should be used?
89A data scientist is using SageMaker to build a model for fraud detection. The dataset is highly imbalanced. Which THREE techniques should be applied to address class imbalance?
90A research lab is using SageMaker to train deep learning models on a custom dataset stored in S3. Each training job uses a single ml.p3.2xlarge instance. Recently, training jobs have been failing intermittently with 'NetworkError: Connection reset by peer' during the data download phase. The data scientist notices that the dataset is 50GB and the network throughput is low. The training script uses the default S3 download method (boto3) to copy data from S3 to the local instance storage. Which solution should the data scientist implement to resolve the issue?
91A media company uses SageMaker to train a recommendation model. The training data is stored in an S3 bucket with versioning enabled. The data pipeline updates the training data daily by overwriting objects with new data. Recently, the model's performance degraded, and the team suspects that the training data was corrupted on a specific day. They want to train the model using the data from a previous version. How can the team retrieve the previous version of the training data?
92A machine learning team is using Amazon SageMaker to train a PyTorch model on a dataset that is 500 GB in size. The training job runs on a single ml.p3.2xlarge instance, but the training takes over 48 hours, which exceeds the maximum allowed time. The team wants to reduce training time to under 24 hours. They are open to using multiple instances and have budget for up to 4 instances. The dataset is stored in Amazon S3 and can be split into shards by a key. The model architecture must remain unchanged. What should the team do?
93A data scientist is using Amazon SageMaker to train a TensorFlow model on a dataset that includes sensitive personal information (PII). The data is stored in Amazon S3 with server-side encryption using AWS KMS (SSE-KMS). The training job fails with an Access Denied error when trying to read from S3. The data scientist has already verified that the SageMaker execution role has s3:GetObject permissions on the S3 bucket. What additional configuration is needed?
94A data scientist is using Amazon SageMaker to train a model using the built-in XGBoost algorithm. The training job uses a hyperparameter tuning job to optimize hyperparameters. The tuning job has been running for 3 hours and has completed 20 training jobs. The data scientist wants to stop the tuning job early if it is not making progress. What should the data scientist do to accomplish this?
95A data scientist is training a model using Amazon SageMaker with a custom Docker container. The training job fails with an error: 'Resource exhausted: Out of memory'. The training data is stored in S3. What should the data scientist do to resolve this issue?
96A machine learning engineer needs to deploy a model that performs real-time fraud detection. The model must be highly available and scalable. Which AWS service should be used to host the model?
97A data scientist is using Amazon SageMaker Debugger to monitor training jobs. The training loss is decreasing but then suddenly spikes. What is the most likely cause and how should it be addressed?
98A company wants to perform automated hyperparameter tuning for a model. Which Amazon SageMaker feature should be used?
99A machine learning engineer is building a pipeline using Amazon SageMaker Pipelines. The pipeline has multiple steps including data preprocessing, training, and evaluation. Which statement about SageMaker Pipelines is correct?
100A company is using Amazon SageMaker to deploy a model for real-time inference. The endpoint receives variable traffic and the company wants to optimize cost while maintaining responsiveness. Which scaling policy should be used?
101A machine learning engineer is setting up a training job in Amazon SageMaker. Which THREE components are required to define a training job? (Choose three.)
102A company is training a deep learning model on Amazon SageMaker. The training job is failing with an out-of-memory error. Which SageMaker feature should the company use to resolve this issue without changing the instance type?
103A company is using SageMaker to train a large NLP model. The training job is taking too long due to high I/O wait time. The data is stored as CSV files in S3. Which optimization should the company implement to reduce I/O wait time?
104A machine learning team is using SageMaker to build a model. They need to track hyperparameter tuning experiments, compare results, and visualize metrics. Which SageMaker feature should they use?
105A company has deployed a model on SageMaker for real-time inference. The endpoint is experiencing high latency during traffic spikes. Which action should the company take to reduce latency?
106A data scientist is using SageMaker to train a model with a custom algorithm. The training script uses TensorFlow and runs on GPU instances. The training job fails with 'CUDA_ERROR_OUT_OF_MEMORY'. What is the most likely cause?
107A company wants to use SageMaker to host multiple models behind a single endpoint to reduce costs. Which SageMaker feature should they use?
108A machine learning team is using SageMaker to train a model. They want to ensure that the training data is encrypted at rest in the S3 bucket and that the data is also encrypted during transit. Which configuration should they use?
109A company is using SageMaker to train a model with a large dataset that is stored in S3. The training job is taking a long time due to high I/O latency. The team has already converted the data to RecordIO format. What should they do next to reduce I/O latency?
110Which TWO of the following are benefits of using SageMaker Managed Spot Training? (Select TWO.)
111Which THREE of the following are valid ways to deploy a model using SageMaker? (Select THREE.)
112An IAM policy attached to a SageMaker notebook role is shown in the exhibit. A data scientist is trying to run a training job from the notebook, but the job fails with an access denied error. The training job needs to read data from 'my-bucket' and write output to 'my-bucket'. What is the most likely cause of the failure?
113A machine learning engineer needs to deploy a TensorFlow model to a SageMaker endpoint. The model expects a specific input format. The engineer has the model artifacts stored in an S3 bucket. Which step is REQUIRED to deploy the model?
114A company wants to use Amazon Rekognition to detect objects in images stored in an S3 bucket. The images are uploaded by users. Which IAM policy statement is necessary to allow Rekognition to read from the bucket?
115A team has deployed a SageMaker endpoint for a sentiment analysis model. The model was trained on text data from social media. After deployment, the team notices that the model's accuracy has dropped significantly after 3 months. Which action should the team take to detect and address this issue?
116Which TWO actions can reduce inference latency for a SageMaker real-time endpoint? (Choose 2.)
117Which TWO SageMaker features can be used to monitor and debug training jobs? (Choose 2.)
118Refer to the exhibit. A data scientist is training a PyTorch model on a SageMaker ml.p3.2xlarge instance (16 GB GPU memory). The training fails with the shown error. Which change should the scientist make to resolve the error?
119Refer to the exhibit. An administrator has attached this IAM policy to a user. The user tries to start a SageMaker training job that uses a custom Docker image from Amazon ECR. The training job fails with an access denied error. What is the MOST likely reason?
120A machine learning engineer is deploying a model using Amazon SageMaker. The model is a PyTorch model that performs real-time inference with low latency requirements. The engineer wants to use automatic scaling based on the number of concurrent requests. Which SageMaker feature should be used to achieve this?
121A company is using Amazon SageMaker to build a binary classification model. The dataset is highly imbalanced, with 95% negative class and 5% positive class. Which technique should be used to address the class imbalance?
122A data scientist is training a model using Amazon SageMaker and wants to track hyperparameter tuning jobs, training jobs, and model metrics. The team also needs to compare experiments visually. Which AWS service should be used?
123A company is using Amazon SageMaker to train a model. The training data is stored in an S3 bucket in a different AWS account. Which IAM policy configuration is required to allow SageMaker to access the data?
124A machine learning engineer is deploying a model on SageMaker and needs to ensure that the endpoint can handle a sudden spike in traffic. The engineer expects traffic to increase by 10x during a promotional event. Which scaling strategy should be used?
125A data scientist is using Amazon SageMaker to train a model and wants to use a custom Docker container for training. The container requires access to a private Amazon ECR repository. Which IAM role configuration is needed?
126A company wants to use Amazon SageMaker to train a model using data that is updated daily. The training data is stored in an S3 bucket, and the team wants to automate the training process whenever new data arrives. Which AWS service should be used to trigger the SageMaker training job?
127A data scientist is training a model using Amazon SageMaker and wants to reduce the training time. The training job uses a single GPU instance. Which THREE actions can reduce training time?
128A company wants to deploy a machine learning model on Amazon SageMaker and needs to monitor the model's performance in production. Which TWO AWS services can be used to set up monitoring?
129A data scientist is deploying a PyTorch model to Amazon SageMaker for real-time inference. The model runs on a large instance but inference latency is too high. Which action is MOST likely to reduce latency without sacrificing accuracy?
130A team is using Amazon SageMaker to train a linear regression model on a dataset with 10 features. After training, they notice the model has high bias. Which action is MOST likely to reduce bias?
131A data scientist is using Amazon SageMaker to train a model. The training job uses a custom Docker image stored in Amazon ECR. The training job fails with an error 'CannotPullContainerError'. Which TWO actions should the data scientist take to resolve this issue? (Choose TWO.)
132A company is deploying a machine learning model using Amazon SageMaker. To reduce costs, they want to use SageMaker Managed Spot Training. Which THREE conditions must be met for the training job to use spot instances? (Choose THREE.)
133A data engineer is building a data pipeline for a machine learning project using Amazon SageMaker. The raw data is stored in Amazon S3. Which TWO steps are essential to ensure data privacy and security before training? (Choose TWO.)
134An IAM policy attached to a SageMaker execution role is shown in the exhibit. When a data scientist tries to create a training job that writes logs to CloudWatch Logs, the job fails. What is the MOST likely reason?
135An engineer sees the error in the exhibit when trying to deploy a model from a model registry in SageMaker. What is the MOST likely cause?
136A data scientist needs to deploy a trained model to Amazon SageMaker for real-time inference. The model is stored as a .tar.gz file in Amazon S3. Which AWS service is used to create a SageMaker endpoint?
137A machine learning engineer is using Amazon SageMaker to train a model. The training job is taking too long. The engineer suspects the data loading is a bottleneck. Which action would MOST effectively diagnose the issue?
138A company is using Amazon SageMaker to run a hyperparameter tuning job. The tuning job uses Bayesian optimization. Which THREE statements about Bayesian optimization are correct? (Choose THREE.)
139A data scientist wants to deploy a PyTorch model for real-time inference. Which SageMaker deployment option provides the lowest latency for single-digit millisecond responses?
140A team is training a large NLP model using SageMaker. The training job fails with an OutOfMemory error. The instance type is ml.p3.2xlarge with 61 GB GPU memory. Which action should the team take to resolve the issue without changing the model architecture?
141A data scientist uses SageMaker to train a model. The training job takes 10 hours, but the team needs to reduce costs. Which approach is MOST cost-effective?
142A company deploys a SageMaker endpoint for real-time inference. After a week, the response latency increases from 50 ms to 500 ms. CPU utilization is at 30%. What is the most likely cause?
143A team needs to automatically retrain a model every week using new data. Which SageMaker feature is designed to schedule and automate this workflow?
144A model deployed on a SageMaker endpoint is producing predictions that are consistently biased against a certain demographic. Which step should the team take FIRST to address this issue?
145A machine learning team is using SageMaker to train a model with a custom Docker container. The training script runs locally but fails on SageMaker with a 'Permission denied' error when writing to /opt/ml/model. What is the likely cause?
146A company wants to monitor SageMaker endpoints for data drift. Which TWO services can be used together to detect and alert on drift?
147A data scientist needs to deploy a model with a custom inference container. Which THREE requirements must the container meet for SageMaker hosting?
148A machine learning team is using SageMaker Pipelines to orchestrate a multi-step workflow. The pipeline fails with a 'ThrottlingException' when submitting a training job. Which TWO actions can reduce the likelihood of throttling?
149A data scientist is training a deep learning model on Amazon SageMaker using the built-in Object Detection algorithm. The training job is failing with a 'ResourceLimitExceeded' error when trying to launch multiple GPU instances. Which of the following is the MOST likely cause?
150A machine learning team is deploying a real-time inference endpoint on Amazon SageMaker for a model that requires low latency (<100 ms). The model is a PyTorch model with custom pre- and post-processing logic. The team uses a SageMaker Model with a custom inference container. After deployment, they observe that the endpoint takes over 500 ms for the first request, but subsequent requests are fast (~50 ms). What is the MOST likely cause?
151A data scientist is performing hyperparameter tuning using Amazon SageMaker Automatic Model Tuning (AMT). The job uses a random search strategy. After 20 training jobs, the best objective metric value has plateaued. The data scientist wants to explore more of the hyperparameter space. Which action should the data scientist take?
152A machine learning engineer is deploying a model to an Amazon SageMaker endpoint. The model requires GPU for inference. Which instance type should be selected?
153A data scientist is using Amazon SageMaker to train a model with a custom Docker container. The training script reads data from an S3 bucket and writes the model artifact to an S3 bucket. The training job fails with a 'NoSuchKey' error. What is the MOST likely cause?
154A company is using Amazon SageMaker to train a machine learning model. The training job is configured to use the File mode to download data from S3 to the training instances. The training data is stored in a single S3 bucket with multiple prefixes. Which TWO actions are required to ensure the training job can access the data? (Choose TWO.)
155A company is deploying a machine learning model using SageMaker. The model is a PyTorch model that requires GPU for inference. The company wants to minimize costs while ensuring low latency. Which instance type should be used for the SageMaker endpoint?
156A company uses SageMaker to host a real-time inference endpoint. The endpoint is receiving a large number of requests, but the latency is higher than expected. The data scientist observes that the CPU utilization is low but memory utilization is high. Which action should be taken to reduce latency?
157A data scientist is using SageMaker Ground Truth to create a labeled dataset for object detection. After the labeling job completes, the scientist notices that the output manifest file contains incorrect labels. What is the most efficient way to correct these labels?
158A company is using SageMaker to host a model that performs real-time fraud detection. The model receives high request volumes with occasional spikes. The company wants to ensure that the endpoint can handle spikes without throttling while minimizing cost. Which scaling strategy should be used?
159A data scientist is using SageMaker to train a model using the built-in XGBoost algorithm. The training job fails with the error 'AlgorithmError: Framework error: No module named 'xgboost''. What is the most likely cause?
160A data scientist is using SageMaker to train a model and wants to track experiments, including hyperparameters and metrics. Which TWO actions should the scientist take to set up experiment tracking? (Choose TWO.)
161A data scientist is training a model using SageMaker and wants to use spot instances to reduce costs. Which THREE considerations should the scientist evaluate? (Choose THREE.)
162A data scientist wants to deploy a PyTorch model for real-time inference with latency under 100 ms. Which AWS service is most suitable?
163A data scientist needs to run a one-time SQL query on a large dataset in S3 to create a training dataset. The query involves aggregations and joins. Which service is most suitable?
164A data scientist needs to store and version machine learning models, along with metadata such as hyperparameters and metrics. Which AWS service is designed for this purpose?
165A company is using SageMaker to train a linear learner algorithm. The training log shows that the algorithm converges but the final loss is still high. Which change is most likely to improve the model?
166A company wants to automate the retraining of a model weekly using new data. The training script is in a SageMaker notebook. Which implementation is most maintainable?
167A data scientist needs to select a model training infrastructure that supports distributed training across multiple GPUs and provides automatic model parallelism. Which TWO AWS services should the scientist consider?
168A training job log shows this error. The training instance is an ml.m5.large with 8 GB EBS storage. The training data is 500 MB, and the model size is expected to be 200 MB. What is the most likely cause?
169A company is using Amazon SageMaker to train and deploy a fraud detection model. The model is a gradient boosting machine (GBM) trained on a dataset with 10 million rows and 50 features. The training job runs on an ml.m5.2xlarge instance with 8 vCPUs and 32 GB memory. The training completes successfully, and the model is deployed to a real-time endpoint. After deployment, the inference latency is around 200 ms per request, which is acceptable. However, after a week, the company observes that latency increases to over 1 second during peak hours (12:00-13:00 UTC). CloudWatch metrics show CPU utilization on the endpoint instance reaches 95% during these peaks. The endpoint is configured with a single ml.m5.large instance. The company wants to maintain latency under 500 ms during peak hours without incurring unnecessary cost during off-peak hours. Which solution should the company implement?
170A data scientist is training a TensorFlow model on a single GPU instance. The training is taking too long. Which AWS service should be used to reduce training time by distributing the workload across multiple GPUs?
171A machine learning team is using Amazon SageMaker to train a model with a custom algorithm packaged in a Docker container. The training job fails with the error 'Error: Unable to locate sagemaker-training toolkit.' What is the MOST likely cause?
172A data scientist needs to perform hyperparameter optimization for a model. Which AWS service provides built-in hyperparameter tuning jobs?
173A company is deploying a model to an Amazon SageMaker endpoint for real-time inference. The model requires a GPU for low-latency predictions. Which instance type should be chosen?
174A machine learning engineer is using AWS Step Functions to orchestrate a SageMaker training job followed by a Lambda function for post-processing. The training job completes successfully, but the Lambda function fails with a timeout error. What is the MOST likely cause?
175A company wants to track and compare metrics from multiple machine learning experiments. Which Amazon SageMaker feature should be used?
176A company is using SageMaker to host a model for real-time inference. They notice that the endpoint's latency increases over time. The model is stateless and the inference code does not log any errors. What is the MOST likely cause?
177Which TWO AWS services can be used to deploy a machine learning model for serverless inference? (Choose 2.)
178A financial services company uses Amazon SageMaker to train a fraud detection model. The training data is stored in an S3 bucket encrypted with AWS KMS. The SageMaker training job is configured to use a custom Docker container that reads data from S3 and writes model artifacts back to S3. The training job fails with the error: 'Unable to write model artifact to s3://my-bucket/output/model.tar.gz. Access Denied.' The IAM role used by the training job has the following permissions: s3:GetObject and s3:PutObject on the bucket, and kms:Decrypt on the KMS key. The training job is not using a VPC. What is the MOST likely cause of the failure?
179A media company uses SageMaker to deploy a real-time inference endpoint for content recommendation. The model is a PyTorch model that uses GPU. The endpoint is deployed with an ml.p3.2xlarge instance. Over time, the endpoint's latency increases significantly during peak hours. The company has enabled auto scaling based on CPU utilization. However, the latency spikes occur even when CPU utilization is low. The model is stateless and the inference code is efficient. What is the MOST likely cause of the latency spikes?
180A startup is using SageMaker to train a model using the built-in XGBoost algorithm. The training job runs successfully but the resulting model performs poorly on the test data. The data scientist suspects overfitting. The training data is relatively small (10,000 rows). Which action should be taken to reduce overfitting?
181A data scientist is training a deep learning model on a large dataset using Amazon SageMaker. The training job is taking too long and the scientist wants to reduce the training time by distributing the workload across multiple GPUs. Which SageMaker feature should be used to achieve this?
182A data science team is using Amazon SageMaker to train a model. The training job is failing with an 'OutOfMemory' error. The team is using a p3.2xlarge instance with 61 GB of memory. They need to resolve this issue as quickly as possible. Which action should they take?
183A company uses Amazon SageMaker to deploy a model for real-time predictions. The model is updated weekly. The company wants to ensure that the new model version is gradually rolled out to a small percentage of traffic before full deployment, and that it can be rolled back quickly if issues are detected. Which deployment strategy should be used?
184A machine learning engineer is building a pipeline to preprocess data and train a model using Amazon SageMaker. The data is stored in Amazon S3 and the preprocessing step is computationally intensive. The engineer wants to minimize costs while ensuring that the preprocessing step does not fail due to instance termination. Which instance type should be used for the preprocessing step?
185A data scientist is using Amazon SageMaker to build a custom training algorithm. The algorithm requires a specific library that is not included in the default SageMaker containers. The scientist wants to create a custom container that includes this library. Which TWO steps are required? (Choose TWO.)
186Refer to the exhibit. An IAM policy is attached to an IAM role used by a SageMaker training job. The training job fails with an access denied error when trying to write model artifacts to an S3 bucket. What is the most likely cause?
187A company has deployed a machine learning model on Amazon SageMaker for real-time inference. The endpoint uses a single ml.c5.xlarge instance. Recently, the traffic has increased, and the endpoint is returning HTTP 503 (Service Unavailable) errors during peak hours. The CloudWatch metrics show that the CPU utilization is consistently above 90% during peak times, and the Invocations metric shows that requests are being throttled. The data science team has already optimized the model to reduce inference time by 20%, but the errors persist. The company needs to resolve the issue without increasing costs significantly. Which course of action should be taken?
188A machine learning engineer is designing an automated ML pipeline for training and deploying models. The pipeline must include data validation, model training, hyperparameter tuning, and model deployment. The engineer wants to use AWS services that integrate well and provide version control. Which THREE services should be combined to achieve this? (Choose THREE.)
189A company is using Amazon SageMaker to host a real-time inference endpoint for a natural language processing model. The endpoint is configured with an ml.m5.large instance. After deployment, the company observes that the inference latency is higher than expected, and the endpoint is experiencing CPU utilization near 100% during peak hours. The model is a PyTorch model that uses a transformer architecture. The company wants to reduce latency without increasing cost significantly. Which approach should the company take?
190A company is deploying a machine learning model using AWS Lambda for real-time inference. The model is a large ensemble model that takes approximately 500 MB of memory. The Lambda function is configured with 1024 MB of memory and a timeout of 15 seconds. The company observes that the function frequently times out during inference. The company wants to keep using Lambda for its serverless benefits. Which solution should the company implement to reduce inference time?
191A company is using Amazon SageMaker Ground Truth to build a training dataset for an image classification model. The company has a large number of unlabeled images stored in Amazon S3. The data science team wants to use a private workforce consisting of internal employees to label the images. The team creates a labeling job with a private workforce. After starting the job, the team notices that the labeling tasks are not being assigned to any workers. The workers have been added to the private workforce and have received their login credentials. What is the MOST likely cause of this issue?
192A company is using Amazon SageMaker to train a deep learning model. The training job uses a script that reads data from Amazon S3 using the SageMaker SDK's `s3_input` method. The training job runs on a single ml.p3.2xlarge instance. The data scientist notices that the GPU utilization is very low during training, often below 20%. The training dataset is large, approximately 50 GB, stored as TFRecord files in S3. What is the MOST likely cause of low GPU utilization?
193A company uses Amazon SageMaker to deploy a model for real-time inference. The model is a linear regression model that was trained using the SageMaker built-in Linear Learner algorithm. The endpoint is configured with an ml.m5.large instance. After deployment, the company notices that the endpoint returns incorrect predictions. The training data was normalized, but the inference requests send raw feature values without normalization. What should the company do to fix the issue?
194A company is using Amazon SageMaker to train a model using a custom Docker container. The training script writes model artifacts to the `/opt/ml/model` directory. The training job completes successfully, but the model artifacts are not uploaded to the S3 output path specified in the training job. The company has verified that the SageMaker execution role has the necessary S3 permissions. The Docker container is built using a base image that is not one of the official SageMaker Docker images. What is the MOST likely reason for the failure to upload model artifacts?
195A machine learning engineer is deploying a real-time inference service on Amazon SageMaker. The model must serve predictions with minimal latency, and the engineer needs to monitor the endpoint for data quality drift and model quality drift. The engineer also wants to receive automated alerts when drift exceeds a threshold. Which two actions should the engineer take to meet these requirements? (Choose two.)
196A company is running a hyperparameter tuning job on Amazon SageMaker to optimize a neural network. The tuning job is using Bayesian optimization and has completed 20 training jobs. The best model so far has an accuracy of 92%. The company wants to reduce the cost of the tuning job without significantly impacting the final model performance. Which action should a machine learning engineer take?
197A machine learning engineer is deploying a real-time inference endpoint on Amazon SageMaker. The model artifacts are stored in an Amazon S3 bucket that is encrypted with AWS Key Management Service (KMS) using a customer managed key. The SageMaker execution role has permissions to access the S3 bucket but does not have any KMS permissions. When the engineer attempts to create the endpoint, the deployment fails with an error indicating that the model data cannot be downloaded. Which action should the engineer take to resolve this issue?
198A machine learning engineer is deploying a model to an Amazon SageMaker endpoint that will receive a sudden spike in traffic. The model is hosted on a single ml.m5.large instance and the engineer wants to ensure the endpoint can automatically scale to handle the increased load without manual intervention. The engineer also wants to minimize latency. Which solution should the engineer implement?
199A data scientist is training a deep learning model using Amazon SageMaker. The training job uses a large dataset stored in Amazon S3 and runs on multiple GPU instances. The data scientist notices that the training job is slow because the GPUs are underutilized, and the training data loading is a bottleneck. The dataset consists of many small files. Which solution will improve GPU utilization and reduce training time?
200An ML engineer trains a model on Amazon SageMaker using a dataset stored in Amazon S3. The training job completes successfully, but the engineer notices that the model artifacts are not saved to the specified S3 output path. The training job's IAM role has full access to the input S3 bucket but only read access to the output bucket. What is the most likely cause of the missing model artifacts?
201A machine learning engineer is deploying a real-time fraud detection model on Amazon SageMaker. The model is a PyTorch neural network that requires GPU acceleration. The engineer needs to ensure that the endpoint automatically scales the number of instances based on the number of invocations per instance. The engineer has already created the SageMaker model and endpoint configuration. Which scaling policy should the engineer configure to meet these requirements?
202A data scientist has trained a linear regression model using Amazon SageMaker and wants to deploy it for real-time predictions. The model expects a CSV input with a single record per request. The data scientist needs to create an endpoint that can be invoked with low latency. Which SageMaker deployment option should they use?
203A company uses Amazon SageMaker to train a deep learning model for image classification. The training job is using a single p3.2xlarge instance and is taking too long. The dataset is stored in Amazon S3 and consists of 1 million images. The company wants to reduce training time without changing the model architecture. Which of the following should be done to achieve this?
204A data scientist is training a deep learning model using Amazon SageMaker. The training data is stored in Amazon S3 and is approximately 2 TB in size. The training job is taking a long time to start because the data is being downloaded from S3 to the training instance's Amazon EBS volume. The data scientist wants to reduce the startup time and improve training performance. Which two actions should the data scientist take? (Choose two.)
205A machine learning engineer is using Amazon SageMaker to train a model. The training job is configured to use a Spot instance to reduce costs. The engineer wants to ensure that the training job can resume from the last saved checkpoint if the Spot instance is interrupted. Which SageMaker feature should the engineer use?
206A machine learning engineer is using Amazon SageMaker to train a model on a dataset that contains sensitive financial information. The company requires that the data be encrypted at rest and in transit, and that the training job uses a specific AWS KMS customer managed key for encryption. The engineer has configured the training job to use the KMS key for the output volume. What additional step must the engineer take to ensure that the training data in Amazon S3 is encrypted with the same KMS key?
207A machine learning engineer must retrain a model nightly on data that grows by roughly 4 GB per day. The current training job reads 90 GB of CSV files from Amazon S3 and the input pipeline is the bottleneck, with GPU utilization below 20%. The engineer wants to reduce training time without changing the algorithm. Which action should the engineer take?
208A machine learning engineer has trained a model using Amazon SageMaker and wants to deploy it for real-time inference. The model artifacts are stored in Amazon S3. The engineer needs to create a SageMaker endpoint that can serve predictions with low latency. Which step is required to deploy the model?
209A company has deployed a real-time inference endpoint on Amazon SageMaker using a custom container. The endpoint is experiencing high latency and occasional timeouts during peak traffic. The model is CPU-bound and the container is configured with one CPU core and 2 GB of memory. The company wants to reduce latency and improve throughput without changing the model. Which action should a machine learning engineer take?
210A machine learning engineer has trained a SageMaker model and registered it in the Amazon SageMaker Model Registry. Compliance requires that every production endpoint only serve model versions that have been reviewed and approved, and that no unreviewed artifact can ever be deployed to production. The engineer needs to wire the deployment pipeline so that only approved model versions are deployed. Which approach meets these requirements with the LEAST operational effort?
211A data scientist has a trained scikit-learn model artifact in Amazon S3 and needs a fully managed HTTPS endpoint that returns predictions in real time with automatic patching of the underlying infrastructure. The team has no Docker expertise and wants the fewest possible steps. Which SageMaker deployment approach should the data scientist use?
212A data scientist is using Amazon SageMaker to train a model on a large dataset. The training job is taking longer than expected, and the data scientist suspects that the input data pipeline is the bottleneck. The dataset consists of many small files stored in Amazon S3. The data scientist wants to improve the training throughput by optimizing data input. Which action should the data scientist take?
213A machine learning engineer is using Amazon SageMaker to train a model. The training job is taking longer than expected, and the engineer suspects that the data loading is the bottleneck. The training data is stored in Amazon S3 as many small files. Which of the following is the most efficient way to improve data loading performance?
214A company runs a real-time SageMaker endpoint backed by an Auto Scaling policy. Traffic is highly variable: it is near zero overnight and spikes sharply every weekday at 09:00. Users report timeouts during the first few minutes of each spike because the endpoint scales out too slowly. The engineer must reduce cold-start scale-out latency while keeping costs low during idle periods. Which change is MOST effective?
215A company operates a SageMaker real-time endpoint that serves a fraud-detection model. Compliance requires that predictions be explainable for individual transactions, and the data science team wants to detect when incoming feature distributions shift away from the training baseline. The team wants to use SageMaker built-in capabilities rather than custom code. Which two capabilities should they enable? (Choose two.)
216A machine learning engineer is deploying a scikit-learn model to a SageMaker endpoint. The model requires a Python preprocessing function that transforms raw JSON input into a NumPy array before prediction. The engineer wants to minimize inference latency and avoid writing a custom Docker container. Which SageMaker feature should be used?
217A machine learning team wants to automatically detect when a deployed SageMaker model's predictions start to drift from the original training data distribution. They need to monitor feature distributions and receive alerts when significant deviations occur. Which SageMaker feature should they use?
218A data scientist is using Amazon SageMaker to train a model. The training job must write model artifacts to an Amazon S3 bucket. The data scientist wants to ensure that the training job can access the S3 bucket securely without embedding long-term AWS credentials in the training container. Which mechanism should the data scientist use to grant the training job access to the S3 bucket?
219A company has deployed a real-time inference endpoint on Amazon SageMaker using a custom container. The endpoint receives a steady stream of requests but occasionally experiences spikes in traffic. The company wants to automatically scale the endpoint based on the number of invocations per instance to handle the load efficiently while minimizing costs. Which scaling policy should the company configure?
220A company has a production SageMaker real-time endpoint serving a fraud detection model. The model was trained on transaction data from several months ago, and the fraud patterns have shifted. The team wants to detect when the live request data diverges from the training data distribution so they can trigger retraining, without building custom monitoring infrastructure. Which solution should they implement?
221A machine learning engineer has trained a model using Amazon SageMaker and wants to deploy it to a real-time endpoint. The model artifacts are stored in Amazon S3. The engineer needs to create a SageMaker model, configure an endpoint configuration, and then create the endpoint. Which sequence of API calls should the engineer use to deploy the model?
222A machine learning engineer has deployed a real-time SageMaker endpoint backed by three production variants. The engineer wants to shift 10% of live inference traffic from the current model to a newly trained model that is already registered and hosted on the same endpoint, while keeping the remaining 90% on the current model. The change must take effect immediately without redeploying the endpoint. Which action should the engineer take?
223A company is deploying a real-time inference endpoint on Amazon SageMaker. They need to ensure the endpoint can automatically scale based on incoming traffic and that the scaling policy targets a specific utilization metric. Which two actions should they take? (Choose two.)
224A company wants to deploy a trained SageMaker model to production with the ability to shift a small percentage of live traffic to a new model version for testing. The company needs to monitor the new model's performance without affecting the majority of users. Which deployment strategy should be used?
225A data scientist is using Amazon SageMaker to run a hyperparameter tuning job for a deep learning model. The tuning job is configured to maximize the validation accuracy. After several training jobs, the data scientist notices that the tuning job is not exploring a wide range of hyperparameter values and is converging to a suboptimal configuration. Which action should the data scientist take to improve the tuning job's exploration?
226A machine learning engineer needs to store and version large datasets used for training models in Amazon SageMaker. The datasets are updated frequently, and the engineer wants to track changes and ensure reproducibility of training jobs. Which AWS service should they use?
227A machine learning engineer has a SageMaker training job that reads training data from an Amazon S3 bucket in the same AWS account. The job is running in a VPC and needs to download the data quickly. The engineer has configured the training job to use a private subnet with a NAT gateway. Data transfer is slow. Which action will improve the data transfer speed?
228A company is deploying a machine learning model to a SageMaker endpoint. The model requires a GPU for inference and must handle occasional bursts of traffic. The company wants to minimize cost while ensuring that the endpoint scales automatically. Which scaling policy should the company use?
229A company uses Amazon SageMaker to train a recommendation model. The training dataset grows daily and is stored in Amazon S3. The data science team wants to avoid re-downloading the entire dataset for each training job, so they configure SageMaker Pipe mode with an augmented manifest file. Which statement accurately describes how Pipe mode streams data to the training algorithm?
230A company has a real-time inference endpoint on Amazon SageMaker that uses a custom container. The endpoint is experiencing high latency during peak traffic. The company wants to reduce latency by enabling GPU acceleration and ensuring that the model server can handle multiple concurrent requests efficiently. The custom container currently runs a Flask application that processes one request at a time. Which solution will meet these requirements with the LEAST operational overhead?
231A machine learning engineer is deploying a real-time inference endpoint using Amazon SageMaker. The model is a TensorFlow model that requires GPU acceleration for low-latency predictions. The engineer must ensure that the endpoint automatically scales based on the number of invocations per instance, and that it scales in a cost-effective manner. The endpoint will receive traffic that varies throughout the day. Which solution should the engineer implement?
232A machine learning engineer is using SageMaker Pipelines to orchestrate a training workflow. The pipeline includes a processing step that outputs a dataset to S3, followed by a training step that consumes that dataset. The engineer wants to ensure that the training step uses the output from the processing step. How should the engineer configure the training step?
233A data scientist is using Amazon SageMaker to train a model on a large dataset stored in Amazon S3. The training job is expected to run for several hours. The data scientist wants to reduce the cost of the training job without compromising model accuracy. The dataset is already in S3 in CSV format. Which two actions should the data scientist take to reduce training cost? (Choose two.)
234A machine learning engineer is deploying a real-time inference endpoint on Amazon SageMaker. The model must handle occasional spikes in traffic and the engineer wants to minimize cost during periods of low traffic. The engineer configures automatic scaling for the endpoint using Application Auto Scaling. Which scaling policy configuration will best meet these requirements?
235A machine learning engineer is deploying a model to a SageMaker endpoint. The model requires a custom inference script that preprocesses input data. The engineer wants to minimize cold start latency and ensure the endpoint can handle sudden traffic spikes. Which deployment strategy should the engineer use?
236A data scientist is using Amazon SageMaker to train a model and needs to monitor the training job's resource utilization and detect anomalies such as high CPU usage or memory leaks. The data scientist wants to receive alerts when the training job's CPU utilization exceeds 90% for more than 5 minutes. Which AWS service should be used to set up these alerts?
237A machine learning engineer has trained a SageMaker model and wants to deploy it for real-time inference. The model artifacts are stored in Amazon S3, and the engineer needs to ensure that the endpoint can automatically scale based on incoming traffic. The engineer also wants to minimize latency for the inference requests. Which deployment option should the engineer use?
238A company has a SageMaker endpoint that serves a model for real-time predictions. The model was trained on data that includes categorical features. The company wants to monitor the endpoint for data drift and model quality. The company has enabled SageMaker Model Monitor with a baseline. After a week, the company notices that the data drift monitor is flagging violations, but the model's predictions still seem accurate. What is the most likely cause of the violations?
239A company has deployed a SageMaker endpoint for a fraud detection model. The model was trained on historical data and is performing well. However, the company notices that the model's predictions are becoming less accurate over time due to changes in customer behavior. The company wants to detect this degradation and receive alerts when the model's performance drops below a threshold. Which SageMaker feature should the company use?
240A company is using Amazon SageMaker to train a deep learning model. The training job is taking longer than expected, and the data scientist suspects that the input data pipeline is the bottleneck. The training data consists of many small files stored in Amazon S3. The data scientist wants to improve training throughput by optimizing data loading. Which solution will MOST effectively reduce the training time?
241A machine learning engineer is deploying a model to a SageMaker endpoint and needs to monitor it for data drift and model quality. The engineer wants to receive alerts when the model's predictions deviate from expected behavior. Which two actions should the engineer take to set up automated monitoring and alerting? (Choose two.)
242A machine learning engineer needs to deploy a trained scikit-learn model to a SageMaker endpoint for real-time inference. The model artifacts are stored in Amazon S3. The engineer wants to use the SageMaker Python SDK to deploy the model with minimal code. Which method should the engineer use?
243A machine learning engineer is tuning a SageMaker XGBoost model. They configure a hyperparameter tuning job with a maximum of 30 training jobs and set the objective metric as 'validation:auc'. After the tuning job completes, they want to deploy the model that achieved the best validation AUC. Which step will identify the best training job?
244A company runs a SageMaker real-time endpoint that serves a fraud detection model. The security team requires that all inference requests and responses be recorded for audit, and the data science team wants to detect feature drift in the incoming request data over time. The endpoint currently has no monitoring configured. Which two actions should the engineer take to meet both requirements? (Choose two.)
245A data scientist is using Amazon SageMaker to train a model and wants to ensure that the training job can access the training data stored in an Amazon S3 bucket. The S3 bucket is in the same AWS account and Region. What is the SIMPLEST way to grant the necessary permissions?
246A company is using Amazon SageMaker to run a hyperparameter tuning job for an XGBoost model. The tuning job must efficiently search a large hyperparameter space while minimizing training time and cost. Which two strategies should the company use to achieve these goals? (Choose two.)
247A data scientist is preparing a training script that will run in script mode on Amazon SageMaker. The script must be able to read the training dataset that SageMaker copies from Amazon S3, and the data scientist wants to avoid hard-coding any bucket names or local paths. Which approach should the data scientist use?
248A machine learning engineer needs to deploy a model on Amazon SageMaker for real-time inference. The model requires a custom inference script that preprocesses input data. The engineer wants to ensure that the endpoint can automatically scale based on traffic and that the model can be updated without downtime. Which deployment approach should the engineer use?
249A company has a SageMaker training job that reads a large dataset from Amazon S3. The data is stored in a single S3 bucket in the same AWS Region. The training job is taking a very long time to start because SageMaker is downloading the entire dataset to the training instance's local disk before training begins. The data is in CSV format and is used in every epoch. The company wants to reduce the startup time and allow the training job to start training faster. What should the company do?
250A data scientist is using Amazon SageMaker to train a model with a built-in algorithm. They want to monitor the training job's resource utilization, such as CPU and GPU utilization, in real-time. Which AWS service can they use to view these metrics?
251A data scientist is using Amazon SageMaker to train a model and needs to monitor the training job's resource utilization and debug performance issues. The training job is running on multiple instances and uses a custom training script. The data scientist wants to collect metrics such as CPU utilization, GPU utilization, and memory usage, and also wants to capture tensor-level information for debugging. Which two actions should the data scientist take to meet these requirements? (Choose two.)
252A machine learning engineer is deploying a real-time inference endpoint on Amazon SageMaker for a model that requires GPU acceleration. The endpoint must automatically scale based on the number of invocations per instance. The engineer wants to minimize cost while ensuring that the endpoint can handle sudden spikes in traffic. Which scaling policy configuration should be used?
253A company is using Amazon SageMaker to train a model. The training data is stored in an Amazon S3 bucket. The company wants to ensure that the training job can access the data securely without embedding AWS credentials in the training script. The company also wants to follow the principle of least privilege. What should the company do?
254A company is using Amazon SageMaker to train a model and wants to ensure that the training job can access data in an Amazon S3 bucket and write model artifacts to another S3 bucket. The company also wants to monitor the training job's resource utilization. Which two actions should be taken to meet these requirements? (Choose two.)
Deep-dive questions
The most-searched questions in this domain — detailed explanations, worked examples, full answer breakdowns.
Be able to configure and troubleshoot SageMaker training and hosting: pick the right endpoint type, wire IAM roles to S3 correctly, and read AWS CLI errors. The single most important thing is verifying permissions and data paths before blaming the model code.
The Courseiva MLS-C01 question bank contains 254 questions in the Machine Learning Implementation and Operations domain, covering the 20% of the exam attributed to this domain in the official Amazon Web Services blueprint. Click any question to see the full explanation and answer breakdown.
Start with a 10-question focused session to identify your baseline accuracy in this domain. Read every explanation — even for questions you answer correctly — to understand the reasoning. Once you score consistently above 80%, move to a 20–30 question session to confirm depth before moving to the next domain.
Yes — the session launcher on this page draws questions exclusively from the Machine Learning Implementation and Operations domain. Choose 10, 20, 30, or 50 questions for a focused session, or click individual questions to review them one by one.
Save your results, see per-domain analytics, and get readiness scores — free, for every certification.
Sign Up FreeFree forever · Every certification included