MLA-C01 · domain
scenario questions
Practise AWS Certified Machine Learning Engineer Associate MLA-C01 scenario questions practice questions — original exam-style scenarios with answer choices, explanations, and analysis of common mistakes.
Focused practice
Practice scenario questions questions
Scored sessions drawing only from this domain — pick a length below.
Start 20-question practice test →What this domain covers
What to know about scenario questions
scenario questions questions test whether you can apply the concept in context, not just recognise a definition.
How the topic appears in realistic exam-style scenarios.
Which detail in the question changes the correct answer.
How to eliminate plausible but wrong options.
How to connect the question back to the wider exam objective.
Watch out for
Common scenario questions exam traps
- ▸Answering from memory before reading the full scenario.
- ▸Missing a constraint such as cost, availability, security, scope or command context.
- ▸Choosing a broad answer when the question asks for the most specific fix.
- ▸Ignoring why the wrong options are tempting.
Question index
All scenario questions questions (665)
Click any question to see the full explanation, or start a practice session above.
A machine learning engineer is using Amazon SageMaker Feature Store to manage features for a fraud detection model. The engineer needs to ensure that the feature group can serve both batch and real-time predictions. The feature group is configured with an online store enabled. Which additional configuration is required to support batch predictions?
Medium2Refer to the exhibit. A data scientist used a SageMaker training job with a custom Scikit-learn script. The training job failed with the error shown. What is the most likely cause of this failure?
Hard3A company is building a CI/CD pipeline for ML models using AWS CodePipeline and SageMaker. The pipeline should include steps to automatically retrain, evaluate, and deploy models. Which THREE components are essential for this pipeline? (Choose three.)
Hard4A team needs to deploy a model that has compliance requirements to log all inference requests and responses for auditing. The model will be served using a real-time endpoint. How can they achieve this without custom code?
Hard5A financial institution is training a fraud detection model using SageMaker. The dataset is highly imbalanced, with only 0.1% fraudulent transactions. The team wants to use SageMaker Automatic Model Tuning to find the best hyperparameters. They notice that the tuning job spends most of its time on configurations that predict all transactions as non-fraudulent. Which hyperparameter should they tune to directly address this issue?
Hard6A machine learning engineer is using Amazon SageMaker Data Wrangler to prepare a dataset with a categorical feature that has over 5,000 distinct values (high cardinality). The engineer needs to transform this feature into a form suitable for a gradient boosting model while preserving as much information as possible. Which transform should be applied?
Hard7An MLOps team is designing a SageMaker Pipeline to automate model retraining. The pipeline must: (1) run training only if new training data is available, (2) register the model in SageMaker Model Registry only if evaluation metrics exceed a threshold, (3) deploy the approved model to a staging endpoint automatically. Which THREE steps should they include? (Choose THREE.)
Hard8A company wants to run inference on a large dataset stored in S3 using a pre-trained model. The inference can tolerate latency from minutes to hours, and they want a fully managed solution that autoscales to handle large volumes. Which SageMaker inference option is most suitable?
Medium9A data scientist wants to compare the performance of two model versions (V1 and V2) in production by splitting traffic between them. They want to gradually increase the percentage of traffic to the new version while monitoring metrics. Which SageMaker feature enables this?
Easy10A team is training a PyTorch model using SageMaker. They have a custom training script that requires specific Python packages not included in the SageMaker default PyTorch container. Which approach should they use?
Medium11A data scientist needs to deploy a single ML model that will serve real-time predictions with low latency (under 10 ms) for a high-traffic web application. The model fits in memory and requires GPU acceleration. Which SageMaker inference option is MOST suitable?
Easy12A machine learning team is preparing numerical features for a linear regression model. Feature 'A' ranges from 0 to 1000, feature 'B' ranges from 0 to 1, and feature 'C' ranges from -10000 to 10000. The team wants to ensure that feature scales do not affect the model's coefficients and that the features are bounded between 0 and 1. Which transformation should they apply?
Hard13A company wants to implement a retraining pipeline that automatically triggers when SageMaker Model Monitor detects data drift. The retraining job should use the latest approved pipeline version in SageMaker Pipelines. Which approach meets these requirements?
Medium14A company uses AWS Glue ETL jobs to transform data for machine learning. They have a dataset with a column 'income' that is heavily right-skewed. Which transformation should be applied to make the distribution more Gaussian-like?
Hard15A data scientist is using SageMaker Autopilot to automatically build a binary classification model on a balanced dataset. They want to understand the relationship between the input features and the model predictions. Which feature in SageMaker Autopilot should they use?
Medium16A company has a model that receives low traffic but needs to handle sudden spikes. Which deployment option is most cost-effective?
Easy17A data engineer needs to ingest streaming clickstream data from a website into an S3 data lake for ML training, with the ability to run real-time aggregations before storage. Which combination of AWS services meets these requirements?
Medium18A machine learning team is training a large natural language processing model on Amazon SageMaker using the SageMaker Hugging Face container. The training job runs on multiple instances and uses Managed Spot Training to reduce costs. However, the job frequently gets interrupted by Spot interruptions, causing long training times. What should the team do to mitigate this issue?
Hard19A machine learning engineer has trained a model using SageMaker and wants to deploy it to a real-time endpoint. The engineer needs to specify the model artifacts, the inference code, and the environment. Which SageMaker resource should the engineer create first?
Easy20A machine learning team is developing a deep learning model for image classification. They observe that the training loss decreases rapidly but the validation loss starts increasing after a few epochs. Which strategy should they implement to address this issue?
Hard21A data scientist notices that a SageMaker endpoint is returning HTTP 5XX errors under high load. The endpoint uses a single ml.m5.large instance. The team wants to reduce these errors without changing the instance type. What is the most cost-effective step?
Easy22A company uses SageMaker Model Monitor for data quality. They notice that monitoring jobs are failing intermittently with constraint violations. Upon review, they see that some features have different data types in production compared to the baseline (e.g., string instead of integer). Which type of drift is this?
Hard23A team is training a PyTorch model using SageMaker with a custom training script. They want to track hyperparameters and metrics across multiple experiments. Which service should they use?
Medium24A machine learning engineer is using SageMaker Automatic Model Tuning (AMT) to optimize a model. They want to ensure the tuning job explores the hyperparameter search space efficiently and stops poorly performing trials early. Which two strategies should they use? (Choose two.)
Medium25A machine learning engineer is using SageMaker Pipelines to automate the training and deployment of a model. The pipeline includes a processing step for feature engineering, a training step, and a model registration step. The engineer wants to ensure that the pipeline is reproducible and that the model artifacts are versioned. Which two actions should be taken? (Choose two.)
Medium26A data scientist wants to track the lineage of models, datasets, and training jobs in SageMaker. Which SageMaker feature should they use to capture these relationships as artifacts and actions?
Easy27A data scientist needs to label a large dataset of product images for a classification model. They want to reduce labeling costs by prioritizing uncertain samples. Which Amazon SageMaker Ground Truth feature should they use?
Easy28A data science team wants to automate the retraining of a model when data drift is detected. Which TWO AWS services should they use in combination to achieve this? (Choose TWO)
Easy29A team wants to deploy a new model using a canary deployment strategy on SageMaker. Which TWO configurations are necessary? (Choose two.)
Medium30A company uses Amazon SageMaker Data Wrangler for data preparation. The data science team wants to automatically detect potential bias in their dataset before training a model. Which feature of Data Wrangler should they use?
Medium31A company has 10 TB of log data in compressed JSON format stored in Amazon S3. The data needs to be processed and transformed into a structured format for machine learning. The processing requires complex transformations, including parsing nested JSON and joining with a reference table. The company wants to minimize infrastructure management. Which approach should the company use?
Easy32A data scientist needs to split a time-series dataset into training and testing sets while avoiding data leakage from future values. Which splitting technique should the data scientist use?
Easy33A company is adopting Amazon SageMaker Pipelines to automate their ML workflow. They want to choose three key benefits that SageMaker Pipelines provides over traditional manual scripts and ad-hoc steps. Which THREE benefits are correct?
Easy34A company is using SageMaker Model Registry to manage model versions. They want to automatically deploy the latest approved model to production after retraining. Which approach is best?
Medium35A data scientist needs to ensure that the same train/test split is used across multiple experiments for reproducibility in SageMaker. Which approach should they take?
Medium36A company stores its model training data in Amazon S3. To meet compliance requirements, all data in transit between the S3 bucket and SageMaker must be encrypted. What should the company enforce?
Easy37A company is using Amazon SageMaker Data Wrangler to prepare a dataset with over 200 features. The dataset includes a categorical feature with more than 10,000 unique values (high cardinality). The ML engineer wants to transform this feature into a numeric representation suitable for a linear model without increasing dimensionality too much. Which built-in transform in Data Wrangler should the engineer use?
Hard38A data engineer is preparing a large training dataset stored in Amazon S3 as many small Parquet files, and a SageMaker training job that reads directly from S3 is spending most of its time on the input channel rather than on model computation. The engineer needs to improve the input throughput without changing the model code or the training algorithm. Which action should the engineer take?
Medium39A company uses SageMaker Autopilot to build a binary classification model. The generated leaderboard shows an ensemble model as the best candidate. The team needs a model that can be deployed for real-time inference with latency < 10ms. What should they do?
Hard40A machine learning engineer is training a tabular regression model using the SageMaker built-in XGBoost algorithm. They want to reduce overfitting and improve generalization without changing the algorithm. Which SageMaker hyperparameter should they tune to control the fraction of features randomly sampled per tree?
Medium41A company deploys a real-time inference endpoint on SageMaker for a customer-facing application. Traffic patterns are unpredictable and sometimes spike. The endpoint must scale automatically to handle load while minimizing cost. Which approach should the company take?
Medium42A machine learning team is building a fraud detection model. They have a dataset with a categorical feature 'merchant_id' that has over 10,000 unique values. Which feature engineering technique should they apply to 'merchant_id' to reduce dimensionality while retaining predictive power?
Medium43A machine learning engineer is deploying a model to a SageMaker endpoint for real-time inference. The model must return predictions within 100 milliseconds for 95% of requests. The engineer wants to monitor the endpoint's latency and automatically roll back if latency exceeds the threshold. Which combination of SageMaker features should be used?
Medium44A company wants to automate remediation when a SageMaker endpoint's latency exceeds a threshold for more than 5 minutes. The team needs to be notified and a Lambda function should be invoked to scale up the endpoint. Which combination of services should be used?
Easy45A company is using SageMaker to train a model with a custom container. The training script requires a specific version of a Python library that is not included in the default SageMaker containers. How should they provide this library?
Hard46A data scientist is preparing a large dataset (50 GB) for training a TensorFlow model on SageMaker. The dataset consists of many small CSV files. Training is slow due to I/O bottlenecks. Which data preparation strategy most effectively accelerates training?
Hard47A data scientist is using SageMaker Autopilot for a regression problem. They want to see which data preprocessing steps Autopilot applied. Which TWO sources can they use to find this information?
Medium48A data scientist wants to version and manage trained models, require approval before deployment, and enable cross-account deployment. Which SageMaker feature provides these capabilities?
Easy49A company uses SageMaker Pipelines to orchestrate their ML workflow. They notice that if a pipeline step fails due to a transient error (e.g., a brief network issue), the entire pipeline fails and they must manually rerun from the beginning. They want to automatically retry failed steps a few times before failing. What should they do?
Hard50A data scientist is using SageMaker to train a deep learning model with the PyTorch estimator. They want to log custom scalar metrics such as validation accuracy and loss during training so they can monitor the job in SageMaker. Which approach should they use to emit these metrics from the training script?
Medium51A machine learning engineer is using a SageMaker training job with a custom training script. They need to save the trained model artifacts to Amazon S3 so that the model can be deployed later. Which parameter in the SageMaker estimator should they configure to specify the S3 location for model artifacts?
Medium52A machine learning engineer is training a model using SageMaker's built-in XGBoost algorithm. The training job fails with an error indicating insufficient memory. Which parameter should be adjusted to reduce memory usage?
Easy53A team deploys a model with SageMaker and notices that the model returns inconsistent results during inference. They suspect a mismatch in feature transformation between the training pipeline and the inference pipeline. Which SageMaker feature can help compare the feature distributions?
Medium54A financial services company is training a fraud detection model using SageMaker. The dataset is highly imbalanced, with only 0.2% fraudulent transactions. The team wants to optimize the model for recall at a fixed precision of 90%. They are using the SageMaker built-in XGBoost algorithm with binary:logistic objective. Which evaluation metric should they monitor during training and hyperparameter tuning?
Hard55A data engineer needs to prepare a large dataset for machine learning. The data is stored in an Amazon RDS MySQL database and needs to be transformed and moved to an S3 bucket in Parquet format for use with SageMaker. Which AWS service is most suitable for this extraction, transformation, and loading (ETL) task?
Easy56A company is using SageMaker to serve a model for real-time predictions. They want to test a new model version by routing a small percentage of live traffic to it while the rest goes to the current model. They also need to compare performance metrics. Which TWO actions should they take? (Select TWO.)
Medium57A machine learning engineer wants to set up a retraining pipeline that triggers when model quality degrades. Which TWO components are essential for this automated retraining pipeline? (Select TWO)
Easy58A company uses Amazon SageMaker Pipelines for automated retraining. The pipeline includes a processing step that runs a Python script. The script uses the boto3 library to call an AWS service, but the calls are being throttled. What is the MOST effective way to address this within the pipeline?
Medium59A machine learning engineer wants to monitor a deployed model for data drift. Which SageMaker feature should they use to automatically detect drift in the input data distribution compared to the training data baseline?
Easy60A machine learning engineer has trained a scikit-learn model and saved it as model.joblib in Amazon S3. The engineer wants SageMaker to host the model for real-time inference without writing a custom container or inference script, because the model uses only standard predict behavior. Which deployment approach should the engineer use?
Medium61A team uses SageMaker Pipelines for CI/CD. The training step fails due to insufficient memory. How to fix without rewriting code?
Hard62A company is deploying a foundation model using SageMaker JumpStart. They want to minimize inference costs while maintaining low latency. Which TWO strategies should they consider? (Select TWO)
Hard63A company operates multiple AWS accounts with SageMaker workloads. They need to implement governance and security controls for model monitoring and maintenance. Which THREE actions should they take to meet compliance requirements?
Hard64A fraud-detection team runs a SageMaker real-time endpoint. Compliance requires that every inference request be logged with its full request and response payloads, and that a security engineer be able to prove later which requests were captured. The team enables SageMaker Model Monitor data capture with a capture percentage of 100. Where are the captured records stored, and what must be configured so the records are encrypted with a customer-managed key rather than an AWS-managed key?
Medium65A company uses an Amazon SageMaker endpoint for real-time inference. The security team requires that all traffic between the endpoint and the client application be encrypted in transit. Which configuration ensures this?
Easy66A company is deploying a machine learning model using SageMaker hosting. They need to support multiple versions of the model for A/B testing. Which TWO actions are required to set up the A/B test? (Choose two.)
Medium67A company is using Amazon SageMaker Ground Truth to create a labeled dataset for object detection in images. The team wants to minimize labeling costs while maintaining high accuracy. Which feature should they use to achieve this?
Easy68A data scientist is building a model to predict customer churn based on historical data. The dataset has 10 features and 100,000 records, and the target is binary. Which algorithm is most appropriate for this binary classification problem?
Easy69A data scientist wants to use SageMaker Clarify to analyze bias during training of a binary classification model. Which TWO types of bias metrics can SageMaker Clarify compute? (Select TWO.)
Medium70A healthcare company is deploying a model for predicting patient outcomes. The model must be deployed across multiple AWS accounts to meet compliance requirements. Each account has its own Amazon SageMaker endpoint. The company wants to centralize monitoring of model performance without exposing data across accounts. Which solution should the company use?
Hard71A machine learning engineer is setting up automated retraining for a model using SageMaker Pipelines. The pipeline should trigger when a data drift alert is received from Model Monitor. Which event source should the engineer use to initiate the pipeline?
Hard72An ML engineer needs to orchestrate a multi-step workflow that includes data preprocessing on Spark, model training on SageMaker, and deployment to a production endpoint. They require tight integration with other AWS services and the ability to add custom logic. Which AWS service should they use alongside SageMaker?
Medium73A company is training a deep learning model for object detection using SageMaker. The training is very slow and the GPU memory is insufficient for the batch size. The team wants to scale across multiple GPUs efficiently. Which THREE actions should they take? (Choose THREE.)
Hard74A healthcare company has deployed a SageMaker model that predicts patient risk scores. The security team requires that all access to the model's endpoint be authenticated and authorized, and that every invocation be traceable to a specific user or application for audit purposes. The team also wants to enforce least privilege so that only specific applications can invoke the endpoint. Which TWO actions should the machine learning engineer take to meet these requirements? (Choose two.)
Hard75A team runs a SageMaker Pipeline that trains a model and registers it in the Model Registry. Compliance requires that the pipeline run automatically every time new labeled data lands in S3, and that each run record the exact S3 data prefix, the training image URI, and the git commit hash as lineage metadata. The engineer wants the least operational overhead. Which approach meets these requirements?
Hard76A data scientist is using SageMaker to train a custom PyTorch model for image classification. They want to use SageMaker Debugger to detect training issues. Which TWO built-in rules are most relevant for detecting common training problems? (Select TWO.)
Medium77A company has a SageMaker endpoint that serves a recommendation model. The security team wants to ensure that the model artifacts stored in Amazon S3 are encrypted at rest and that access to the S3 bucket is limited to the SageMaker execution role only. The team also wants to receive alerts if the bucket policy is changed. Which combination of actions should the machine learning engineer take?
Easy78A team is using SageMaker to run a large-scale distributed training job for a language model. They are using SageMaker's Pipe mode to stream data from S3 to reduce IO. They observe that the training throughput is lower than expected, and the CPU utilization is high while GPU utilization is low. The training script uses PyTorch's DataLoader with num_workers=0. The data preprocessing is minimal. Which change is most likely to improve GPU utilization?
Hard79A machine learning engineer is using Amazon SageMaker Debugger to monitor a training job for a deep neural network. They receive a rule alert indicating 'exploding gradients'. Which action should they take to address this issue?
Hard80A data scientist is training a deep learning model using SageMaker and wants to use distributed training across multiple GPUs to reduce training time. Which TWO actions should the scientist take to configure distributed training? (Select TWO.)
Medium81A machine learning engineer is configuring auto-scaling for a SageMaker real-time endpoint. The endpoint is expected to have steady traffic during business hours and low traffic at night. The engineer wants to minimize costs by scaling in during low traffic, but the model container has a long start-up time (about 5 minutes). Which scaling policy should the engineer use to prevent request drops during sudden traffic spikes?
Medium82A machine learning engineer needs to optimize a trained TensorFlow model for deployment on edge devices with limited compute. Which SageMaker feature should they use to compile the model for target hardware?
Easy83A team is preparing text data for a natural language processing (NLP) model. They have a corpus of customer reviews. Which THREE preprocessing steps are essential to reduce noise and improve model performance?
Hard84An ML engineer wants to use MLflow on SageMaker to track experiments and log metrics. They have set up MLflow on an EC2 instance. How can they best integrate MLflow tracking with SageMaker training jobs?
Medium85A machine learning engineer is using Amazon SageMaker Data Wrangler to create a data preparation pipeline. The pipeline includes multiple transforms such as handling missing values, scaling, and encoding. The engineer wants to export the prepared data directly to a feature group in Amazon SageMaker Feature Store for reuse in training and inference. Which export option should the engineer choose?
Medium86A fraud detection team runs a SageMaker real-time endpoint that logs every request and response to an Amazon S3 bucket. Compliance requires that the model's prediction inputs and outputs be encrypted at rest with a customer-managed AWS KMS key, and that the endpoint be able to read the capture bucket only when necessary. The team enables data capture and specifies a KMS key on the endpoint configuration. Which additional configuration is required for the captured data written to S3 to be encrypted with that customer-managed key?
Medium87A data engineer needs to assess the quality of a dataset containing customer information. The dataset has missing values, outliers, and duplicate records. Which TWO AWS services can be used to perform data quality assessment? (Select TWO.)
Medium88A machine learning team needs to deploy a model that was built using scikit-learn. They want to use SageMaker for hosting. Which approach should they take?
Easy89A startup wants to deploy a model that has variable traffic patterns, with some periods of no traffic and occasional spikes. They want to pay only for what they use and do not want to manage instances. Which SageMaker inference option should they choose?
Medium90A data scientist is preparing a dataset stored in Amazon S3 for a SageMaker training job. The dataset contains missing values in several columns. The scientist wants to impute missing values with the mean of each column. Which SageMaker built-in algorithm or processing method should be used to perform this imputation efficiently?
Easy91A data scientist is using SageMaker Experiments to track multiple training runs. They want to compare different hyperparameter configurations and visualize the impact on model accuracy. What should they use to track hyperparameters?
Medium92A machine learning engineer at a retail company is monitoring a production model that predicts inventory demand. The model's prediction accuracy has dropped significantly over the past week. The engineer checks the model's input data and notices a new product category was introduced with a different distribution. Which concept is most likely causing the performance degradation?
Easy93A data engineer needs to prepare a dataset for a fraud detection model. The dataset contains a highly skewed numerical feature with extreme outliers. The engineer decides to apply a logarithmic transformation to this feature before training. Which SageMaker Data Wrangler transform should be used to apply the logarithmic transformation?
Medium94A team is using Amazon SageMaker for feature engineering. They have a dataset with a column 'TransactionDate' in string format (e.g., '2023-01-15 10:30:00'). They need to create features: year, month, day, hour, and day_of_week. What is the most efficient way to do this in a SageMaker processing job?
Medium95A team is preparing text data for sentiment analysis. They have a large corpus of customer reviews. They want to convert the text into numerical features using a technique that captures word importance relative to the whole corpus. Which feature extraction method should they use?
Medium96A company is building a real-time fraud detection system. They need to store features with historical context for model training and also support low-latency lookups for inference. Which THREE configurations should they set up in Amazon SageMaker Feature Store? (Select THREE.)
Hard97A machine learning team needs to ensure that all model training and inference jobs within SageMaker Studio run in a private network without internet access. The team also requires that inter-container traffic within the same training job be encrypted. Which configurations should they combine?
Hard98A company uses Amazon SageMaker Ground Truth to label a dataset for object detection. To reduce labeling costs, they want to use active learning. Which configuration should they set up in Ground Truth?
Hard99A company is using SageMaker Autopilot to automatically build a regression model on a dataset. They want to understand which features are most important for the model's predictions. Which feature of Autopilot can provide this insight?
Medium100A company uses Amazon SageMaker Ground Truth to create a labeled dataset. They want to monitor the accuracy of human labelers during the labeling process. Which metric should they track?
Hard101A data science team deploys a PyTorch model on Amazon SageMaker for real-time inference. The model requires GPU for low latency. Which instance type is MOST cost-effective while meeting the GPU requirement?
Easy102A data engineer needs to perform feature selection on a dataset with 500 numeric features to train a regression model. The engineer wants to remove features that are redundant or have low predictive power. Which TWO techniques should the engineer consider? (Select TWO.)
Medium103A data scientist is working on a binary classification problem and wants to use AWS Glue for data preparation. The dataset has missing values in several numeric columns. Which imputation strategy is MOST appropriate for the scientist to apply in AWS Glue ETL?
Easy104A data science team uses SageMaker to train and deploy models. They need to track model lineage, including datasets, training jobs, and model versions, to ensure reproducibility. Which THREE actions should they take? (Select THREE)
Medium105A data engineer is using Amazon SageMaker Processing to run a data preprocessing script on a dataset with 500 million rows. The script runs out of memory on a single ml.r5.24xlarge instance. The engineer needs to modify the processing job to handle the dataset size. Which approach is most cost-effective and scalable?
Hard106A machine learning team needs to deploy a new model version for A/B testing, gradually shifting traffic from the old version to the new version over 24 hours. Which deployment strategy should they use?
Medium107A data science team wants to track the lineage of models, including datasets, training jobs, and endpoints, for reproducibility and audit. They need a solution that captures relationships between artifacts automatically during training and deployment. Which service should they use?
Medium108A data scientist wants to track hyperparameters, metrics, and artifacts for multiple training runs in SageMaker. They need to compare runs and identify the best performing model. Which SageMaker feature should they use?
Medium109A machine learning engineer is using Amazon SageMaker Processing to preprocess a large dataset. The processing job runs a custom Python script that uses the pandas library to read multiple CSV files from an S3 input prefix. The script must write the processed output to a different S3 prefix. Which configuration of the ProcessingInput and ProcessingOutput parameters is correct for this scenario?
Hard110During data preparation for a regression model, a data scientist notices that two features have a Pearson correlation coefficient of 0.95. The scientist is concerned about multicollinearity. Which action should be taken to address this issue?
Medium111A company needs to perform time-series forecasting on historical sales data. Which SageMaker built-in algorithm is BEST suited for this task?
Easy112A team uses SageMaker Pipelines to automate model retraining. After a successful pipeline run, they want to register the new model version in the SageMaker Model Registry so that it can be reviewed for approval. Which step type should they add to the pipeline?
Easy113A gaming company uses a SageMaker endpoint for real-time player churn prediction. The model is updated weekly. After a recent retraining, the team notices that the endpoint's predicted probabilities for churn have shifted dramatically: the average predicted probability dropped from 0.3 to 0.05. The team suspects concept drift (the relationship between features and target changed) rather than data drift. They have SageMaker Model Monitor set up for data drift and quality metrics, but not for bias or explainability. The team needs to confirm concept drift and take corrective action. Which approach should the team take FIRST?
Medium114An ML team has deployed a model to a SageMaker real-time endpoint and wants to set up automated monitoring for model quality. Which TWO elements are required to configure SageMaker Model Monitor for model quality? (Select TWO.)
Medium115A data engineer must prepare a 4 TB Parquet dataset stored in Amazon S3 for a SageMaker training job that runs on 8 ml.p4d.24xlarge instances. The engineer wants the fastest possible data throughput during training while minimizing per-epoch I/O overhead. The dataset is immutable for the duration of the training run. Which approach BEST meets these requirements?
Medium116A team uses SageMaker Model Monitor to track data quality. They notice that the monitor's constraint violations are increasing but the model performance remains good. What should they do?
Medium117An ML team uses AWS Step Functions to orchestrate a retraining pipeline triggered by EventBridge when new training data arrives. The pipeline includes a SageMaker training job and a model evaluation. If evaluation fails, the team wants to send an alert. How should they implement this?
Hard118An ML engineer monitors a SageMaker endpoint for data drift. They set up SageMaker Model Monitor to compare inference data against a baseline created from the training dataset. The monitoring schedule runs daily and reports violations. Which monitoring type should be configured to detect if the distribution of a numerical feature in real-time inference data differs significantly from the training distribution?
Medium119A company needs to deploy a large language model (LLM) on SageMaker with the Triton Inference Server to maximize GPU utilization and reduce latency. They have an NVIDIA A100 GPU. Which SageMaker inference option supports Triton?
Hard120A machine learning team at a retail company has deployed a product recommendation model using Amazon SageMaker. The model is updated weekly with new data. Recently, the team noticed that the model's accuracy on a holdout evaluation set has been declining over the past month. The data pipeline that feeds the training job has not changed. The team suspects data drift. They have SageMaker Model Monitor enabled on the inference endpoint and have set up Amazon CloudWatch metrics for feature distribution distances. Upon reviewing the CloudWatch dashboards, they see that the feature distribution distance metric for the most important feature 'product_category' has increased significantly. However, the team is unsure if this is the root cause. Which remediation step should the team take FIRST?
Easy121A data science team uses SageMaker Pipelines for automated training. They need to conditionally register a model only if evaluation metrics exceed a threshold. Which pipeline step type should they use after the evaluation step?
Medium122A machine learning engineer is preparing a dataset for a binary classification model. The dataset has 10,000 rows and 200 features, with 5% positive class. The engineer suspects class imbalance may affect model performance. Which TWO actions should the engineer take to mitigate imbalance? (Choose 2.)
Medium123An ML engineer is using Amazon SageMaker Automatic Model Tuning (AMT) to optimize hyperparameters for a gradient boosting model. The tuning job is taking a long time and has completed many training jobs. The engineer wants to stop training jobs that are unlikely to improve the objective metric. What should they configure?
Medium124A company trains a model daily using Amazon SageMaker and uses the model for real-time inference. They want to detect data drift between the training data and the inference data to decide when to retrain. Which AWS service should they use for this purpose?
Medium125A machine learning engineer wants to store, share, and manage features for multiple ML models across an organization. The features need to be accessible for both real-time inference (low-latency) and batch training. Which AWS service should the engineer use?
Easy126A data engineer needs to split a time-series dataset into training and validation sets for a forecasting model. Which split method should be used to avoid data leakage?
Easy127A data scientist has a 40 GB CSV dataset in Amazon S3 that will be used to train a SageMaker model. The training script reads the data with pandas, and the scientist wants to reduce both storage cost and training-time I/O without changing the logical schema. Which data preparation action should be taken?
Easy128A team is deploying a model that requires GPU acceleration for inference. They are using an Amazon SageMaker real-time endpoint. The model is a large language model (LLM) that does not fit on a single GPU. Which configuration should they use to minimize latency while fitting the model?
Hard129Which SageMaker feature allows you to automatically tune hyperparameters using Bayesian optimization?
Easy130A data engineer wants to transform a categorical feature with 1,000 possible values into numerical features for a linear model. Which feature engineering technique is most appropriate for this high-cardinality feature?
Easy131A company wants to ensure that only authorized users and services can invoke a SageMaker real-time endpoint. Which AWS service can be used to manage access control?
Easy132A data team is using Amazon SageMaker Data Wrangler to prepare a dataset. They need to detect potential bias in the data before training a model. Which feature of Data Wrangler should they use?
Medium133A team deploys a PyTorch model on Amazon SageMaker for real-time inference. They notice that inference latency is higher than expected. They suspect the serialization format used for input data is inefficient. Which approach would MOST likely reduce latency?
Medium134A machine learning engineer is using SageMaker Automatic Model Tuning to optimize hyperparameters for a regression model. The objective metric is RMSE. The training job is costly, and the engineer wants to find a good configuration quickly. Which tuning strategy should they use?
Hard135A machine learning team is using Amazon SageMaker to train a model. They notice that the training job is taking longer than expected and the logs show repeated warnings about 'loss not decreasing'. Which SageMaker feature should they use to diagnose and visualize the training process?
Medium136Refer to the exhibit. A team has configured data capture for a SageMaker endpoint. The endpoint is returning predictions but no captured data appears in the S3 bucket. What is the most likely cause?
Easy137An ML engineer is preparing a time-series dataset for a forecasting model that predicts daily sales for the next 30 days. The dataset contains 3 years of daily sales data. Which data splitting strategy should the engineer use to evaluate the model's performance on future data?
Hard138A company wants to secure access to a SageMaker real-time endpoint. Which TWO actions should be taken? (Select two.)
Medium139A company runs a batch inference job on 10 TB of image data stored in S3. Each image needs to be processed by a GPU-accelerated model. The job is not time-sensitive and cost is the primary concern. Which SageMaker option is MOST appropriate?
Medium140A company wants to use SageMaker to serve real-time predictions with a model that has a large memory footprint. They need to ensure the endpoint can handle traffic spikes. Which scaling policy should they use?
Easy141A team is training a PyTorch model using SageMaker and wants to use their own custom training container with a specific PyTorch version. Which approach should they use?
Medium142A data engineer is preparing a dataset for a k-means clustering algorithm. The features have different scales: age (18-100), income ($20k-$200k), and number of purchases (0-50). Without scaling, which feature will dominate the distance calculations?
Easy143A data scientist is using SageMaker to train a linear regression model on a dataset with a large number of features. They notice that the model's training time is long and want to speed it up by using a more efficient algorithm. They decide to use the SageMaker built-in Linear Learner algorithm. Which of the following is a key advantage of using the Linear Learner algorithm in SageMaker for this scenario?
Easy144Refer to the exhibit. A data scientist configured an automatic model tuning job for a classification model. The tuning job completed after 20 training jobs, but the best validation accuracy was only 0.65. What is the most effective way to potentially improve the result?
Medium145A social media company is processing a real-time stream of user activity data from Amazon Kinesis Data Streams to train a machine learning model for content recommendation. The raw data includes user ID, timestamp, content ID, interaction type (like, share, comment), and device type. The data scientists need to aggregate features per user over a sliding window of 7 days, including counts of interaction types, unique content IDs engaged, and a moving average of interaction timestamps. The aggregated data will be used to update a user embedding model. The streaming data volume is approximately 500 records per second, and the company uses an AWS Glue streaming ETL job for transformation. However, the Glue job is failing frequently with high latency and checkpoint errors. The team needs a more robust solution to prepare the streaming data features. Which approach should the team take?
Hard146A machine learning engineer is deploying a model using SageMaker and needs to ensure that the endpoint can automatically scale based on traffic patterns. Which TWO actions should the engineer take? (Choose two.)
Medium147An ML engineer needs to create a feature store that supports both low-latency online inference and large-scale offline training. The features are updated hourly from a streaming source. Which Amazon SageMaker Feature Store configuration should the engineer use?
Easy148A company wants to forecast monthly sales that show clear seasonality. Which algorithm is most suitable?
Hard149A company uses SageMaker Model Monitor's feature attribution drift monitoring with SHAP. They receive an alert that the average SHAP value for a particular feature has increased significantly compared to the baseline. The feature's input distribution has not changed. What does this likely indicate?
Hard150A machine learning team is building a feature store using Amazon SageMaker Feature Store. They need to store features that support both real-time inference (low latency) and historical training. Which configuration should they choose?
Hard151A data engineer is preparing a dataset in Amazon SageMaker Data Wrangler for a binary classification model. The dataset contains missing values in several numeric columns, and the engineer wants a reusable, reproducible transformation that can be applied identically to the training data and to future inference data. The engineer plans to export the transformation and integrate it into a SageMaker Pipeline. Which TWO actions should the engineer take to ensure the imputation is consistent between training and inference? (Choose two.)
Medium152A company has a trained machine learning model that needs to be deployed as a real-time inference endpoint on Amazon SageMaker. The endpoint must automatically scale based on incoming traffic. Which SageMaker feature should be used?
Easy153A company wants to detect anomalies in login events from a large user base, focusing on unusual patterns that may indicate compromised accounts. Which SageMaker built-in algorithm is most suitable for this task?
Easy154A data scientist has a 200 GB Parquet dataset in Amazon S3 that will be used to train a SageMaker model. The training script reads the data with the SageMaker training toolkit's File mode, and the job currently spends a long time downloading before training begins. The team wants to reduce startup time without changing the training algorithm. Which change should the data scientist make?
Easy155A company wants to use a pre-trained NLP model from SageMaker JumpStart for sentiment analysis. Which step is required to make predictions?
Hard156A company wants to version and track ML models, with an approval workflow for promoting models from staging to production. Which SageMaker feature should they use?
Easy157A financial services company is deploying a model for loan approval. They must ensure that the model's predictions do not show bias against protected groups. They plan to monitor for bias drift after deployment. Which SageMaker feature should they use?
Medium158A machine learning engineer runs a training job and notices the loss is NaN after a few steps. Which SageMaker Debugger rule can help identify this issue?
Medium159A machine learning engineer is training a deep learning model on SageMaker using the PyTorch estimator. The training job fails with an error indicating that the GPU memory is exhausted. The engineer wants to reduce memory usage without changing the model architecture. Which SageMaker feature should the engineer use?
Hard160A financial institution uses SageMaker to train and deploy models. They need to track every experiment, model version, and deployment step for audit purposes. Which SageMaker feature should they use to capture the full lineage of artifacts, actions, and contexts?
Medium161A team is deploying a real-time inference endpoint in SageMaker. The model requires access to an S3 bucket containing customer data, which is encrypted with SSE-KMS. The team needs to ensure that the endpoint can decrypt the data. Which IAM role configuration is necessary?
Hard162A data science team is using SageMaker Experiments to track hyperparameters and metrics for a model training project. They need to compare multiple trials and identify the best model. Which THREE actions are part of a typical workflow? (Select THREE.)
Hard163A data scientist wants to train a binary classification model using Amazon SageMaker with a built-in algorithm that performs well on tabular data. Which algorithm should they choose?
Easy164A company is deploying a large NLP model on SageMaker for real-time inference. They want to reduce inference latency and cost by optimizing the model for the target hardware. The model is trained in PyTorch. Which SageMaker feature should they use to compile the model for best performance on the chosen instance?
Medium165A machine learning engineer must grant a data scientist the least-privilege permissions needed to invoke one specific SageMaker real-time endpoint from their own AWS account, and to view that endpoint's CloudWatch metrics without being able to modify the endpoint. The endpoint ARN is known. Which TWO IAM policy statements should the engineer include? (Choose two.)
Medium166A healthcare analytics team stores model artifacts and training datasets in Amazon S3 and uses SageMaker. An internal audit finds that some S3 buckets containing protected health information are missing encryption and that access is granted broadly. The team must remediate quickly and prevent future misconfiguration. Which combination of actions should the ML engineer take FIRST?
Medium167Refer to the exhibit. A data scientist reviews the CloudWatch Logs from an Amazon SageMaker real-time endpoint. What is the MOST likely root cause of the NaN output?
Easy168A machine learning engineer runs a SageMaker Processing job that must load a 200 GB dataset from S3, compute statistics, and write a small summary to S3. The job repeatedly fails with an out-of-disk-space error on the processing instance. Which change is MOST likely to resolve the failure?
Hard169A machine learning engineer uses Amazon SageMaker Data Wrangler to preprocess a dataset. After applying a transform, the engineer wants to export the data to a feature group in Amazon SageMaker Feature Store for reuse in training and inference. Which export option should they choose?
Medium170A data scientist is preparing text data for a sentiment analysis model using Amazon SageMaker. Which two data preprocessing techniques are commonly used when working with text data for natural language processing? (Choose two.)
Medium171A company is training a deep learning model using SageMaker and wants to reduce the time spent on data loading from Amazon S3 during training. The training dataset consists of many small files. Which approach is MOST effective to accelerate data loading?
Hard172A data scientist wants to track feature definitions, share them across teams, and serve features for both training and real-time inference. Which AWS service provides these capabilities?
Easy173Which TWO of the following are best practices for deploying machine learning models on SageMaker? (Select TWO.)
Medium174A machine learning engineer is performing feature selection for a regression model with 200 features. The dataset has 10,000 samples. The engineer wants to remove irrelevant features while keeping those that have a strong non-linear relationship with the target. Which feature selection method is best suited for this requirement?
Hard175Refer to the exhibit. A data scientist ran a training job using a custom algorithm container. The job failed with the error shown. What is the most likely cause?
Easy176A team is using AWS Step Functions to orchestrate a machine learning workflow that includes data preprocessing, training, and model evaluation. The team wants to run the workflow whenever new data arrives in an S3 bucket. Which approach should they use to trigger the Step Functions workflow?
Medium177A data engineer is building a data pipeline for a machine learning model that requires both structured and unstructured data. The structured data (customer demographics) is in Amazon RDS, and the unstructured data (customer support chat logs) is in Amazon S3 as JSON files. The engineer needs to combine these datasets into a single training dataset stored in S3 in Parquet format. They must also perform feature engineering such as text vectorization on the chat logs. The pipeline should be serverless and cost-effective. Which approach should they use?
Medium178A machine learning engineer is preparing a training job on SageMaker with a custom Docker container. Which TWO actions are required to use the container with SageMaker? (Choose TWO.)
Medium179A data scientist is building a binary classification model on a highly imbalanced dataset where the positive class represents only 1% of the data. The scientist needs to train the model using Amazon SageMaker's built-in XGBoost algorithm. Which strategy should be used to address the class imbalance?
Medium180An ML team uses AWS Step Functions to orchestrate a multi-step inference pipeline: data preprocessing, model inference, and postprocessing. The pipeline runs on demand for single records. The team notices that the pipeline occasionally fails due to timeouts in the preprocessing step. They want to implement retries with exponential backoff and a maximum retry count of 3 for that step. How should they configure this?
Medium181A data scientist is preparing a dataset for a linear regression model. The features have different scales: one feature ranges from 0 to 1000, another from 0 to 1, and a third from -5 to 5. The scientist wants to ensure that all features contribute equally to the model. Which TWO scaling techniques should the scientist consider? (Select TWO.)
Medium182An engineer runs: aws sagemaker describe-endpoint --endpoint-name my-endpoint and receives the exhibit output. The engineer wants to update the endpoint to use a new model version stored in ECR with tag ':2'. Which step is necessary to perform the update?
Medium183A team wants to evaluate a binary classification model for credit risk. They need to understand the trade-off between false positives and false negatives. Which TWO metrics should they use? (Select TWO.)
Medium184A data engineer needs to catalog metadata from multiple data sources across the organization for use in ML workflows. Which AWS Glue component should be used to store and manage this metadata?
Easy185A data scientist needs to split a dataset into training, validation, and test sets. The dataset has a categorical target variable with imbalanced class distribution. Which splitting technique ensures that each subset has a similar proportion of each class?
Medium186A data scientist is preparing a dataset for a binary classification model. The dataset has 10,000 samples, but the positive class represents only 2% of the data. The data scientist needs to train a model that will be evaluated on a hold-out test set that preserves the original class distribution. Which data preparation strategy is MOST appropriate?
Medium187A company uses SageMaker Pipelines to automate model retraining. The pipeline runs daily but sometimes fails due to data quality issues. What is the best design to handle this?
Medium188A data engineer is preparing a dataset for a SageMaker training job. The dataset contains a timestamp column and is stored in Amazon S3 as CSV files. The engineer needs to ensure that the training job reads the data efficiently and that the data is partitioned by date to improve query performance in Amazon Athena. Which action should the engineer take?
Medium189A company wants to use SageMaker Autopilot to automatically build a binary classification model. Which output does Autopilot provide to help understand model decisions?
Medium190A machine learning engineer is using AWS Glue ETL to transform a large dataset stored in Amazon S3. The transformation involves joining two tables on a high-cardinality column and aggregating results. The job is running slowly and the engineer needs to improve performance. Which optimization technique should the engineer apply?
Medium191A data scientist runs this pipeline but the Train step fails with "ResourceLimitExceeded". What is the most likely cause?
Medium192A data scientist has trained a model that achieves 95% accuracy on the training set but only 70% on the test set. Which of the following is the most likely cause?
Medium193A company deploys a model using SageMaker and enables data capture for monitoring. After a week, they notice that the captured data is not being written to the specified S3 bucket. The endpoint is running and invocations are successful. What is the most likely cause?
Hard194Refer to the exhibit. An IAM policy is attached to a user to allow invoking a SageMaker endpoint. A developer tries to call the endpoint from a laptop with IP 203.0.113.5 and receives an access denied error. What is the most likely reason?
Hard195A data scientist is working with a dataset containing a categorical feature 'country' with 200 unique values. They plan to use a linear regression model. Which encoding method is most suitable to avoid the dummy variable trap while maintaining interpretability?
Medium196An ML engineer is setting up monitoring for a SageMaker endpoint. Which THREE metrics should be monitored to detect performance issues? (Select THREE.)
Medium197A SageMaker endpoint is logging an error when processing inference requests that require database access. What is the most likely cause?
Medium198A data engineer is using AWS Glue to prepare a dataset for ML. The engineer wants to split the dataset into training and testing sets while preserving the distribution of the target variable. Which TWO methods achieve this goal? (Select TWO)
Easy199A data scientist is using SageMaker Data Wrangler to prepare a dataset for a binary classification model. The dataset contains a mix of numerical and categorical features. The scientist wants to perform feature engineering to improve model performance. Which TWO actions are appropriate for handling categorical features in Data Wrangler? (Choose two.)
Medium200Which SageMaker built-in algorithm should be used for forecasting time series data with seasonal patterns?
Easy201A data scientist is using Amazon SageMaker Data Wrangler to prepare a dataset for classification. They want to detect potential bias in the data before training. Which SageMaker service should they use in conjunction with Data Wrangler to detect bias?
Medium202A retail company uses a SageMaker Model Monitor data quality monitor on a real-time endpoint. The monitor's baseline was generated from a training dataset in which the "promo_code" feature was often null. In production the feature is now populated for nearly every record, and the monitor reports violations even though model accuracy has not degraded. The team wants the monitor to stop flagging this expected change without disabling monitoring entirely. What should they do?
Hard203A machine learning engineer trains a binary classifier and obtains an accuracy of 95% on the test set. The dataset is imbalanced with 95% positive class. What is the most important metric to evaluate the model's performance?
Easy204A financial services company must ensure that a SageMaker model deployed to a real-time endpoint only produces predictions consistent with a fairness constraint on a protected attribute, and that any violation is detected within minutes and triggers an alert to the compliance team. The model is already deployed and monitored for data quality. Which approach should the machine learning engineer implement?
Hard205A data science team detects that a deployed model's prediction accuracy is degrading over time due to concept drift. They need to implement a retraining strategy. Which THREE actions are recommended best practices for handling concept drift?
Medium206A company's ML pipeline runs in multiple AWS accounts (dev, test, prod). They want to enforce that only approved models from a central Model Registry can be deployed to the production account. Which combination of services is MOST appropriate to implement this governance?
Hard207A data scientist needs to ingest streaming customer clickstream data from a website into an S3 data lake for ML training. The data must be delivered within 1 minute of ingestion, and JSON records must be converted to Parquet. Which AWS service combination should be used?
Medium208A machine learning engineer wants to reduce training costs by using excess EC2 capacity. Which instance purchasing option should they choose for SageMaker training jobs?
Easy209A data scientist is training a large transformer model using SageMaker's model parallelism library. The training job is failing with an out-of-memory (OOM) error. Which two actions can help resolve the OOM error? (Choose two.)
Hard210A data engineer is designing a feature engineering pipeline using Amazon SageMaker Feature Store. The team needs to support both real-time inference (millisecond latency) and batch training jobs that require access to historical feature values at specific points in time. Which configuration should the engineer choose?
Medium211A data scientist wants to train a model on SageMaker using a custom PyTorch script, then register the best model in the SageMaker Model Registry. The training job is part of a SageMaker Pipeline. Which pipeline step should be used to register the model?
Medium212A company wants to enable cross-account access to a SageMaker model endpoint. The model is in Account A, and Account B needs to invoke it. Which TWO steps are required? (Select TWO)
Hard213A team is building a recommendation system and wants to store and serve features for online and offline models. The features include user statistics (updated daily) and movie metadata (static). The team needs low-latency inference for real-time recommendations and wants to reuse features across multiple models. Which AWS service should the team use to store, manage, and serve these features?
Medium214A machine learning engineer needs to split a time-series dataset for a forecasting model. The data spans 3 years of daily sales. Which splitting strategy should they use to avoid look-ahead bias?
Medium215A team is building a time-series forecasting model for daily sales data. They want to evaluate model performance using cross-validation while respecting the temporal order of the data. Which data splitting strategy should they use?
Hard216A team is using Amazon SageMaker Data Wrangler to prepare a large dataset. They need to detect potential bias in the data before training. Which capability of Data Wrangler should they use?
Hard217Refer to the exhibit. A data engineer investigates why a SageMaker endpoint is returning errors. The endpoint configuration has been updated to point to a new model version. What is the MOST likely cause of the error?
Medium218A company is using SageMaker Pipelines to orchestrate their ML workflow. They have a Condition step that checks if a model's accuracy exceeds 0.9. If true, they want to register the model in the model registry; otherwise, they want to run a retraining step. Which step type should they use for the decision?
Medium219A company wants to deploy 50 small models (each ~100 MB) for real-time inference. They need to minimize hosting costs while maintaining low latency. Which SageMaker hosting option is most cost-effective?
Medium220An ML team uses Amazon SageMaker Data Wrangler to prepare a dataset for a binary classification model. They suspect the dataset might contain bias against a certain demographic group. They want to detect and visualize potential bias before training the model. Which feature of SageMaker should they use?
Medium221A retail company stores training datasets, model artifacts, and feature data in Amazon S3. An auditor requires that all objects be encrypted at rest with keys the company controls and that key usage be independently auditable. The team wants minimal operational overhead. Which approach should the ML engineer recommend?
Easy222A team deploys a machine learning model using an Amazon SageMaker endpoint. They need to monitor for data drift and model quality issues. Which AWS services or features should they use? (Choose THREE.)
Medium223A company uses Amazon SageMaker Data Wrangler to prepare data for ML. The dataset contains a timestamp column and sensor readings from IoT devices. The data scientist needs to create features such as moving averages and rolling statistics over time windows. Which Data Wrangler transformation type should be selected?
Hard224A machine learning engineer is preparing a training dataset in Amazon SageMaker for a binary classification model. The dataset is stored as a single CSV file in Amazon S3 and contains 12 categorical features with high cardinality (thousands of unique values each). The engineer wants to avoid the curse of dimensionality and reduce training time while preserving predictive power. Which preprocessing approach should be used with the SageMaker built-in XGBoost algorithm?
Medium225A team is building a SageMaker Pipeline that trains a model and then registers it in the SageMaker Model Registry. They want the pipeline to automatically deploy the model to a real-time endpoint only after a human approves the model package. Which two actions should the team take to implement this approval-gated deployment? (Choose two.)
Hard226A data scientist is building a text classification model using a pre-trained BERT model from the Hugging Face library on SageMaker. The scientist wants to fine-tune the model on a custom dataset. Which TWO steps are necessary to set up the fine-tuning job? (Select TWO.)
Medium227A data science team has trained a PyTorch model for real-time inference and needs to deploy it on AWS with GPU acceleration while minimizing cold-start latency. Which SageMaker inference option should they choose?
Easy228A team wants to use a custom PyTorch training script in SageMaker. They need to install additional Python packages not included in the base PyTorch container. Which approach should they take?
Medium229Which TWO SageMaker Pipelines steps are essential for automating a complete ML workflow from data processing to model deployment? (Choose 2.)
Medium230A data scientist is working with a dataset that contains missing values in several numeric features. The data scientist wants to impute the missing values with the median of each feature. Which Amazon SageMaker Data Wrangler transformation should be used?
Easy231A team is deploying a TensorFlow model on a SageMaker real-time endpoint with automatic scaling. They set the scaling policy to target an average CPU utilization of 50%. However, during traffic spikes, the endpoint experiences high latency and 503 errors. The instance type is ml.c5.large. What should the team do to resolve this while minimizing cost?
Hard232A company deploys a machine learning model as a SageMaker real-time endpoint. They need to implement a mechanism to automatically roll back to the previous model version if performance degrades after a deployment. Which approach should they use?
Hard233After deploying a model to a SageMaker endpoint, the operations team notices high inference latency. They suspect it is due to insufficient instance capacity. Which first step should they take to diagnose the issue?
Medium234A machine learning engineer is using SageMaker to train a model with the built-in LightGBM algorithm. The engineer wants to use early stopping to prevent overfitting. The training job is configured with a validation dataset. Which hyperparameter should be set to enable early stopping?
Hard235A company wants to automate its machine learning pipeline using AWS CodePipeline and Amazon SageMaker. The pipeline should train a model, evaluate it, and if the evaluation passes, register the model in the SageMaker Model Registry. Which service should the company use to orchestrate the training and evaluation steps?
Easy236A team receives alerts that their SageMaker endpoint latency has increased significantly. They check CloudWatch metrics and see Invocations rising, but ModelLatency remains stable. Which metric should they investigate to find the source of the increased latency?
Medium237A data science team wants to host 50 different models for a recommendation engine. Each model is small (under 100 MB) and traffic patterns are unpredictable. They need to minimize cost and operational overhead. Which approach should they take?
Hard238A company uses SageMaker Clarify to detect bias in their training data. They find that the model has a high disparate impact for a protected attribute. What should they do to mitigate this bias during training?
Medium239A machine learning engineer is deploying a model to a SageMaker endpoint and wants to ensure that the model's predictions can be explained. The engineer needs to understand which features contributed most to each prediction. Which SageMaker feature should be used?
Hard240A media company runs a real-time recommendation model on a SageMaker endpoint. Traffic triples every evening between 18:00 and 22:00 and drops to near zero overnight, and the team wants to cut costs without underserving evening users. They want the endpoint to scale out automatically as invocations rise and scale back in when traffic falls. Which solution should they implement?
Medium241A team uses SageMaker Clarify to monitor bias drift on a deployed model. They have defined a baseline with training data and set up a monitoring schedule. After one month, they receive a violation report indicating that the post-training metrics have deviated from the baseline. What does this violation indicate?
Medium242A team is using Amazon SageMaker Processing for data preprocessing. They have a Parquet dataset in Amazon S3. Which configuration will provide the most efficient reading of the dataset during processing?
Medium243A machine learning engineer wants to automatically track hyperparameters, metrics, and artifacts for multiple training runs. Which SageMaker feature should they use?
Easy244A financial services company deploys a fraud detection model on a SageMaker real-time endpoint. The inference logic includes a pre-processing step that requires access to a DynamoDB table for user metadata. The model container is a custom Docker image. How should the team grant the endpoint access to DynamoDB?
Hard245A data scientist is performing text preprocessing for a sentiment analysis model. The dataset contains many stop words and rare words. Which combination of preprocessing steps will reduce dimensionality and improve model performance?
Medium246An organization uses SageMaker Studio and needs to restrict Studio's internet access while allowing users to install custom packages from a private PyPI mirror hosted in a VPC. Which networking configuration should they use?
Hard247An organization needs to ensure that all data used for inference on a SageMaker endpoint is encrypted at rest. The endpoint uses a SageMaker-provided container. Which configuration should be applied?
Easy248A data engineer is setting up a Glue ETL job to process a large dataset stored in Amazon S3. The job needs to read data in Parquet format, apply a filter, and write the results back to S3 in Parquet. The engineer wants to minimize the cost and runtime. Which optimization technique is MOST effective?
Easy249A company is running a SageMaker endpoint serving multiple models. They need to monitor for data drift and model quality. Which THREE actions are necessary? (Choose three.)
Hard250A financial services company trains multiple models on SageMaker and needs to track hyperparameters, metrics, and artifacts for each experiment. Which SageMaker feature should they use to organize and compare experiments?
Medium251A company operates an e-commerce platform that uses a machine learning model to recommend products to users. The model is deployed on an Amazon SageMaker endpoint with automatic scaling enabled based on average CPU utilization. The model was trained on historical data and is updated weekly. Recently, the platform experienced a flash sale event that caused a sudden spike in traffic. During the event, the endpoint's latency increased dramatically, and many requests timed out. After the event, the team reviews the CloudWatch metrics and notices that the CPU utilization never exceeded 70%, and the scaling policy was triggered but instances took several minutes to become available. The team wants to prevent similar issues in future flash sales. Which course of action would be MOST effective?
Hard252An ML team wants to deploy a model that was trained using XGBoost in SageMaker. They want to use the built-in XGBoost algorithm container for inference. Which inference option requires the least custom code?
Easy253A company wants to serve a large ensemble of models using NVIDIA Triton Inference Server on SageMaker for high throughput GPU inference. Which SageMaker inference option supports this?
Hard254A data engineer is using Amazon SageMaker Data Wrangler to create a data preparation flow for a dataset with 500 columns, many of which are highly correlated. The goal is to reduce dimensionality while preserving interpretability. Which built-in transform in Data Wrangler should be applied?
Hard255A retailer runs a nightly batch scoring job that processes 40 GB of transaction data and writes predictions to S3. Occasionally a single partition is corrupt, causing the entire job to fail after several hours. The team wants the job to skip the corrupt partition, log which partition failed, and still complete processing of the remaining data with minimal changes to their existing SageMaker Processing job. Which change should they make?
Medium256A machine learning engineer is using SageMaker Automatic Model Tuning (AMT) to optimize hyperparameters for a random forest model. The engineer notices that the tuning job is taking too long and many hyperparameter combinations are being evaluated but not improving the objective metric. Which action should the engineer take to make the tuning more efficient?
Hard257A team uses SageMaker Pipelines to train and register a model. They want to conditionally run a hyperparameter tuning step only if the data quality check passes. Which pipeline step type should they use to branch the execution?
Hard258Which SageMaker built-in algorithm is specifically designed for time series forecasting?
Easy259A company has a SageMaker real-time endpoint that serves predictions. They want to set up automated monitoring and remediation for when the number of 5XX errors exceeds a threshold. Which TWO steps should they take? (Choose TWO.)
Medium260A machine learning team is deploying a model to a SageMaker endpoint and needs to implement A/B testing between two model versions. They want to split traffic 80/20 and monitor performance metrics for each variant. Which two actions should they take? (Choose two.)
Hard261A company uses AWS Glue ETL jobs to clean and transform data from S3 before training. The data contains a column with 40% missing values. The column is normally distributed. Which imputation strategy should the data engineer use?
Medium262A data scientist is preparing a CSV dataset in Amazon S3 for a SageMaker training job. Several rows contain missing values in numeric feature columns, and the chosen algorithm cannot handle NaNs. The scientist wants a repeatable, code-based transformation that runs inside a SageMaker Processing job before training. Which step is the MOST appropriate?
Easy263A company needs to deploy a new model version to a SageMaker real-time endpoint. They want to route 5% of traffic to the new version initially to monitor for errors before full rollout. Which deployment strategy should they use?
Medium264A machine learning engineer observes that a SageMaker training job fails with the error shown in the exhibit. What is the most likely cause of the failure?
Medium265A machine learning engineer needs to run a one-time scoring job over 500 GB of data stored in Amazon S3 using a trained model, and the results must be written back to S3. There is no requirement for a persistent HTTPS endpoint. Which SageMaker feature should the engineer use?
Easy266A company requires that all SageMaker notebook instances be created within a private VPC without internet access. Which configuration step is mandatory?
Easy267A company is fine-tuning a large language model using LoRA on SageMaker. They want to reduce GPU memory usage during training. Which configuration change would help?
Hard268A machine learning engineer is building a pipeline to preprocess text data for a sentiment analysis model. The data consists of customer reviews. The engineer wants to convert the text into numerical features while preserving the semantic meaning of words. Which technique should be used?
Medium269Refer to the exhibit. A user launches a SageMaker notebook instance with this lifecycle configuration. What happens?
Easy270A data scientist trains a binary classification model using SageMaker and obtains an AUC of 0.95 on the test set. However, the precision-recall curve shows low precision for high recall thresholds. The business requires a model that performs well on the minority class. Which metric should the team primarily optimize during hyperparameter tuning?
Hard271A data engineer is preparing a dataset for a time series forecasting model. The dataset contains a timestamp column and a target variable. The engineer wants to create additional features such as lag values and rolling averages. Which SageMaker Data Wrangler transform should be used to generate these time series features?
Medium272A data science team deployed a model on Amazon SageMaker and enabled Model Monitor to detect data drift. After a week, they receive alerts indicating that the distribution of a key feature has shifted significantly. However, the model's accuracy on the recent production data remains high. Which action should the team take next?
Medium273A data science team deploys a regression model using Amazon SageMaker. After one week, the model's prediction accuracy drops significantly. The team needs to detect this degradation automatically and trigger retraining. Which AWS service should they use to monitor the model's performance over time and set up alerts?
Easy274A data scientist notices that a production model's accuracy has degraded over the past week. The training data distribution remains unchanged, but the relationship between features and the target has shifted. Which type of drift is occurring, and which monitoring approach should be used?
Medium275A company needs to detect bias in a pre-trained model before deployment. They want to compute metrics like disparate impact and equal opportunity difference. Which AWS service should they use?
Hard276A data engineer needs to prepare a large dataset (10 TB) stored in Amazon S3 for a training job on SageMaker. The data is in CSV format, but the training algorithm expects Parquet for performance. The engineer must transform the data with minimal cost and without writing custom code. Which service should be used?
Medium277A data scientist needs to create a feature group in Amazon SageMaker Feature Store for real-time recommendations. Which TWO configurations are required? (Select TWO.)
Medium278A data scientist is training a binary classifier on a highly imbalanced dataset (1:100 class ratio). The dataset contains 500,000 rows and 30 features. The data is stored in S3 in Parquet format. The data scientist wants to use SageMaker's built-in XGBoost algorithm. Which data preparation technique should the data scientist apply to best address the class imbalance without causing data leakage?
Hard279A machine learning team deploys a fraud detection model on a SageMaker endpoint. The model's predictions are used in real-time. The team wants to monitor for data drift by comparing incoming data distributions against a baseline created from the training data. Which SageMaker capability should they use?
Medium280An ML team uses SageMaker Pipelines to automate model retraining. They want to skip redundant training steps when input data has not changed. Which feature should they enable?
Hard281A company plans to deploy a large foundation model using SageMaker JumpStart. They are concerned about costs because the model will be used intermittently. Which deployment option is MOST cost-effective for intermittent traffic?
Medium282Refer to the exhibit. The training job failed. What is the MOST likely cause?
Hard283A team is deploying a model using SageMaker Pipelines. They have defined a pipeline with steps: preprocessing, training, evaluation, and conditional registration. The evaluation step produces a JSON file with metrics. If accuracy > 0.9, the model is registered; else, the pipeline fails. Which TWO statements about this pipeline are correct? (Choose TWO.)
Hard284A machine learning engineer is deploying a real-time inference endpoint on Amazon SageMaker AI for a fraud detection model. The model must serve predictions with consistent latency under 50 ms and the team expects traffic to fluctuate unpredictably, with occasional bursts. The engineer wants to automatically adjust the number of instances based on actual workload while minimizing cost during idle periods. Which SageMaker AI feature should the engineer configure?
Medium285A data scientist is training a deep learning model on SageMaker and notices that the training loss oscillates and does not converge. They want to debug this issue. Which SageMaker feature can they use to monitor and analyze the training process?
Easy286A team has a large number of models that need to be deployed for batch inference weekly. They want to minimize cost and management overhead. Which approach is MOST efficient?
Medium287A company wants to deploy a foundation model from SageMaker JumpStart with the lowest possible inference cost, given that latency requirements are flexible. They have a mix of traffic volumes. Which approach should they take?
Medium288A company wants to deploy a machine learning model using infrastructure as code to ensure reproducibility. They need to define the SageMaker Studio domain, user profiles, and the endpoint configuration. Which tool should they use?
Medium289An ML engineer creates a SageMaker inference pipeline with two containers: a preprocessor and a predictor. The preprocessor is a lightweight Python script that transforms input data. How should the engineer structure the endpoints to ensure both containers run sequentially?
Medium290Which SageMaker feature compiles a trained model into an optimized binary for a specific hardware target (e.g., Intel, ARM, NVIDIA, or edge devices) to improve inference performance?
Easy291A team is training a large language model using PyTorch on SageMaker. They need to reduce training time. The model has 10 billion parameters. Which distributed training strategy should they use?
Medium292A machine learning engineer is configuring a SageMaker Processing job that runs a custom container to compute bias metrics on a dataset containing personally identifiable information. The job reads input data from one S3 bucket and writes reports to another, and the security team requires that the container has no outbound internet access and that the input and output buckets are reached without traversing the public internet. Which configuration satisfies these requirements?
Hard293A company deploys a model for fraud detection. They want to monitor if the model's predictions become less accurate over time due to changes in the underlying data distribution, but they do not have immediate access to ground truth labels. Which type of drift should they monitor as a proxy?
Medium294A company is training a deep learning model with SageMaker and wants to reduce training time by using pipe mode instead of file mode for a large dataset stored as TFRecord files in Amazon S3. After switching the estimator's input mode to Pipe, the training job fails immediately with a dataset format error. The data scientist confirms the files are valid TFRecords and that the same script works with File mode. What is the most likely cause?
Hard295A company needs to serve real-time predictions from a large ensemble of three deep learning models, each requiring different inference environments (PyTorch, TensorFlow, MXNet). Which SageMaker endpoint type supports running multiple inference containers together?
Medium296A company wants to build a customer service chatbot that answers questions about their internal policy documents. The documents are updated monthly, and the team cannot afford to retrain a model each time. Which approach is MOST appropriate?
Medium297An organization stores raw data in Amazon S3 as CSV files. They need to perform serverless data transformation and convert the data to Parquet format for efficient ML training. Which AWS service is most appropriate?
Easy298A team is building a fraud detection model using SageMaker and wants to detect anomalies in user login events. Which SageMaker built-in algorithm is specifically designed for anomaly detection in event-based data?
Medium299A retail company has a SageMaker model that predicts customer churn. The model was trained on data that included a 'customer_zipcode' feature. After deployment, the data science team notices that the model's predictions for certain zip codes have become less accurate over time. They suspect that the relationship between zip code and churn has changed due to a recent relocation of a major employer. Which SageMaker monitoring capability should they use to detect this type of drift?
Medium300A data scientist is preparing text data for sentiment analysis. They need to convert the text into numerical features while reducing the impact of common words. Which feature extraction method should they use?
Medium301A financial services company operates a real-time inference endpoint for a fraud detection model on Amazon SageMaker. The model was trained on historical transaction data from 2023. Over the past month, the model's precision has dropped from 92% to 78%, while recall remains high at 95%. The data science team suspects data drift and has already enabled SageMaker Model Monitor with data capture and a baseline from the training data. The latest monitoring report indicates no statistically significant drift in any of the input features. The team also verified that the inference code and model artifact have not changed. Despite the stable feature distributions, the model is misclassifying an increasing number of legitimate transactions as fraudulent (false positives). The business is concerned about the impact on customer experience. What is the best course of action?
Hard302Refer to the exhibit. The data scientist wants to update the endpoint to use a new model version without downtime. Which approach should they use?
Easy303A team wants to secure SageMaker endpoints for a healthcare application. They must ensure data is encrypted at rest and in transit, and that the endpoint can only be accessed from within a VPC. Which THREE steps should they take? (Select THREE)
Medium304A data scientist is using Amazon SageMaker Data Wrangler to create a data preparation flow. After completing the flow, the scientist wants to export the processed data to a feature group in Amazon SageMaker Feature Store for reuse in multiple training jobs. Which export option should the scientist choose in Data Wrangler?
Medium305A data scientist needs to store training data in Amazon S3 and wants to optimize read performance for iterative training jobs. Which S3 feature should they use?
Easy306A fraud-detection team trains a model in SageMaker and wants to shift 10 percent of live prediction traffic to a newly retrained model to compare accuracy before a full cutover. The endpoint already serves the current model on one production variant. They need the endpoint to route a controlled fraction of requests to the new model without changing the client application. What should they do?
Medium307A team deploys a SageMaker real-time endpoint and configures it with an auto scaling policy targeting a variant. During a flash sale, traffic spikes and the team notices that the number of instances increases, but the average model latency still climbs above the target. The team wants the scaling behavior to react faster to sudden bursts without over-provisioning during steady periods. Which change should they make to the scaling policy?
Easy308A company ingests streaming transaction data from multiple sources using Amazon Kinesis Data Streams. The data must be transformed (e.g., JSON parsing, data type conversions) and then stored in Amazon S3 for ML training. The transformation logic may change over time. Which approach provides the greatest flexibility and ease of maintenance?
Medium309A company wants to use SageMaker to fine-tune a foundation model for a text generation task using RLHF (Reinforcement Learning from Human Feedback). Which THREE components are required in the RLHF pipeline?
Hard310A data scientist is using Amazon SageMaker Data Wrangler to prepare a dataset. They need to identify potential bias in the data before training. Which SageMaker feature should they use?
Easy311A machine learning engineer has trained a model in SageMaker and wants to deploy it to a real-time endpoint for low-latency inference. The model artifacts are stored in Amazon S3, and the engineer needs to create the endpoint with the least operational effort. Which sequence of actions should the engineer take?
Easy312A data scientist needs to run a hyperparameter tuning job for a PyTorch model using SageMaker. They want to use Hyperband for efficient resource allocation. Which tuning strategy should they select in the HyperparameterTuner?
Medium313A data science team deploys a real-time inference endpoint on Amazon SageMaker. They want to monitor for data drift in the input features over time. Which AWS service should they use to capture and analyze the input data distribution?
Easy314A data science team needs to deploy a PyTorch model for real-time inference with low latency. The model requires GPU acceleration. Which SageMaker endpoint configuration should they use?
Easy315A data scientist discovers that a dataset for binary classification contains 95% negative samples and 5% positive samples. Which technique is MOST appropriate to address the class imbalance?
Easy316A company wants to serve 200 different PyTorch models. Each model is small (under 1 GB) and only a fraction are used at any time. To minimize cost and management overhead, which SageMaker inference option should be used?
Medium317A machine learning engineer is evaluating a binary classification model for detecting fraudulent transactions. The dataset is highly imbalanced, and the cost of false negatives (missing a fraud) is very high. Which two evaluation metrics should the engineer consider? (Choose two.)
Hard318An ML team is using Amazon SageMaker Feature Store to serve features for both real-time inference and batch training. They need to ensure that training data uses feature values as they were at the time of each event. Which type of query should they use?
Medium319A company deploys a real-time inference endpoint with auto-scaling using a target tracking policy based on average Invocations per instance. They notice that during a traffic spike, the endpoint scales out too late, causing increased latency. They want to scale proactively before the spike. Which strategy should they implement?
Hard320A company wants to deploy a PyTorch model on SageMaker for real-time inference. Which two steps are required? (Select TWO.)
Medium321A company is building a sentiment analysis model for customer reviews. The text data contains many typos, abbreviations, and informal language. Which text preprocessing step would be most beneficial to support the ML model?
Medium322A company is using SageMaker Debugger to monitor a training job for a deep learning model. They want to detect when gradients become extremely large, which may cause training instability. Which built-in rule should they use?
Medium323A machine learning engineer is using Amazon SageMaker Experiments to track multiple training runs. They want to compare the performance of different hyperparameter configurations visually. Which SageMaker tool provides an interactive interface to compare experiments?
Easy324A team is fine-tuning a foundation model using reinforcement learning from human feedback (RLHF) on SageMaker. They have a dataset of human preferences. Which SageMaker capability is most suitable for the reward model training step?
Hard325A data scientist needs to annotate a large dataset of images for an object detection model. The team wants to minimize manual labeling effort and cost. Which Amazon SageMaker feature should they use?
Easy326A company is setting up a data pipeline to ingest streaming clickstream data from their website for real-time analytics and machine learning. The data must be reliably ingested, transformed, and stored in Amazon S3 for batch processing. Which combination of AWS services should be used?
Easy327A team is fine-tuning a Hugging Face transformer model on SageMaker. They need to use a custom training script with the Hugging Face Estimator. Which SageMaker feature does this represent?
Medium328A machine learning team notices an increase in 5XXError count for a SageMaker endpoint. They want to set up automated remediation. Which THREE actions should they take? (Select THREE)
Medium329A data scientist needs to version and manage multiple models for a team of five. The team frequently experiments with different algorithms and hyperparameters. They need a centralized registry to store, deploy, and compare model versions. Which AWS service should the data scientist use?
Easy330A data scientist suspects that a deep learning model is overfitting. They enable SageMaker Debugger and want to detect overfitting automatically. Which built-in rule should they use?
Medium331A machine learning engineer is monitoring a deployed model for data drift. The input features are a mix of categorical and numerical columns. The baseline is from the training data. Which SageMaker Model Monitor feature should they enable to detect changes in the distribution of each feature over time?
Medium332A machine learning team deploys a model for loan approval. They want to monitor data drift on the real-time endpoint using SageMaker Model Monitor. Which set of actions should they take to set up data quality monitoring?
Medium333A company is using SageMaker endpoints for inference. To reduce costs, they want to use Automatic Scaling. However, they observe that scaling up takes several minutes, causing latency spikes during traffic bursts. What should they do to mitigate this?
Medium334Which SageMaker built-in algorithm is best suited for detecting anomalous login attempts based on IP addresses and user behavior?
Easy335A company needs to deploy a model that processes large payloads (up to 1 GB) asynchronously. The results should be written to S3, and the team needs SNS notifications upon completion. Which SageMaker inference option is MOST suitable?
Easy336A data engineer needs to ingest streaming data from IoT devices into Amazon S3 for machine learning. The data arrives continuously and must be available for querying within minutes. Which service should be used to collect and deliver the streaming data to S3?
Easy337A company uses Amazon SageMaker Feature Store to manage features for real-time inference. They need to store customer transaction data that is updated frequently and must be available for low-latency lookups. Which type of Feature Store should be used?
Medium338A company runs an online retail business and wants to build a product recommendation system. They have a dataset of customer purchases stored in Amazon S3 as CSV files. The dataset includes columns: 'customer_id', 'product_id', 'purchase_date', 'quantity', 'price', and 'category'. The data science team plans to use Amazon SageMaker to train a factorization machines model. During data exploration, they discover that the 'category' column has 1,200 unique values, and many categories appear only a few times. The 'product_id' column has 50,000 unique values. They want to include both features in the model. The team is concerned about the high cardinality of these features. Which approach should they take to prepare these features for the factorization machines model?
Medium339A company deploys a large NLP model on a SageMaker real-time endpoint using an ml.p3.2xlarge instance. To reduce inference cost without sacrificing throughput, they want to compile the model for their target hardware. Which service should they use?
Hard340A company ingests daily log data into an S3 bucket. They need to update the existing ML training dataset with new data without reprocessing the entire history. Which two strategies should they adopt? (Choose two.)
Easy341A healthcare company uses Amazon SageMaker to deploy a real-time inference endpoint for a diagnostic model. The endpoint is configured with a single ml.p3.2xlarge instance. The model processes patient data and returns a risk score. Recently, the endpoint has been experiencing intermittent 504 errors along with increased latency. The team uses Amazon CloudWatch to monitor the endpoint's InvocationsPerInstance and ModelLatency metrics. They observe that InvocationsPerInstance is well below the throttling threshold, but ModelLatency shows periodic spikes lasting 5-10 seconds. The endpoint's CPU utilization remains below 60%, but memory utilization occasionally spikes to 90% during those spikes. The team has checked the inference code and found no obvious memory leaks or performance bottlenecks in the custom logic. The model itself is a deep neural network hosted using Apache MXNet. The team suspects that the issue might be related to resource contention or an external dependency. What should the team do FIRST to diagnose and resolve the issue?
Hard342A machine learning engineer is preparing a dataset for training a SageMaker built-in Linear Learner model for binary classification. The dataset contains a highly imbalanced target with only 2% positive examples. They want to improve the model's ability to detect positives without collecting more data. Which SageMaker Linear Learner hyperparameter should they adjust to assign more weight to the positive class?
Hard343A machine learning model is deployed on SageMaker and its predictions are used in a production application. The model's accuracy has degraded over time. What is the most likely cause?
Medium344A data scientist is training a model on text data from customer reviews. The dataset contains a mix of English and Spanish reviews. The scientist wants to convert the text into numerical features for a classification model. Which approach is MOST appropriate for this multilingual dataset?
Medium345A fraud-detection team runs a SageMaker real-time endpoint in a production account. Their security team requires that the endpoint be reachable only from within a specific Amazon VPC and that access to invoke the endpoint be governed by identity-based policies with least privilege. Which TWO configurations should the ML engineer implement to meet these requirements? (Choose two.)
Hard346An ML engineer needs to monitor the operational health of a SageMaker endpoint, specifically the time taken for the container to process an inference request and the overhead added by SageMaker. Which two CloudWatch metrics should they examine?
Easy347Refer to the exhibit. A data scientist is trying to use AWS Glue to read data from the S3 bucket `ml-data-bucket`. The Glue job fails with an access denied error. What is the most likely cause?
Easy348A team is evaluating classification models for a medical diagnosis application. The cost of a false negative is much higher than the cost of a false positive. Which metric should be optimized during model selection?
Medium349A machine learning team trains a model in SageMaker and wants to track every step — from dataset version to hyperparameters to final model artifact — for reproducibility and audit compliance. Which SageMaker feature should they use?
Medium350A data scientist trained a model using SageMaker and wants to automate the retraining process when new data becomes available. Which AWS service is best suited to trigger a SageMaker training job based on an S3 event?
Easy351A data scientist is using SageMaker Experiments to track multiple training runs. They want to compare the F1 scores across runs. Which component should they use to log the F1 score?
Medium352A company is fine-tuning a large language model using reinforcement learning from human feedback (RLHF). Which THREE components are typically required?
Hard353Refer to the exhibit. A SageMaker training job logs show training AUC increasing but validation AUC plateauing at 0.880. What is the most likely issue?
Hard354A company deploys a deep learning model to a real-time SageMaker endpoint. After deployment, users report high inference latency. Which action is the MOST effective first step to reduce latency?
Easy355A company is training a binary classifier in SageMaker and observes that the training loss decreases but validation loss increases after a few epochs. What is the most likely issue?
Easy356Which feature scaling method is most robust to outliers in the data?
Medium357A company has a dataset of 2 billion records stored as text files in Amazon S3. The data is partitioned by year and month. The data science team wants to read only the last 6 months of data for model training using SageMaker. To minimize data scanned and reduce costs, which approach should the team use?
Easy358A team monitors a production endpoint and notices a sudden increase in 5XXError count. Which of the following is the most likely cause?
Medium359A data scientist is performing feature selection for a linear regression model and wants to remove features that are highly correlated with each other to reduce multicollinearity. Which technique is BEST suited for this purpose?
Easy360A team is using SageMaker Pipelines to train a model. The pipeline has multiple steps: data processing, training, evaluation, and registration. They use a Condition step to evaluate the model's accuracy and if it exceeds a threshold, register the model. They run the pipeline and the training step succeeds, but the pipeline fails at the Condition step with an error: 'Unable to evaluate condition: the property 'Accuracy' does not exist.' The evaluation step output is a JSON file with key 'accuracy'. What is the most likely cause?
Hard361A company is building a real-time inference pipeline for an ML model. The raw data arrives in JSON format via Amazon Kinesis Data Streams. Before invoking the SageMaker endpoint, the data must be preprocessed to match the training data format. Which THREE steps should be included in the preprocessing function? (Select THREE)
Hard362A team uses SageMaker Clarify to monitor bias drift in production. They schedule weekly analysis. After a month, Clarify reports a significant increase in a bias metric. What should the team do first?
Medium363A financial services firm is training a fraud detection model using SageMaker. The dataset is highly imbalanced (0.1% fraudulent transactions). The model currently achieves 99.9% accuracy but only catches 5% of fraud cases. Which metric should the team prioritize to evaluate model performance?
Hard364A data engineer is ingesting streaming clickstream data from a website into Amazon S3 for ML training. The data arrives at a rate of 10,000 events per second, and the team needs near-real-time availability with minimal transformation. Which AWS service should the engineer use to ingest the data into S3 with the LEAST operational overhead?
Hard365A company has deployed a machine learning model on Amazon SageMaker and wants to automatically detect when the distribution of input features deviates significantly from the training data distribution. Which SageMaker feature should they use?
Easy366A company wants to use SageMaker Clarify to analyze bias in their training data and model predictions. Which TWO types of bias can Clarify detect? (Choose TWO.)
Easy367A data science team is using Amazon SageMaker to train and deploy a binary classification model. They want to continuously monitor the model for data drift in production. Which combination of AWS services and SageMaker features should they use to implement automated drift detection with minimal operational overhead?
Medium368A data scientist is using Amazon SageMaker Data Wrangler to prepare a dataset. The dataset contains a column with date strings in the format 'YYYY-MM-DD'. The data scientist wants to extract the year, month, and day as separate features. Which Data Wrangler transform should be used?
Medium369A machine learning engineer is using Amazon SageMaker Data Wrangler to prepare a dataset for a regression model. The dataset contains a categorical feature with high cardinality (over 10,000 unique values). The engineer wants to encode this feature efficiently without creating thousands of binary columns, which would explode the feature space. Which encoding technique should the engineer use in Data Wrangler?
Hard370A company deploys a model for fraud detection. They need to monitor for bias after deployment, specifically whether the model's false positive rate changes across demographic groups over time. Which SageMaker feature should they use?
Hard371A data scientist is using SageMaker Automatic Model Tuning to optimize hyperparameters for an XGBoost model. They want to maximize AUC. Which search strategy is MOST appropriate for efficient exploration?
Medium372A machine learning engineer is preparing data for a SageMaker training job and needs to split a large dataset into training, validation, and test sets while avoiding data leakage from the same entity appearing in multiple splits. The dataset contains multiple rows per customer, and the target is customer churn. Which TWO strategies are appropriate? (Choose two.)
Hard373A company is building a fraud detection model on credit card transactions. The dataset contains a column 'merchant_id' with 50,000 unique values, many with low frequency. The team wants to avoid overfitting while preserving predictive signal. Which feature engineering approach is most appropriate?
Hard374A company is building a recommender system using implicit feedback (clicks) and explicit feedback (ratings). They plan to use Amazon SageMaker to train a model. The data includes user ID, item ID, timestamp, and rating (if any). Which TWO data preparation steps should the team perform? (Choose TWO.)
Medium375A machine learning engineer is preparing a dataset for training a model on Amazon SageMaker. The dataset contains numerical features with varying scales, and the engineer wants to ensure that all features contribute equally during training. Which data preparation step should the engineer take?
Easy376Which TWO actions are recommended best practices for securing an Amazon SageMaker notebook instance? (Select TWO.)
Easy377A company is training a large NLP model on SageMaker and wants to reduce costs by using Spot Instances. Which TWO configurations should they implement to handle Spot interruptions gracefully?
Medium378A company wants to deploy a new model using a canary deployment strategy on SageMaker. Which two actions should they take? (Select TWO.)
Medium379A team has a SageMaker Pipeline that trains a model and registers it in the Model Registry. They want to automate the deployment of the approved model to a staging environment. Which event-driven approach should they use?
Medium380A company is using SageMaker Automatic Model Tuning to optimize a regression model. They want to minimize the root mean squared error (RMSE). The tuner has completed 20 jobs, and the RMSE has plateaued. Which action should the data scientist take to potentially improve the results?
Medium381A data scientist wants to fine-tune a Llama 2 7B model using SageMaker for a text summarization task. The dataset is 10 GB. The budget is limited, so cost efficiency is important. Which THREE steps should the data scientist take? (Choose THREE.)
Medium382A machine learning engineer is using SageMaker to train a model and wants to automatically stop a training job when the validation loss has not improved for 10 consecutive epochs, while still saving the best model artifacts. The engineer is using the SageMaker training toolkit in a custom container. Which combination of actions should the engineer take?
Medium383A company wants to reduce costs for a real-time inference endpoint that experiences predictable traffic spikes during business hours and low traffic at night. Which auto-scaling policy is MOST cost-effective while maintaining performance?
Easy384A healthcare analytics team trains models in SageMaker and stores artifacts in an S3 bucket that contains protected health information. An auditor asks how the team can prove which training dataset and container image produced the model currently deployed to production, and wants the evidence retained even if someone deletes the training job. Which SageMaker capability should the team rely on to capture and retain this metadata automatically?
Medium385A data scientist is using SageMaker Experiments to track multiple training runs for a PyTorch model. They want to compare metrics across runs and identify the best hyperparameters. Which TWO capabilities should they use? (Choose TWO.)
Medium386A data scientist is preparing a dataset for a multi-class classification problem. The dataset contains a categorical feature with 50,000 unique values (high cardinality). The scientist wants to reduce dimensionality while preserving predictive information. Which TWO approaches are appropriate? (Choose 2)
Hard387A team wants to automatically retrain a model when new labeled data arrives. Which SageMaker feature can orchestrate this workflow?
Easy388A team uses SageMaker ML Lineage Tracking to capture the metadata of their ML workflow. They want to query the lineage to see which model version was trained from a specific dataset. Which Lineage Tracking entity represents the dataset?
Medium389A data scientist wants to normalize a feature to have a range between 0 and 1 for a neural network. Which scaling technique should be applied?
Easy390A machine learning engineer is preparing a dataset for a binary classification model. The dataset has 10,000 samples with a 1:100 class imbalance. The engineer needs to balance the classes before training. Which technique would create a balanced dataset without discarding majority class samples and without generating synthetic data?
Medium391A data scientist is building a time-series forecasting model for daily sales data. The data spans two years. To evaluate the model's performance, the data scientist needs to simulate a realistic rolling forecast scenario. Which data splitting strategy should be used?
Hard392A machine learning engineer needs to standardize features to have zero mean and unit variance before training a support vector machine. Which scaling method should they apply?
Easy393A company has a SageMaker endpoint that uses a trained model to classify images. The endpoint is experiencing high latency and the team suspects it is due to the model size. Which action can the team take to reduce latency without significantly impacting accuracy?
Easy394A hospital deploys a model to predict patient readmission risk. To comply with regulations, they must ensure that the model's predictions do not show bias against any demographic group over time. Which service should they use for ongoing monitoring?
Hard395A data scientist is using SageMaker Data Wrangler to prepare a large dataset. The data contains duplicate rows, which could bias the model. Which built-in step in Data Wrangler can automatically detect and remove duplicates?
Medium396A financial services company is developing a fraud detection model using Amazon SageMaker. They have a dataset with 10 million transactions, each with 300 features. The dataset is highly imbalanced (0.1% fraud). They have performed feature engineering and now need to split the data for training, validation, and test sets. The data is stored in CSV files in Amazon S3. They plan to use SageMaker's built-in XGBoost algorithm. To ensure proper evaluation and avoid data leakage, which data splitting strategy should they use?
Hard397A company runs a regression model to predict house prices. They have 50 features including 'zip_code' (high cardinality), 'square_footage', and 'year_built'. They want to select the most important features to reduce overfitting. Which feature selection method is computationally efficient for high-dimensional data and can handle multicollinearity?
Hard398A team uses SageMaker Ground Truth to create labeled datasets. They need to ensure labeling jobs are cost-effective. Which TWO measures should they take? (Select TWO.)
Easy399A machine learning team is building a model to predict customer churn. They have historical data that includes customer activity logs, each with a timestamp. The team wants to ensure that the training data does not contain any data leakage from the future. Which approach should they take when preparing the training and validation datasets?
Hard400A company stores its raw IoT sensor data in Amazon S3. The data is in CSV format and contains timestamps, sensor IDs, and readings. A data engineer needs to catalog this data for discoverability and querying by other team members. Which AWS service should they use to create a searchable metadata catalog?
Easy401A company is using SageMaker to train a neural network for image classification. The training job is taking too long. The team wants to reduce training time without sacrificing model accuracy. Which approach should they recommend?
Medium402A model deployed on SageMaker is returning inaccurate predictions for certain customer segments. The team suspects data drift. Which SageMaker feature should they use to continuously monitor input data distribution?
Hard403A team wants to apply a custom container for inference on SageMaker. The container needs to implement a web server that responds to API requests. Which protocol and port must the container listen on to be compatible with SageMaker hosting?
Easy404A startup wants to deploy a containerized ML application that includes both a model inference server and a preprocessing component in the same endpoint. Which SageMaker endpoint type supports running multiple containers?
Medium405A company uses Amazon SageMaker to deploy a real-time inference endpoint. They notice increased latency in predictions during peak hours. Which should they investigate first to address the issue?
Easy406A company uses SageMaker to train a model. They want to ensure that training data is encrypted at rest and in transit, and that only authorized users can access the training artifacts. Which three steps should they take? (Choose three.)
Hard407A retail company is building a machine learning model to predict customer churn. The data engineering team has extracted customer transaction data from Amazon Aurora and stored it as CSV files in Amazon S3. The data includes customer IDs, transaction amounts, timestamps, and product categories. A data scientist discovers that the dataset contains several missing values in the 'transaction_amount' column for about 15% of the records. The data scientist also notices that the 'customer_id' column has some duplicate entries. The team wants to prepare the data for training a churn model using Amazon SageMaker. The data is approximately 50 GB in size. What should the data scientist do to handle the missing values and duplicates efficiently while preparing the data for training?
Easy408A data scientist is preparing a large dataset for training a machine learning model. The dataset contains missing values in several columns. Which approach is the MOST efficient for handling missing values in a large dataset using AWS services?
Easy409A data engineer needs to ingest streaming clickstream data from a website into an S3 data lake for ML training. The data arrives continuously and must be written to S3 in near real-time. Which AWS service is best suited for this task?
Easy410A machine learning team deploys a custom container image for an Amazon SageMaker training job. The container needs to access an S3 bucket that contains sensitive data. The team wants to follow the principle of least privilege. How should the team grant access?
Medium411A company wants to update an existing SageMaker real-time endpoint to serve a new model version. They need to route a small percentage of traffic to the new version initially and monitor for errors before switching fully. Which deployment pattern supports this?
Easy412During model training on Amazon SageMaker, the training job fails with a 'ResourceLimitExceeded' error. What is the most likely cause?
Medium413A team wants to track and compare multiple machine learning experiments, including hyperparameters, metrics, and artifacts. They are using Amazon SageMaker. Which AWS service or feature should they use to achieve this?
Easy414An organization wants to ensure that only approved model versions can be deployed to production. They use the SageMaker Model Registry to track model versions. How can they enforce that only approved models are deployed?
Medium415A machine learning engineer is training a SageMaker job with the TensorFlow estimator and wants to automatically capture model training metadata such as loss curves and accuracy for later comparison, without writing any custom code. Which SageMaker feature should they enable?
Medium416A company uses Amazon SageMaker to deploy a model for real-time inference. They want to perform A/B testing between two model versions. Which TWO actions should the company take to set up A/B testing? (Choose TWO.)
Medium417A machine learning engineer trains a binary classifier in SageMaker and the model outputs class probabilities. The business requires that the model achieve at least 90% recall on the positive class, while keeping precision above 70%. The engineer uses the default threshold of 0.5 when deploying. Which approach should the engineer take to meet these requirements?
Medium418A company wants to maintain multiple versions of a trained model in a central repository and track metadata such as training metrics, hyperparameters, and approval status. Which SageMaker feature should they use?
Easy419Refer to the exhibit. A data scientist configured SageMaker Debugger to monitor training for overfitting. However, the rule never triggers even though the model appears to be overfitting. What is the most likely reason?
Hard420A financial services company uses Amazon SageMaker to deploy a fraud detection model for real-time inference. The model is deployed on an ml.m5.large instance with a SageMaker real-time endpoint. The endpoint has an auto scaling policy configured using a custom scaling policy based on average CPU utilization, with scale out threshold at 70% and scale in threshold at 30%. During a flash sale event, the traffic to the endpoint spikes tenfold within minutes. The endpoint fails to handle the load, resulting in increased latency and timeouts. The data science team needs to improve the scalability of the endpoint to handle sudden traffic spikes. Which solution should the team implement?
Hard421A company is building a real-time fraud detection system using Amazon Kinesis Data Streams. The data must be joined with a reference table (e.g., customer profile) that is stored in Amazon DynamoDB and updated frequently. The enriched data will be used for ML predictions. Which THREE AWS services should the company use to build this streaming pipeline? (Select THREE.)
Medium422A company is using Amazon SageMaker to train a model on sensitive customer data. The security team requires that all data be encrypted in transit and at rest, and that the training job does not have internet access. Which configuration should the team use to meet these requirements?
Easy423A machine learning engineer wants to deploy a pre-trained foundation model for text summarization using SageMaker JumpStart. Which of the following is a primary cost consideration when deploying such a model?
Easy424A data engineer is building a feature store using Amazon SageMaker Feature Store. The team needs to store features that are updated frequently and require low-latency retrieval for real-time inference. Which type of store should the engineer use?
Easy425A healthcare company deploys a model to predict patient readmission risk. The model was trained on historical data and is now showing signs of concept drift. The team needs to implement a monitoring solution that can detect drift and automatically retrain the model when drift is detected. Which THREE steps should the team take to build this solution? (Choose THREE.)
Hard426A team wants to use SageMaker Clarify to monitor bias in their production model predictions. They have configured a bias drift monitor. What does SageMaker Clarify compare to detect bias drift?
Medium427A data scientist is using SageMaker built-in XGBoost algorithm for a regression problem. Which metric is most appropriate as the objective metric for hyperparameter tuning?
Easy428An ML engineer needs to split a dataset into training, validation, and test sets. The dataset has a time-based column that should not be leaked. Which split method is most appropriate?
Easy429A data engineer is designing an ETL pipeline using AWS Glue to transform raw data from S3 into a curated set for ML training. The data contains personally identifiable information (PII) that must be masked before being used by data scientists. Which TWO actions should the engineer take? (Choose TWO.)
Medium430A data scientist is training a binary classification model using Amazon SageMaker. The dataset has a severe class imbalance (95% negative, 5% positive). The model achieves 99% accuracy but fails to identify positive cases correctly. Which action should the data scientist take to improve the model's ability to detect positive cases?
Medium431A company deploys a model using SageMaker real-time endpoint with auto scaling. They observe that during a traffic spike, the endpoint quickly scales up to 10 instances, but after the spike, it takes a long time to scale down, leading to high costs. The scaling policy is based on a simple average CPU utilization threshold. Which adjustment would optimize the scaling down behavior?
Hard432An MLOps engineer is setting up a SageMaker endpoint for a model that performs inference on large images. The model is containerized and expects input in a specific format. The team wants to preprocess the images (resize and normalize) before passing them to the model. What is the most efficient way to implement this?
Medium433A data engineer is using AWS Glue to run an ETL job that joins two large datasets and writes the output to S3 for ML training. The job is failing due to out-of-memory errors. Which THREE actions can help resolve this issue? (Select THREE.)
Hard434A healthcare analytics team trains models in Amazon SageMaker and needs an immutable, queryable record of which dataset version and training job produced each registered model version, so an auditor can trace a deployed model back to its inputs months later. Which SageMaker capability should they rely on?
Easy435A machine learning engineer is preparing a dataset for training a SageMaker model. The dataset contains missing values in several numerical features. The engineer wants to handle these missing values during the training pipeline. Which two methods are valid ways to handle missing values in SageMaker? (Choose two.)
Medium436A company is using SageMaker to train a model for image classification. The training dataset contains 100,000 labeled images. The team wants to use a pre-trained model to reduce training time. Which SageMaker feature should they use?
Medium437A company uses SageMaker Neo to compile a trained model for deployment on edge devices. What is the primary benefit of using Neo?
Easy438A company wants to track the lineage of their ML models for reproducibility and auditability. Which THREE services or features should they use together to achieve this? (Choose THREE.)
Medium439A data science team needs to deploy a trained PyTorch model for real-time inference with sub-100ms latency. The model fits on a single GPU. Which SageMaker inference option is MOST cost-effective while meeting the latency requirement?
Easy440A company wants to reduce costs for a SageMaker real-time endpoint that has variable traffic. Which feature allows the endpoint to automatically adjust instance count based on demand?
Easy441A financial services company has a SageMaker real-time endpoint serving a fraud detection model. Compliance requires that all inference requests and responses be logged with the ability to detect anomalous input feature distributions over time. The team wants a managed solution that captures request/response payloads to Amazon S3 and automatically computes statistics and constraints against a baseline. Which combination of SageMaker features should they enable?
Medium442An ML engineer is fine-tuning a foundation model using RLHF on SageMaker. Which THREE components are essential for this workflow? (Select THREE.)
Hard443A data scientist is training a deep learning model on Amazon SageMaker and notices that the training loss decreases but the validation loss starts increasing after a certain number of epochs. The model is likely overfitting. Which SageMaker feature can they use to detect and diagnose this issue during training?
Medium444A company is building a fraud detection model on an imbalanced dataset (99% legitimate, 1% fraudulent). To improve recall on the minority class, they want to resample data. Which combination of techniques should they use?
Medium445A company wants to deploy a scikit-learn model to a SageMaker AI real-time endpoint. The model must be loaded from a custom Python module that contains preprocessing logic not present in the built-in scikit-learn container. The team wants to minimize operational overhead and does not need to change system-level libraries. Which approach should the engineer take?
Medium446An ML engineer has a real-time SageMaker endpoint serving a fraud-detection model. The team wants to release a new model version to a small percentage of live traffic first, monitor CloudWatch metrics for accuracy regressions, and roll back quickly if performance degrades. They also want the production and candidate variants to share the same endpoint so latency comparisons are apples-to-apples. Which SageMaker deployment strategy should they use?
Medium447A team wants to fine-tune a pre-trained Hugging Face transformer model for text classification using SageMaker. They have a custom training script. Which SageMaker estimator should they use?
Easy448A financial services company uses SageMaker to train a fraud detection model. They have imbalanced data with 1% fraud. They trained a Gradient Boosting model using SMOTE for oversampling and achieved 99% accuracy on the test set, but the fraud recall is only 10%. The data scientist is concerned about the model's performance. Which change is most likely to improve fraud recall without sacrificing too much precision?
Medium449A machine learning team has a model that needs to serve predictions with very low latency (under 10 ms) for a real-time web application. The model is a small ensemble of three neural networks that fits in memory. Which SageMaker inference option is MOST appropriate?
Medium450A company deployed a machine learning model on an Amazon SageMaker real-time endpoint. Over several weeks, they notice that inference latency has been gradually increasing, especially during peak business hours. The model and instance type have remained unchanged. What is the most likely cause of the increased latency?
Easy451A data science team deploys a regression model to Amazon SageMaker for real-time inference. After one month, the model's prediction errors increase significantly, but data distributions remain unchanged. Which monitoring approach is MOST suitable for detecting this issue?
Easy452A team needs to deploy a new model version to production while minimizing risk. They want to route 5% of live traffic to the new model and 95% to the current model, and then gradually increase the new model's traffic. Which SageMaker deployment pattern should they use?
Medium453A machine learning engineer is setting up an Amazon SageMaker notebook instance. The instance needs to access a private S3 bucket that contains training data. The notebook instance is in a VPC. Which combination of steps will grant access to the S3 bucket? (Choose TWO.)
Easy454An ML platform team must orchestrate a workflow that trains a model, evaluates it against a baseline, and only registers the model if evaluation passes. If evaluation fails, the workflow must notify the data science channel and stop without registering. The team wants the orchestration logic to be expressed as code, versioned in git, and integrated with SageMaker training jobs and Model Registry. Which approach best fits these requirements?
Hard455A company uses Amazon SageMaker Ground Truth to label images for object detection. They want to minimize labeling costs while maintaining high accuracy. Which feature should they enable?
Medium456A trained model needs to be deployed for real-time inference with low latency. Which AWS service is best suited for this?
Medium457A machine learning engineer is preparing a training dataset stored in Amazon S3 for a SageMaker training job. The data is in CSV format, and the engineer wants to ensure that the training job can access the data efficiently and securely. The S3 bucket is in the same AWS Region as the SageMaker training job. The engineer needs to provide the training job with the necessary permissions to read the data. Which of the following is the MOST secure and appropriate way to grant the training job access to the S3 bucket?
Medium458A data science team is using Amazon SageMaker Pipelines to orchestrate a multi-step workflow that includes data preprocessing, training, and model evaluation. They want to reuse the preprocessed data across multiple pipeline executions without re-running the preprocessing step if the source data hasn't changed. What should they configure?
Hard459A data scientist is training a large model on SageMaker and wants to reduce training time by using multiple GPUs. The model is small enough to fit on a single GPU but training is slow. Which SageMaker feature should be used?
Medium460A data scientist is using SageMaker built-in XGBoost algorithm for a binary classification task. Which objective metric is MOST appropriate for SageMaker Automatic Model Tuning to maximize?
Easy461A machine learning engineer is training a deep learning model on SageMaker and notices that the training loss decreases rapidly in the first few epochs but then plateaus. The validation loss starts increasing after 10 epochs. Which action should the engineer take to improve generalization?
Medium462A machine learning engineer manages a SageMaker Model Monitor schedule for a real-time endpoint. The monitor's baseline was computed from a training dataset with a categorical feature named region. In production, a new category value appears that was never seen in training, and the monitor begins reporting violations. The engineer wants the monitor to flag only the appearance of unknown categories without treating normal distribution shifts in known categories as violations. Which approach should the engineer take?
Hard463An ML engineer needs to compile a trained TensorFlow model to run efficiently on a target edge device with an ARM CPU. Which AWS service should they use?
Easy464A data scientist wants to track the training and validation accuracy of a SageMaker training job over time. They need to visualize these metrics in Amazon CloudWatch. Which action should they take?
Easy465A company wants to monitor its Amazon SageMaker real-time endpoint for data quality issues. Which TWO actions should the company take?
Easy466A company uses SageMaker JumpStart to deploy a foundation model for a summarization task. They want to minimize costs while still meeting a latency requirement of under 2 seconds. Which option should they consider?
Medium467A financial services company ingests transaction data from multiple sources into an S3 data lake. They want to use AWS Glue to catalog this data and make it queryable by Amazon Athena. The data schema changes frequently as new sources are added. Which AWS Glue feature should they enable to automatically detect and update the schema?
Medium468An ML team uses SageMaker to deploy a model for real-time inference. They want to monitor and improve cost efficiency. Which THREE actions should they take? (Select THREE.)
Hard469A data science team deploys a machine learning model to a SageMaker endpoint for real-time inference. They need to monitor the model for feature distribution drift over time to ensure the model's predictions remain accurate. Which AWS service should they use?
Easy470A machine learning engineer needs to split a dataset for binary classification where the positive class represents only 2% of the data. Which data splitting strategy ensures that both training and test sets maintain the same class proportion as the original dataset?
Easy471A data scientist is training a binary classification model using Amazon SageMaker. The dataset is highly imbalanced (95% negative class, 5% positive class). The model is evaluated on a held-out test set, and the F1 score is 0.12. The data scientist wants to improve the F1 score. Which two actions should the data scientist take? (Choose two.)
Medium472A company runs a real-time fraud detection model on a SageMaker endpoint. The model is updated weekly, and each update must be validated against live traffic without affecting existing predictions. The team wants to compare the new model's performance against the current model using a small percentage of incoming requests, while ensuring that the current model continues to serve the majority of traffic. Which SageMaker deployment strategy should they use?
Medium473A hospital's ML team deploys a diagnostic model to a SageMaker real-time endpoint. Compliance requires that every inference request and response be recorded for later auditing, and the records must be retrievable months later. The team needs a low-effort way to capture this data. Which approach should they use?
Easy474A company is using SageMaker Pipelines to automate a multi-step ML workflow. The pipeline includes data preprocessing, training, and model evaluation. The team wants to ensure that if the evaluation step fails, the pipeline stops and sends an alert to the operations team. Which SageMaker Pipelines feature should they use?
Easy475Which THREE steps should be taken to optimize a large-scale distributed training job on SageMaker? (Choose 3.)
Hard476A data science team uses SageMaker Pipelines to orchestrate their ML workflow. They noticed that even when source data hasn't changed, the pipeline re-runs all steps, wasting compute time. What should they enable to avoid redundant runs?
Hard477A company is building a time series forecasting model using SageMaker DeepAR. The raw data is a CSV with columns: timestamp, item_id, and value. What is the correct data format required for DeepAR training?
Medium478A data scientist is preparing a dataset for a machine learning model that predicts customer churn. The dataset contains a column 'CustomerID' that is a unique identifier. What should the data scientist do with this column before training the model?
Easy479A team uses SageMaker Pipelines to automate retraining. They want to skip the training step if the data has not changed since the last run. Which feature should they enable?
Medium480A data scientist is using SageMaker Experiments to track multiple training runs. They want to compare runs based on the objective metric and visualize performance. Which THREE steps should they perform? (Choose THREE.)
Hard481A company wants to deploy a machine learning model that was trained on-premises using TensorFlow. The model is a TensorFlow SavedModel. The company uses AWS and wants to minimize operational overhead. Which deployment option meets these requirements?
Easy482A machine learning engineer deploys a new model version to a SageMaker endpoint with production variants. They want to gradually shift traffic from the old model to the new model, monitoring for errors, and automatically roll back if the error rate exceeds 5%. Which deployment pattern should they use?
Hard483A data scientist wants to quickly build a binary classification model without writing any code. Which SageMaker feature is MOST suitable?
Easy484A data scientist needs to ingest streaming clickstream data from a website into an S3 data lake for ML training. The data must be processed in near real-time and partitioned by hour. Which AWS service combination should be used?
Medium485A machine learning engineer needs to split a dataset into training, validation, and test sets for a SageMaker training job. The dataset is stored in Amazon S3 as a single CSV file. The engineer wants to ensure that the splits are reproducible and that the test set is never used during training or hyperparameter tuning. Which approach should the engineer use?
Easy486A machine learning team stores training data in an Amazon S3 bucket and wants to catalog it so that Amazon Athena and Amazon SageMaker Feature Store can discover the schema. The data is partitioned by year, month, and day in Hive-style prefixes, and new partitions are added daily. A data engineer must ensure new partitions are automatically discoverable without manual intervention. Which solution meets these requirements?
Medium487A company has 200 small PyTorch models that are each used infrequently but need to be available for real-time inference. To minimize costs, they want to host all models on a single endpoint. Which SageMaker feature should they use?
Medium488An organization wants to automate ML retraining using an event-driven architecture. Which THREE services should they combine? (Select THREE.)
Medium489A machine learning engineer is building a real-time fraud detection pipeline using Amazon Kinesis Data Streams. The data must be prepared (e.g., feature engineering, normalization) before being fed into a SageMaker endpoint. Which TWO steps should the engineer implement to ensure low-latency data preparation?
Medium490A company wants to deploy a trained XGBoost model for batch inference on a large dataset stored in S3. The inference job should be cost-effective and does not require real-time responses. Which SageMaker inference option should they use?
Easy491A machine learning engineer needs to deploy a new version of a model gradually, initially sending 5% of traffic to the new version and 95% to the current version, while monitoring for errors. Which deployment pattern should they use?
Easy492A media company uses SageMaker endpoints to serve a model that predicts video engagement. They have two production variants: Variant A (ml.c5.large) for regular traffic and Variant B (ml.c5.xlarge) for burst traffic. They use weighted routing (90% to A, 10% to B). Recently, during peak hours, Variant A's latency increase causes many requests to time out. The metrics show that both variants are under similar CPU load, but the number of concurrent requests to Variant A is very high. The team wants to ensure that burst traffic is handled properly without manual intervention. What should they do?
Medium493A retail company runs a SageMaker real-time endpoint serving a demand forecasting model. Security policy requires that all inference requests travel over the AWS private network and never traverse the public internet, and that the endpoint cannot be invoked from outside the company VPC. The endpoint already uses a customer-managed KMS key for volume encryption. Which TWO configurations should the engineer apply to meet these requirements? (Choose two.)
Hard494A company deploys a SageMaker model using AWS KMS for encryption at rest. They have a compliance requirement to rotate the KMS key every year without causing downtime for the inference endpoint. Which approach should they take?
Hard495Which TWO data storage options are commonly used by Amazon SageMaker Feature Store for offline and online storage?
Easy496A data scientist is trying to create a SageMaker endpoint configuration with 6 instances of ml.c5.large for a production variant. The creation fails with the error shown in the exhibit. Which action should the data scientist take to resolve this issue?
Hard497A data scientist is training a model with SageMaker and needs to reduce the cost of a long-running training job that can tolerate interruptions. The job uses a custom training script and reads data from Amazon S3. The data scientist wants the job to resume from the last saved state if the underlying compute is reclaimed. (Choose two.)
Hard498A machine learning engineer observes that model performance on a SageMaker endpoint has degraded over the past week. Ground truth labels are available with a 2-day delay. The engineer wants to automatically trigger a retraining pipeline when prediction quality drops below an acceptable threshold. Which approach is most appropriate?
Medium499A team wants to monitor the number of requests and latency of their SageMaker endpoint using a unified dashboard. Which AWS service should they use to create a custom dashboard with these metrics?
Easy500A team uses SageMaker real-time endpoints for inference. They want to deploy a new model version and compare its performance with the current version under live traffic without affecting user experience. Which method should they use?
Hard501A company is deploying a large language model (LLM) to a SageMaker endpoint. They want to minimize inference latency and cost by using GPU acceleration and model parallelism. The model is too large to fit on a single GPU. Which SageMaker feature should they use?
Hard502A data engineer stores raw ML training data in Amazon S3 and needs to catalog the schema, track partition changes, and make the data queryable by Amazon Athena without running ETL. Which AWS service should the engineer use?
Easy503A data engineer needs to convert a JSON dataset to Parquet format for efficient querying with Amazon Athena. The JSON files are in an S3 bucket. Which service can perform this conversion with minimal coding?
Easy504A company wants to automatically trigger a retraining pipeline when concept drift is detected in their deployed model. Which combination of services should they use?
Medium505A company uses SageMaker endpoints with auto-scaling. The endpoint is experiencing high latency during peak hours. The metrics show CPU utilization is low but memory is high. What is the most likely cause?
Medium506Which TWO actions are recommended best practices when preparing training data for a machine learning model in AWS? (Choose two.)
Easy507A team is using SageMaker Pipelines to automate retraining and deployment. They want to trigger the pipeline automatically when new training data is available in an S3 bucket. Which approach should they use?
Medium508A machine learning team runs a SageMaker AI Pipeline that trains a model and registers it in the SageMaker AI Model Registry. A separate deployment process must promote the model to production only after a human reviewer approves the model version. The team wants to automate the promotion so that approval in the Model Registry triggers the deployment without manual intervention. Which combination of steps should the engineer implement?
Hard509Refer to the exhibit. A user is unable to invoke a SageMaker endpoint. The IAM policy shown is attached to the user. Which permission is missing to allow invocation?
Easy510A company uses Amazon SageMaker Feature Store to store features for a real-time recommendation model. The feature data is updated continuously, and the model must use the most recent feature values for each user at inference time. Which type of Feature Store should the company use for serving features to the model?
Hard511A machine learning team needs to monitor a deployed model for both data drift and concept drift. Which TWO approaches should they implement? (Select TWO.)
Medium512A machine learning engineer is deploying a model to a SageMaker endpoint that must handle occasional large payloads up to 1 GB. The inference time can take up to 10 minutes. The team wants to minimize cost and avoid idle compute. Which deployment option is most appropriate?
Medium513An ML engineer runs the CLI command shown in the exhibit. However, the training job fails immediately with an error: 'Unable to assume role'. What is the most likely cause?
Easy514A company is using Amazon SageMaker to deploy a model for real-time inference. The model requires access to a private S3 bucket that contains reference data. The company wants to ensure that the endpoint can access the S3 bucket without using a public internet connection. Which TWO actions should they take? (Select TWO.)
Easy515A data engineer is using Amazon SageMaker Ground Truth to create a labeled dataset for an object detection task. The dataset contains millions of images, and the labeling budget is limited. Which approach can reduce labeling costs while maintaining high model accuracy?
Medium516A company wants to deploy a PyTorch model on SageMaker using the NVIDIA Triton Inference Server for GPU acceleration. They have an existing Triton configuration. Which approach should they take?
Medium517A team uses AWS Auto Scaling for a SageMaker real-time endpoint. They notice that when scaling in, the latest instance is always terminated first, causing disruption to recent requests. How can they configure the scaling policy to terminate the oldest instance first?
Medium518A team is using Amazon SageMaker to train a neural network. They want to minimize training time while effectively exploring the hyperparameter space. Which approach should they use?
Medium519A company needs to update a model in production without any downtime. They currently have a single real-time endpoint serving traffic. Which approach allows them to deploy a new model version and switch traffic gradually while being able to roll back quickly?
Hard520An ML engineer has trained a model and stored the model artifacts in an Amazon S3 bucket in the same AWS Region as the planned SageMaker AI endpoint. During endpoint creation, the engineer must specify the S3 location of the model artifacts. Which permission must the SageMaker AI execution role have for the endpoint to load the model successfully?
Easy521An ML team wants to use Amazon SageMaker Ground Truth to create a labeled dataset for a multi-class image classification task. They have a large set of unlabeled images and want to minimize labeling costs while maintaining high accuracy. Which Ground Truth feature should they enable?
Easy522A data scientist deploys a model and wants to monitor the endpoint's invocation latency. They notice that the CloudWatch metric 'ModelLatency' is high, but 'OverheadLatency' is low. Which statement correctly interprets these metrics?
Medium523A data engineer is preparing a large dataset of 10 TB for ML training on Amazon SageMaker. The data is stored in Amazon S3 as CSV files. To reduce training time and cost, the engineer wants to use a columnar format that is optimized for analytical queries. Which format should the engineer convert the data to?
Easy524A small team runs a SageMaker real-time endpoint in production. They want a low-effort way to know when the endpoint's invocations are failing so they can react quickly, and they want the alert delivered to their on-call channel. Which approach requires the least custom code?
Easy525A data scientist is using SageMaker to train a linear regression model. After training, they evaluate the model on the test set and get an R² of 0.95. However, when they deploy the model to a SageMaker endpoint and run predictions on new data, the predictions are far off. What is the most likely cause?
Easy526A machine learning engineer notices that the latency of a SageMaker endpoint has increased over time. They need to identify which component (model inference vs. pre/post-processing) contributes most to the latency. Which CloudWatch metrics should they examine?
Medium527A healthcare startup has deployed a machine learning model on Amazon SageMaker that predicts patient readmission risks. The model uses sensitive health data stored in an S3 bucket encrypted with AWS KMS. The SageMaker endpoint is configured with an IAM role that has the following policy attached: { "Version": "2012-10-17", "Statement": [ { "Effect": "Allow", "Action": "s3:*", "Resource": "arn:aws:s3:::healthcare-data/*", "Condition": { "Bool": { "aws:SecureTransport": "true" } } }, { "Effect": "Allow", "Action": "kms:Decrypt", "Resource": "*" } ] }. During a security audit, the team discovers that the IAM role's KMS permission is too permissive because it allows decryption of any KMS key in the account. The team needs to modify the policy to follow the principle of least privilege while still allowing the SageMaker endpoint to read the encrypted data. Which modification should the team make?
Hard528A data scientist is training a model using SageMaker and wants to use spot instances to reduce costs. The training job is checkpointed every 5 minutes. However, the job gets interrupted frequently and never completes. What is the MOST likely cause?
Hard529A data engineer is using Amazon SageMaker Data Wrangler to prepare a dataset for a machine learning model. The engineer needs to ensure that the data preparation flow is reproducible and can be integrated into an MLOps pipeline. Which TWO actions should the engineer take to achieve this? (Choose two.)
Medium530A data scientist is using Amazon SageMaker Ground Truth to create a labeled dataset for an object detection model. The dataset contains 1 million images, and the team wants to reduce labeling cost by labeling only the most informative samples. Which feature of Ground Truth should they use?
Medium531A company has deployed a SageMaker real-time endpoint for a model that predicts customer churn. The endpoint uses a single ml.m5.large instance. After deployment, the team notices that during peak hours, the endpoint returns 5xx errors for about 20% of requests. The endpoint has not been configured with any scaling policy. The team needs to resolve this issue with minimal cost increase. Which solution should the team implement?
Easy532A data engineer is building an AWS Glue ETL job that reads raw JSON clickstream events from Amazon S3, flattens nested structures, and writes Parquet to a curated S3 prefix for SageMaker training. The job must run daily on only the newly arrived files and must keep the Glue Data Catalog table current so Athena and SageMaker can query it. Which TWO configurations should the engineer apply to meet these requirements? (Choose two.)
Hard533A machine learning engineer is deploying a pre-trained NLP model on Amazon SageMaker for real-time inference. The model expects input sequences of variable length, and performance is critical. The engineer wants to minimize latency while handling the variable-length inputs efficiently. Which approach should the engineer choose?
Hard534A machine learning engineer is using SageMaker Autopilot for AutoML. Which TWO outputs does Autopilot produce?
Medium535A data science team is using AWS Step Functions to orchestrate a machine learning workflow that includes a SageMaker training job followed by a model deployment. They want to ensure that if the training job fails, the workflow retries up to three times with exponential backoff before sending a notification to an Amazon SNS topic. Which Step Functions feature should they use to implement this?
Medium536A company uses SageMaker Pipelines to automate their ML workflow. They notice that the pipeline reruns all steps even when the input data has not changed. Which feature should they enable to avoid unnecessary recomputation?
Easy537A company uses SageMaker endpoints for real-time inference. They want to automatically scale the number of instances based on the number of outstanding requests. Which auto-scaling policy type should they choose?
Medium538A team is fine-tuning a Hugging Face BERT model for text classification using SageMaker. They want to use the Hugging Face estimator for convenience. Which parameter must be set to use a custom training script?
Medium539A machine learning engineer is training a TensorFlow model using SageMaker with distributed training. They need to implement data parallelism across multiple GPUs. Which SageMaker feature should they use to distribute the training?
Medium540Refer to the exhibit. A data scientist reviews the output of a SageMaker training job. The model has 95% training accuracy and 92% validation accuracy. Which statement is true?
Easy541An e-commerce company uses a SageMaker endpoint to serve a product recommendation model. The model is retrained every month using batch transforms. The ML team has set up a retraining pipeline using SageMaker Processing jobs and Step Functions. Recently, the Step Functions workflow has been failing at the retraining step with an error: 'AccessDeniedException: User: arn:aws:sts::123456789012:assumed-role/RetrainingRole/abc123 is not authorized to perform: s3:GetObject on resource: arn:aws:s3:::training-data/processed/latest.parquet'. The team confirms that the S3 bucket exists and the object is present. The retraining role has the following policy: { "Version": "2012-10-17", "Statement": [ { "Effect": "Allow", "Action": [ "s3:GetObject", "s3:PutObject" ], "Resource": "arn:aws:s3:::training-data/*" } ] }. The team also verifies that the bucket policy does not explicitly deny access. What is the MOST likely cause of the AccessDenied error?
Easy542A company uses SageMaker to deploy a model and wants to perform A/B testing by splitting traffic between two model variants. Which TWO actions should they take? (Select TWO.)
Medium543A data scientist is using SageMaker Autopilot to automatically build a model. Which TWO aspects does Autopilot handle? (Choose TWO.)
Easy544A team is fine-tuning a large language model (LLM) using SageMaker and wants to reduce memory footprint during training. Which technique should they use?
Medium545Which TWO options are recommended best practices for monitoring model performance in production on SageMaker? (Choose 2.)
Medium546A machine learning engineer needs to give a data scientist read-only access to the model artifacts, training metrics, and monitoring reports stored in a single Amazon S3 bucket used by a SageMaker project, while ensuring the data scientist cannot delete or overwrite any object. Which approach follows least-privilege practice?
Easy547A data scientist is using SageMaker Automatic Model Tuning to find the best hyperparameters for a model. They want to reduce the total tuning time for a given number of training jobs. Which tuning strategy should they choose?
Easy548A data science team has trained a model using SageMaker and wants to deploy it to a production endpoint with automatic scaling based on request volume. Which SageMaker feature should they use to configure scaling?
Easy549A team uses MLflow on SageMaker for experiment tracking. They want to automate the retraining of a model when new training data arrives in an S3 bucket. Which combination of services should they use?
Medium550A company uses SageMaker Inference Recommender to select the optimal endpoint configuration. After running the recommender, they receive a recommendation for a specific instance type and initial instance count. What should they do next to optimize costs over time?
Medium551A data scientist creates a feature group as shown in the exhibit. When ingesting data with an 'age' column of integer values, the ingestion fails. What is the most likely cause?
Hard552A machine learning engineer is building a time-series forecasting model to predict daily sales for the next 30 days. The dataset spans two years of daily sales data. To evaluate model performance, the engineer needs to simulate a realistic forecasting scenario where the model is trained on past data and tested on future data without leakage. Which data splitting strategy should they use?
Hard553A company uses SageMaker Studio for collaborative ML development. The security team requires that all SageMaker Studio notebooks run within a VPC and cannot access the public internet. Which configuration should the administrator set?
Easy554A company needs to secure a SageMaker notebook instance that contains sensitive data. Which THREE of the following are effective security measures? (Select THREE.)
Hard555A company collects sensor data from IoT devices. The data arrives with missing timestamps due to network issues. For anomaly detection, the engineer needs to create features that capture rolling statistics over fixed windows. Which data preprocessing step is essential before feature generation?
Medium556A data scientist is using SageMaker Automatic Model Tuning with Hyperband. They want to stop poorly performing trials early to save resources. Which strategy does Hyperband use?
Medium557An organization wants to schedule a retraining pipeline to run every Sunday night. Which AWS service should they use to trigger the pipeline on a schedule?
Easy558A data engineer is using AWS Glue to prepare a dataset for machine learning. The dataset has several columns with outliers. The engineer wants to detect and handle outliers in a scalable manner. Which TWO approaches should the engineer consider? (Select TWO.)
Easy559A machine learning team at a bank is training a binary classification model using SageMaker's built-in XGBoost algorithm on a dataset with 20 million rows and 300 features. They need to reduce training time while maintaining model accuracy. The data is stored in Amazon S3 as CSV files. Which approach should they take to speed up training?
Medium560A team is monitoring a SageMaker endpoint and notices that the average latency (ModelLatency) is increasing over time, but the number of invocations is steady. They suspect that the model's inference code is becoming slower due to memory leaks. Which metric should they also examine to confirm this hypothesis?
Medium561A company is using Amazon SageMaker Data Wrangler to prepare a dataset for training. They have created a data flow with multiple transforms. Which TWO actions can they take to operationalize the data preparation pipeline for production? (Choose 2)
Medium562A machine learning engineer is using SageMaker Debugger to monitor a training job and wants to detect issues early. The engineer wants to receive alerts when the training job is likely to fail due to vanishing gradients and when the loss is not decreasing. Which two actions should the engineer take to achieve this? (Choose two.)
Medium563A team needs to split a time-series dataset for a forecasting model. They want to avoid data leakage and evaluate model performance on future unseen data. Which data splitting strategy should they use?
Medium564A machine learning engineer needs to deploy a TensorFlow model that requires a custom inference environment with specific system libraries. The model will be used in a real-time application with variable traffic. They want to minimize cold start latency. Which SageMaker hosting option should they choose?
Medium565An ML engineer wants to be notified when the average inference latency of a SageMaker endpoint exceeds 500 ms for 2 consecutive evaluation periods. Which AWS service combination should they use?
Easy566A team uses SageMaker Neo to compile a model for deployment on a target device. After compilation, they deploy the compiled model to a SageMaker endpoint using the Neo-optimized container. The endpoint fails to start with error "RuntimeError: Unable to load model". What could be the issue?
Hard567A company wants to use SageMaker Autopilot for a regression problem. They require an explainability report that shows feature importance globally. Which Autopilot feature should they enable?
Medium568A company uses SageMaker to train a model. The training job is failing with an error "ResourceLimitExceeded". What is the most likely cause?
Easy569A machine learning engineer needs to reduce costs when training a large model on SageMaker. They are willing to accept potential interruptions and have checkpointing enabled. Which instance purchasing option should they use?
Easy570A company wants to deploy a model using a serverless inference endpoint that can automatically scale to zero when not in use and has a configurable maximum concurrency. Which SageMaker inference option meets these requirements?
Easy571A company wants to automate the deployment of a SageMaker model into production whenever a new model version is approved in the Model Registry. Which service can be used to trigger the deployment pipeline?
Easy572A machine learning engineer is evaluating a binary classification model that predicts customer churn. The model achieves 95% accuracy, but the engineer suspects class imbalance is causing a misleading metric. Which THREE evaluation steps should the engineer perform to properly assess the model? (Choose THREE.)
Hard573A company uses SageMaker Model Monitor for feature attribution drift monitoring with SHAP. Which THREE prerequisites must be in place before starting the monitoring schedule? (Select THREE)
Medium574A data science team uses Amazon SageMaker Model Monitor to detect data drift in production. They notice that the schema of incoming data (number of features) has changed compared to the training baseline. Which type of monitor is BEST suited to detect this issue?
Medium575A data scientist must join a 50 GB transactional table with a small 5 MB lookup table in AWS Glue before writing Parquet output for SageMaker training. The join currently shuffles the large table across the cluster and the job runs slowly. Which optimization should the data scientist apply?
Medium576An ML team is deploying a model using SageMaker. The model requires GPU inference and must be available in multiple AWS regions for low latency. The team has created a multi-model endpoint with GPU instances. After deployment, they notice high latency spikes when a new model is loaded. What is the most likely cause?
Medium577A financial services company must ensure that all data used by Amazon SageMaker training jobs is encrypted at rest. The company wants to use a customer-managed key (CMK) for the encryption. Which steps are necessary to achieve this? (Choose TWO.)
Hard578A team wants to automate the retraining and deployment of an ML model whenever new labeled data arrives in S3. The workflow includes data preprocessing, training, evaluation, and conditional deployment. Which AWS service is best suited for orchestrating this end-to-end pipeline?
Easy579A data scientist is training an object detection model using SageMaker built-in Object Detection algorithm. They want to visualize the bounding boxes on validation images after training. Which approach should they use?
Medium580A data scientist wants to version control trained models and manage approvals for deployment. Which SageMaker feature should they use?
Easy581A company has a batch transform job in Amazon SageMaker that processes large datasets every night. Recently, the job has been failing sporadically with an out-of-memory error. The data size has not increased. What is the MOST likely cause?
Medium582A company is deploying a SageMaker real-time endpoint and needs to monitor inference latency. Which THREE metrics are available from SageMaker for this purpose? (Choose THREE.)
Medium583A machine learning engineer wants to automatically trigger a retraining pipeline whenever new training data arrives in an S3 bucket. The pipeline uses SageMaker Pipelines. Which AWS service should be used to detect the S3 event and start the pipeline?
Easy584A team built a SageMaker Pipeline that includes a training step and a model evaluation step. They want to automatically register a model in SageMaker Model Registry only if the evaluation metric (accuracy) exceeds 0.9. Which pipeline step should be used to implement this conditional logic?
Medium585An ML engineer is designing a SageMaker Pipeline for model training and registration. They need to ensure that the pipeline can be re-run with different datasets without manual intervention, and that the steps are only re-executed if inputs have changed. Which THREE features should they configure? (Select THREE.)
Hard586Refer to the exhibit. A team configured a SageMaker Model Monitor schedule for data quality. The baseline was created from a training dataset. After running for a day, the monitoring results show frequent violations. What is the most likely cause?
Easy587An ML engineer is fine-tuning a large language model using LoRA on SageMaker. The training is converging slowly, and GPU utilization is low. The engineer suspects the bottleneck is data loading. Which action should the engineer take to improve GPU utilization?
Hard588A data engineer is optimizing Amazon Athena queries on large datasets stored in S3 for machine learning data preparation. Which THREE practices improve query performance?
Hard589Refer to the exhibit. A user has the above IAM policy attached but cannot access files in SageMaker Studio. What additional permission is most likely needed?
Easy590A machine learning engineer needs to deploy a TensorFlow model to Amazon SageMaker and wants to use the built-in TensorFlow Serving container. What should the engineer provide in the model archive?
Easy591A company wants to deploy its trained model to edge devices such as cameras and IoT devices. The model must run efficiently with low latency and minimal memory footprint. Which THREE actions should the company take to prepare the model for edge deployment? (Choose THREE.)
Easy592A machine learning engineer is deploying a model to a SageMaker real-time endpoint that must be accessible only from within a specific Amazon VPC and must not have a public IP address. The engineer also needs to ensure that all data in transit between the endpoint and the calling application is encrypted. Which configuration should the engineer use?
Medium593A machine learning engineer needs to monitor a SageMaker endpoint for data drift and receive alerts when drift is detected. They want to use a fully managed AWS service to schedule the monitoring jobs and send notifications. Which AWS service should they use to orchestrate the monitoring schedule and trigger alerts?
Easy594A machine learning practitioner is building a binary classifier with severe class imbalance (1:1000). They want to use SMOTE for oversampling. What is a potential drawback of applying SMOTE on the entire dataset before splitting into training and test sets?
Hard595A team is training a large language model and needs to split the model layers across multiple GPUs due to memory constraints. Which distributed training strategy should they use?
Medium596A data engineer stores a training dataset in Amazon S3 across a prefix hierarchy organized by year, month, and day, and analysts query it with Amazon Athena. The engineer wants Athena to scan only the relevant date range instead of the entire dataset. Which action should the engineer take?
Easy597A team is migrating their ML infrastructure to AWS and wants to use infrastructure as code to manage SageMaker Studio domains, user profiles, and associated resources. Which services can they use for this purpose? (Select THREE.)
Medium598A team is training a deep learning model on Amazon SageMaker using a custom Docker container. Which three practices should they follow to optimize training performance? (Choose three.)
Medium599A company uses SageMaker Model Monitor to detect bias drift in their real-time inference endpoint. They have collected ground truth labels and want to monitor for bias across different demographic groups. Which type of monitoring should they configure?
Medium600A machine learning engineer is preparing a dataset for binary classification. The target variable has a severe class imbalance (95% negative, 5% positive). Which technique can help address this imbalance during data preparation?
Easy601An ML team is preparing time-series data for a demand forecasting model. They want to evaluate model performance over time without leaking future information into past training windows. Which data splitting strategy is MOST appropriate?
Medium602A team needs to deploy a PyTorch model that uses custom CUDA kernels. They want to use NVIDIA Triton Inference Server on SageMaker for high-performance serving. Which SageMaker configuration is required to use Triton?
Hard603A data scientist is preparing a dataset for training a binary classification model. The dataset has 100,000 rows and 50 features. The target variable is imbalanced, with only 5% positive cases. Which technique should the data scientist apply to address the class imbalance BEFORE training?
Easy604During deployment of a Hugging Face model, the endpoint logs show this error. Which step was likely missed?
Hard605A company is training a deep learning model on Amazon SageMaker. The training job started but has been stuck in 'InProgress' state for an unusually long time with low CPU utilization. The data scientist suspects a bottleneck. What should be the first troubleshooting step?
Medium606A machine learning engineer needs to prepare a dataset containing customer transactions for training a fraud detection model. The dataset includes features such as transaction amount, timestamp, merchant category, and customer ID. The engineer wants to create a feature representing the average transaction amount per customer over the last 7 days. Which approach should be used in Amazon SageMaker Data Wrangler?
Medium607A financial services company is deploying a fraud detection model on SageMaker. To comply with regulations, they must ensure that the model's predictions are not biased against protected groups. They plan to monitor bias drift post-deployment using SageMaker Clarify. Which data inputs are required to configure Clarify's bias drift monitoring?
Hard608A company deploys a model on Amazon SageMaker for real-time inference. The inference latency is too high. The model is a large deep learning model. The company wants to reduce latency without significantly impacting accuracy. Which approach should the company consider?
Medium609A data scientist is building a text classification model using Amazon SageMaker. The dataset is stored as a CSV file in Amazon S3. The scientist wants to use the SageMaker built-in BlazingText algorithm. Which of the following steps are required to prepare the data for training? (Choose TWO.)
Hard610A company wants to reduce costs for a SageMaker real-time endpoint that receives predictable traffic patterns: high during business hours and low at night. The model is a small PyTorch model. Which cost-saving strategy is most suitable?
Easy611A machine learning engineer needs to ingest streaming data from thousands of IoT devices into Amazon S3 for batch training. The data should be available in S3 within minutes of arrival. Which combination of services should the engineer use?
Medium612Which technique is commonly used to handle missing values in a categorical feature?
Easy613An MLOps engineer is building an automated retraining pipeline for a fraud detection model. The model must be retrained weekly, and the new model should only be promoted to production if it meets predefined performance thresholds compared to the current model. Which combination of SageMaker capabilities should the engineer use?
Hard614A machine learning engineer is using SageMaker Pipelines to orchestrate a training workflow. The pipeline includes a processing step that outputs a dataset, which is then used by a training step. The engineer notices that the processing step runs every time the pipeline executes, even when the input data has not changed. The engineer wants to avoid re-running the processing step if the input data and code are unchanged, while ensuring that downstream steps still execute if the processing step is skipped. Which approach should the engineer take?
Hard615A media company stores training data in an S3 bucket encrypted with an AWS KMS customer-managed key. A SageMaker training job runs inside a private VPC subnet with no internet access and must read that bucket. The job currently fails with an access-denied error from S3. Which change most directly resolves the failure while preserving the private-network requirement?
Hard616A team has deployed a real-time inference endpoint and wants to automatically scale based on CPU utilization. Which scaling policy type should they use with Application Auto Scaling for SageMaker endpoints?
Medium617A company uses Amazon SageMaker to train and deploy a machine learning model. After deployment, they notice that the model's accuracy drops significantly over time due to changes in the underlying data distribution. Which monitoring solution should they implement to detect this issue automatically?
Medium618A team uses SageMaker for training. They need to monitor training progress and view metrics like loss and accuracy. Which SageMaker feature should they use?
Easy619A data scientist wants to evaluate the performance of a binary classification model. The dataset is highly imbalanced with only 5% positive class. Which metric should be used to evaluate the model?
Easy620A team has 200 small ML models that need to be served via HTTPS endpoints. Each model is used infrequently, and the team wants to minimize hosting costs. Which SageMaker deployment approach is MOST cost-effective?
Medium621A machine learning team uses SageMaker Pipelines to automate retraining. They want to avoid re-running data processing steps if the data has not changed since the last successful pipeline run. Which built-in feature should they enable?
Medium622A model deployed on a SageMaker endpoint is returning predictions. The team wants to log all predictions to an S3 bucket for auditing. What is the most efficient way to achieve this?
Medium623A data scientist observes that a linear regression model has many irrelevant features. They want to perform feature selection to improve generalization. Which method combines feature selection with model training using a penalty that can shrink coefficients to zero?
Hard624A team is training a large deep learning model on SageMaker using a single ml.p3.16xlarge instance. Training is taking too long. They want to reduce time by distributing across multiple GPUs but are constrained by model size that does not fit in a single GPU memory. Which distributed training strategy should they use?
Medium625A machine learning team is building a model to predict customer churn. The dataset includes a feature 'customer_tenure' with values ranging from 1 to 100 months, and 'monthly_spend' ranging from $10 to $5000. The model will use gradient boosting. Which feature scaling approach is most appropriate?
Hard626A healthcare company is building a model to predict patient readmission rates. The dataset contains a mix of numeric features (age, blood pressure, lab test results) and categorical features (gender, diagnosis code, hospital department). The dataset has 2 million rows. The data is stored in an Amazon S3 bucket, and they use AWS Glue to catalog and preprocess the data. The data scientist notices that the 'diagnosis_code' column has 10,000 unique codes, and 20% of the rows have missing values for 'blood_pressure'. They plan to use a SageMaker built-in XGBoost model. For optimal model performance, which preprocessing steps should they apply using AWS Glue ETL?
Medium627A company deploys a model on SageMaker that serves predictions to a web application. The model's performance degrades over time due to data drift. The company wants to set up continuous monitoring. Which TWO actions should the company take to monitor and retrain the model effectively? (Choose TWO.)
Medium628A company is using AWS Step Functions to orchestrate their ML retraining pipeline. They want to trigger retraining when new data arrives, but only if the model's performance has degraded below a threshold. Which THREE AWS services should they use together to achieve this? (Choose three.)
Medium629A data scientist is building a text classification model using Amazon SageMaker. The dataset is large and includes imbalanced classes. Which three techniques can help improve model performance? (Choose three.)
Medium630Which TWO tools are specifically designed for debugging and analyzing training jobs in SageMaker?
Hard631A data scientist is using Amazon SageMaker Debugger to monitor training metrics. They want to stop training automatically if the model is overfitting. Which action should they take?
Hard632Which SageMaker feature provides AutoML capabilities, including automatic data preprocessing, model selection, and hyperparameter tuning?
Easy633A company uses SageMaker Pipelines to automate their ML workflow. They need to add model versioning and approval workflow. Which THREE steps should they include in their pipeline to achieve this? (Choose THREE.)
Medium634Which THREE components are required to set up automated model retraining in response to performance degradation using Amazon SageMaker? (Select THREE.)
Medium635A company wants to trigger a model retraining pipeline whenever new training data arrives in an S3 bucket. They also need to send a notification to a Slack channel when the retraining completes. Which TWO AWS services should they use to implement this event-driven workflow? (Select TWO.)
Easy636A data scientist is building a regression model to predict house prices. The dataset contains a feature 'neighborhood' with 500 distinct values, and most neighborhoods have fewer than 10 samples. Which approach is MOST appropriate for handling this high-cardinality categorical feature?
Medium637A company has 50 small PyTorch models that are used infrequently for inference. They want to minimize costs while maintaining the ability to serve all models from a single endpoint. Which SageMaker feature should they use?
Easy638A machine learning engineer needs to select features for a regression model. The dataset contains 50 numeric features, and the target variable is continuous. The engineer wants to reduce dimensionality by selecting features that have the strongest linear relationship with the target. Which feature selection method is MOST appropriate?
Medium639Which SageMaker built-in algorithm is designed for time series forecasting?
Easy640A data scientist is evaluating a binary classification model. They have the confusion matrix and want to assess the model's performance comprehensively. Which THREE metrics should they consider? (Select THREE.)
Medium641A machine learning engineer is training a model using SageMaker and wants to set up monitoring to detect if gradients become too large, which could destabilize training. Which SageMaker Debugger built-in rule should they enable?
Medium642A company is using Amazon SageMaker to train a large deep learning model. The training job is taking a very long time. The data scientist suspects that the GPU utilization is low due to inefficient data loading. Which action should the data scientist take to diagnose and address this issue?
Medium643Refer to the exhibit. A team observes that their SageMaker endpoint scales out quickly when load increases, but scales in very slowly when load decreases, causing over-provisioning. What is the most likely cause?
Medium644A machine learning team needs to automatically retrain a model when concept drift is detected in the deployed endpoint's predictions. Which TWO steps should they take? (Choose TWO.)
Medium645A machine learning engineer is building an ML pipeline using Amazon SageMaker. The engineer needs to prepare the data, detect bias in the dataset, and then create features for training. Which TWO AWS services or features should the engineer use? (Choose TWO.)
Medium646A machine learning engineer is developing a text classification model using Amazon SageMaker. The dataset consists of 1 million customer reviews, with labels indicating sentiment (positive, negative, neutral). The engineer uses a pre-trained BERT model from the Hugging Face Model Hub and fine-tunes it on the dataset using SageMaker's Hugging Face estimator with a ml.p3.2xlarge instance. After 2 hours of training, the training job fails with a 'ResourceExhaustedError: CUDA out of memory' error. The error occurs during the forward pass of the first epoch. The engineer confirms that the batch size is set to 32, the maximum sequence length is 512 tokens, and the dataset is stored in a S3 bucket in the same AWS region. The engineer needs to complete fine-tuning without increasing instance costs. Which course of action should the engineer take?
Medium647A company wants to reduce costs for a production SageMaker endpoint that has predictable traffic patterns. They have purchased a Savings Plan. What additional step can they take to further optimize costs while maintaining performance?
Easy648A data scientist uses SageMaker Model Monitor to track feature attribution drift. Which technique does SageMaker Model Monitor use to compute feature attributions?
Medium649A team is using AWS Glue to process streaming data from Amazon Kinesis. The streaming data contains both structured and semi-structured fields. The team needs to flatten the semi-structured fields into columns for downstream ML training. Which Glue feature is BEST suited?
Hard650A machine learning team needs to deploy a PyTorch model that has been compiled with SageMaker Neo to improve inference performance on edge devices. Which TWO statements about SageMaker Neo are correct? (Select TWO.)
Medium651A company uses SageMaker endpoints with auto-scaling based on CPU utilization. During a flash sale, latency increases despite low CPU. What should be done?
Hard652A data scientist needs to evaluate a binary classification model. The dataset is highly imbalanced (5% positive class). Which metric is MOST appropriate for assessing model performance?
Medium653A company wants to track the lineage of their ML models, including the training dataset, hyperparameters, and training job used to produce each model version. Which AWS service should they use?
Easy654A data scientist wants to use SageMaker Autopilot to automatically build a regression model. The dataset contains 200 features and 50,000 rows. Which output does SageMaker Autopilot provide?
Easy655A company has a SageMaker endpoint running a model that provides real-time recommendations. Recently, the model's accuracy has degraded due to data drift. The team wants to automatically retrain the model when a drift metric exceeds a threshold and deploy the new model without downtime. Which architecture should the team implement?
Hard656A data scientist is preparing a dataset for binary classification. The dataset has a target variable with 90% of samples belonging to class 0 and 10% to class 1. Which data splitting strategy should the scientist use to ensure that the training and test sets maintain the same class proportion as the original dataset?
Easy657A company deploys a model with SageMaker and wants to monitor for concept drift. They have noticed that the relationship between input features and the target variable has changed, causing model accuracy to degrade. However, the input data distribution remains stable. Which type of drift is this, and what is the most appropriate response strategy?
Hard658Refer to the exhibit. A data scientist tries to deploy a model from an S3 bucket encrypted with SSE-KMS. What should the administrator do to resolve this?
Medium659A machine learning engineer is using Lasso regression for feature selection. After training, many coefficients become zero. The engineer notices that some features with high mutual information with the target also have zero coefficients. What is the most likely reason?
Hard660A data scientist is using SageMaker to train an XGBoost model for a regression problem. After training, they evaluate the model on a test set and get an RMSE of 10 and an R² of 0.85. Which additional metric would give the MOST insight into the model's average prediction error magnitude?
Medium661An ML team is building a time-series forecasting model for daily sales. They need to split the data into training and validation sets without data leakage, and the validation set should be the most recent 30 days. Which splitting strategy should they use?
Hard662A financial services company needs to deploy a machine learning model for real-time fraud detection. The model must be highly available across multiple Availability Zones and must support automatic scaling based on request volume. The company also needs to perform canary deployments to test new model versions with a small percentage of traffic before full rollout. Which SageMaker feature should they use?
Hard663A machine learning engineer is preparing a training script that must run on multiple GPU instances with SageMaker. The script currently reads the entire training dataset from local disk into memory, which fails on larger datasets. The engineer wants the script to stream training data from the SageMaker training channel path without loading everything into memory. Which approach should the engineer take?
Medium664An organization needs to ensure that all data transmitted between containers in a SageMaker training job is encrypted. In the training job configuration, which setting should they enable?
Hard665A data scientist wants to fine-tune a large language model for a question-answering task. They want to reduce memory usage during training by using a low-rank approximation of the weight updates. Which technique should they use?
MediumOther domains
All MLA-C01 exam domains
Frequently asked questions
- What does the scenario questions domain cover on the MLA-C01 exam?
- scenario questions questions test whether you can apply the concept in context, not just recognise a definition.
- How many questions are in this domain?
- This page lists all 665 scenario questions questions in the MLA-C01 question bank. The actual exam draws from this domain proportionally to its weighting in the official exam blueprint.
- What is the best way to practise this domain?
- Start with a short focused session (10 questions) to identify gaps, then work through explanations. Repeat with a longer session once the weak areas feel solid.
- Can I practise only scenario questions questions?
- Yes — the session launcher on this page filters questions to this domain only. Choose any session length for inline explanations and scoring.