Courseiva

MLA-C01 · domain

ML Model Development

This domain covers building, training, tuning, and evaluating models on SageMaker, roughly a quarter of the MLA-C01 exam. Questions test built-in algorithm selection, Autopilot outputs, handling imbalanced data, and choosing the right metric or job type for a stated business problem.

108 questions22 easy60 medium26 hard

Focused practice

Practice ML Model Development questions

Scored sessions drawing only from this domain — pick a length below.

Start 20-question practice test →

What this domain covers

What to know about ML Model Development

Be able to pick the right SageMaker algorithm, metric, and job configuration for a described problem, then interpret results. The single most important thing: match the evaluation metric to the business objective, especially on imbalanced datasets.

Selecting SageMaker built-in algorithms such as DeepAR for time series forecasting

Using SageMaker Autopilot and reading its model explainability and candidate reports

Choosing evaluation metrics like recall, precision, or F1 for imbalanced classification

Configuring training, tuning, and inference jobs with the SageMaker Python SDK

Watch out for

Common ML Model Development exam traps

  • ▸Optimizing for accuracy on imbalanced data, which hides poor fraud or anomaly recall.
  • ▸Confusing forecasting algorithms with general regression, or picking XGBoost for seasonal time series.
  • ▸Assuming Autopilot only returns a model, ignoring the notebooks and explainability artifacts it also produces.

Question index

All ML Model Development questions (108)

Click any question to see the full explanation, or start a practice session above.

1

A financial institution is training a fraud detection model using SageMaker. The dataset is highly imbalanced, with only 0.1% fraudulent transactions. The team wants to use SageMaker Automatic Model Tuning to find the best hyperparameters. They notice that the tuning job spends most of its time on configurations that predict all transactions as non-fraudulent. Which hyperparameter should they tune to directly address this issue?

Hard
2

A team is training a PyTorch model using SageMaker. They have a custom training script that requires specific Python packages not included in the SageMaker default PyTorch container. Which approach should they use?

Medium
3

A data scientist is using SageMaker Autopilot to automatically build a binary classification model on a balanced dataset. They want to understand the relationship between the input features and the model predictions. Which feature in SageMaker Autopilot should they use?

Medium
4

A team is training a PyTorch model using SageMaker with a custom training script. They want to track hyperparameters and metrics across multiple experiments. Which service should they use?

Medium
5

A machine learning engineer is using SageMaker Automatic Model Tuning (AMT) to optimize a model. They want to ensure the tuning job explores the hyperparameter search space efficiently and stops poorly performing trials early. Which two strategies should they use? (Choose two.)

Medium
6

A machine learning engineer is using SageMaker Pipelines to automate the training and deployment of a model. The pipeline includes a processing step for feature engineering, a training step, and a model registration step. The engineer wants to ensure that the pipeline is reproducible and that the model artifacts are versioned. Which two actions should be taken? (Choose two.)

Medium
7

A company uses SageMaker Autopilot to build a binary classification model. The generated leaderboard shows an ensemble model as the best candidate. The team needs a model that can be deployed for real-time inference with latency < 10ms. What should they do?

Hard
8

A machine learning engineer is training a tabular regression model using the SageMaker built-in XGBoost algorithm. They want to reduce overfitting and improve generalization without changing the algorithm. Which SageMaker hyperparameter should they tune to control the fraction of features randomly sampled per tree?

Medium
9

A machine learning engineer is deploying a model to a SageMaker endpoint for real-time inference. The model must return predictions within 100 milliseconds for 95% of requests. The engineer wants to monitor the endpoint's latency and automatically roll back if latency exceeds the threshold. Which combination of SageMaker features should be used?

Medium
10

A company is using SageMaker to train a model with a custom container. The training script requires a specific version of a Python library that is not included in the default SageMaker containers. How should they provide this library?

Hard
11

A data scientist is using SageMaker Autopilot for a regression problem. They want to see which data preprocessing steps Autopilot applied. Which TWO sources can they use to find this information?

Medium
12

A data scientist is using SageMaker to train a deep learning model with the PyTorch estimator. They want to log custom scalar metrics such as validation accuracy and loss during training so they can monitor the job in SageMaker. Which approach should they use to emit these metrics from the training script?

Medium
13

A machine learning engineer is using a SageMaker training job with a custom training script. They need to save the trained model artifacts to Amazon S3 so that the model can be deployed later. Which parameter in the SageMaker estimator should they configure to specify the S3 location for model artifacts?

Medium
14

A financial services company is training a fraud detection model using SageMaker. The dataset is highly imbalanced, with only 0.2% fraudulent transactions. The team wants to optimize the model for recall at a fixed precision of 90%. They are using the SageMaker built-in XGBoost algorithm with binary:logistic objective. Which evaluation metric should they monitor during training and hyperparameter tuning?

Hard
15

A data scientist wants to use SageMaker Clarify to analyze bias during training of a binary classification model. Which TWO types of bias metrics can SageMaker Clarify compute? (Select TWO.)

Medium
16

A company is training a deep learning model for object detection using SageMaker. The training is very slow and the GPU memory is insufficient for the batch size. The team wants to scale across multiple GPUs efficiently. Which THREE actions should they take? (Choose THREE.)

Hard
17

A data scientist is using SageMaker to train a custom PyTorch model for image classification. They want to use SageMaker Debugger to detect training issues. Which TWO built-in rules are most relevant for detecting common training problems? (Select TWO.)

Medium
18

A machine learning engineer is using Amazon SageMaker Debugger to monitor a training job for a deep neural network. They receive a rule alert indicating 'exploding gradients'. Which action should they take to address this issue?

Hard
19

A data scientist is using SageMaker Experiments to track multiple training runs. They want to compare different hyperparameter configurations and visualize the impact on model accuracy. What should they use to track hyperparameters?

Medium
20

A company is using SageMaker Autopilot to automatically build a regression model on a dataset. They want to understand which features are most important for the model's predictions. Which feature of Autopilot can provide this insight?

Medium
21

A data scientist wants to track hyperparameters, metrics, and artifacts for multiple training runs in SageMaker. They need to compare runs and identify the best performing model. Which SageMaker feature should they use?

Medium
22

A company needs to perform time-series forecasting on historical sales data. Which SageMaker built-in algorithm is BEST suited for this task?

Easy
23

Which SageMaker feature allows you to automatically tune hyperparameters using Bayesian optimization?

Easy
24

A machine learning engineer is using SageMaker Automatic Model Tuning to optimize hyperparameters for a regression model. The objective metric is RMSE. The training job is costly, and the engineer wants to find a good configuration quickly. Which tuning strategy should they use?

Hard
25

A team is training a PyTorch model using SageMaker and wants to use their own custom training container with a specific PyTorch version. Which approach should they use?

Medium
26

A data scientist is using SageMaker to train a linear regression model on a dataset with a large number of features. They notice that the model's training time is long and want to speed it up by using a more efficient algorithm. They decide to use the SageMaker built-in Linear Learner algorithm. Which of the following is a key advantage of using the Linear Learner algorithm in SageMaker for this scenario?

Easy
27

A company wants to detect anomalies in login events from a large user base, focusing on unusual patterns that may indicate compromised accounts. Which SageMaker built-in algorithm is most suitable for this task?

Easy
28

A machine learning engineer runs a training job and notices the loss is NaN after a few steps. Which SageMaker Debugger rule can help identify this issue?

Medium
29

A machine learning engineer is training a deep learning model on SageMaker using the PyTorch estimator. The training job fails with an error indicating that the GPU memory is exhausted. The engineer wants to reduce memory usage without changing the model architecture. Which SageMaker feature should the engineer use?

Hard
30

A data science team is using SageMaker Experiments to track hyperparameters and metrics for a model training project. They need to compare multiple trials and identify the best model. Which THREE actions are part of a typical workflow? (Select THREE.)

Hard
31

A data scientist wants to train a binary classification model using Amazon SageMaker with a built-in algorithm that performs well on tabular data. Which algorithm should they choose?

Easy
32

A company is training a deep learning model using SageMaker and wants to reduce the time spent on data loading from Amazon S3 during training. The training dataset consists of many small files. Which approach is MOST effective to accelerate data loading?

Hard
33

A machine learning engineer is preparing a training job on SageMaker with a custom Docker container. Which TWO actions are required to use the container with SageMaker? (Choose TWO.)

Medium
34

A team wants to evaluate a binary classification model for credit risk. They need to understand the trade-off between false positives and false negatives. Which TWO metrics should they use? (Select TWO.)

Medium
35

A company wants to use SageMaker Autopilot to automatically build a binary classification model. Which output does Autopilot provide to help understand model decisions?

Medium
36

Which SageMaker built-in algorithm should be used for forecasting time series data with seasonal patterns?

Easy
37

A machine learning engineer wants to reduce training costs by using excess EC2 capacity. Which instance purchasing option should they choose for SageMaker training jobs?

Easy
38

A machine learning engineer is preparing a training dataset in Amazon SageMaker for a binary classification model. The dataset is stored as a single CSV file in Amazon S3 and contains 12 categorical features with high cardinality (thousands of unique values each). The engineer wants to avoid the curse of dimensionality and reduce training time while preserving predictive power. Which preprocessing approach should be used with the SageMaker built-in XGBoost algorithm?

Medium
39

A team wants to use a custom PyTorch training script in SageMaker. They need to install additional Python packages not included in the base PyTorch container. Which approach should they take?

Medium
40

A company uses SageMaker Clarify to detect bias in their training data. They find that the model has a high disparate impact for a protected attribute. What should they do to mitigate this bias during training?

Medium
41

A machine learning engineer is deploying a model to a SageMaker endpoint and wants to ensure that the model's predictions can be explained. The engineer needs to understand which features contributed most to each prediction. Which SageMaker feature should be used?

Hard
42

A machine learning engineer wants to automatically track hyperparameters, metrics, and artifacts for multiple training runs. Which SageMaker feature should they use?

Easy
43

A financial services company trains multiple models on SageMaker and needs to track hyperparameters, metrics, and artifacts for each experiment. Which SageMaker feature should they use to organize and compare experiments?

Medium
44

Which SageMaker built-in algorithm is specifically designed for time series forecasting?

Easy
45

A company is fine-tuning a large language model using LoRA on SageMaker. They want to reduce GPU memory usage during training. Which configuration change would help?

Hard
46

A data scientist trains a binary classification model using SageMaker and obtains an AUC of 0.95 on the test set. However, the precision-recall curve shows low precision for high recall thresholds. The business requires a model that performs well on the minority class. Which metric should the team primarily optimize during hyperparameter tuning?

Hard
47

A company needs to detect bias in a pre-trained model before deployment. They want to compute metrics like disparate impact and equal opportunity difference. Which AWS service should they use?

Hard
48

A team is training a large language model using PyTorch on SageMaker. They need to reduce training time. The model has 10 billion parameters. Which distributed training strategy should they use?

Medium
49

A company is training a deep learning model with SageMaker and wants to reduce training time by using pipe mode instead of file mode for a large dataset stored as TFRecord files in Amazon S3. After switching the estimator's input mode to Pipe, the training job fails immediately with a dataset format error. The data scientist confirms the files are valid TFRecords and that the same script works with File mode. What is the most likely cause?

Hard
50

A company wants to build a customer service chatbot that answers questions about their internal policy documents. The documents are updated monthly, and the team cannot afford to retrain a model each time. Which approach is MOST appropriate?

Medium
51

A team is building a fraud detection model using SageMaker and wants to detect anomalies in user login events. Which SageMaker built-in algorithm is specifically designed for anomaly detection in event-based data?

Medium
52

A company wants to use SageMaker to fine-tune a foundation model for a text generation task using RLHF (Reinforcement Learning from Human Feedback). Which THREE components are required in the RLHF pipeline?

Hard
53

A machine learning engineer has trained a model in SageMaker and wants to deploy it to a real-time endpoint for low-latency inference. The model artifacts are stored in Amazon S3, and the engineer needs to create the endpoint with the least operational effort. Which sequence of actions should the engineer take?

Easy
54

A data scientist needs to run a hyperparameter tuning job for a PyTorch model using SageMaker. They want to use Hyperband for efficient resource allocation. Which tuning strategy should they select in the HyperparameterTuner?

Medium
55

A company is using SageMaker Debugger to monitor a training job for a deep learning model. They want to detect when gradients become extremely large, which may cause training instability. Which built-in rule should they use?

Medium
56

A team is fine-tuning a foundation model using reinforcement learning from human feedback (RLHF) on SageMaker. They have a dataset of human preferences. Which SageMaker capability is most suitable for the reward model training step?

Hard
57

A team is fine-tuning a Hugging Face transformer model on SageMaker. They need to use a custom training script with the Hugging Face Estimator. Which SageMaker feature does this represent?

Medium
58

A data scientist suspects that a deep learning model is overfitting. They enable SageMaker Debugger and want to detect overfitting automatically. Which built-in rule should they use?

Medium
59

Which SageMaker built-in algorithm is best suited for detecting anomalous login attempts based on IP addresses and user behavior?

Easy
60

A machine learning engineer is preparing a dataset for training a SageMaker built-in Linear Learner model for binary classification. The dataset contains a highly imbalanced target with only 2% positive examples. They want to improve the model's ability to detect positives without collecting more data. Which SageMaker Linear Learner hyperparameter should they adjust to assign more weight to the positive class?

Hard
61

A data scientist is using SageMaker Experiments to track multiple training runs. They want to compare the F1 scores across runs. Which component should they use to log the F1 score?

Medium
62

A company is fine-tuning a large language model using reinforcement learning from human feedback (RLHF). Which THREE components are typically required?

Hard
63

A financial services firm is training a fraud detection model using SageMaker. The dataset is highly imbalanced (0.1% fraudulent transactions). The model currently achieves 99.9% accuracy but only catches 5% of fraud cases. Which metric should the team prioritize to evaluate model performance?

Hard
64

A company wants to use SageMaker Clarify to analyze bias in their training data and model predictions. Which TWO types of bias can Clarify detect? (Choose TWO.)

Easy
65

A data scientist is using SageMaker Automatic Model Tuning to optimize hyperparameters for an XGBoost model. They want to maximize AUC. Which search strategy is MOST appropriate for efficient exploration?

Medium
66

A company is training a large NLP model on SageMaker and wants to reduce costs by using Spot Instances. Which TWO configurations should they implement to handle Spot interruptions gracefully?

Medium
67

A company is using SageMaker Automatic Model Tuning to optimize a regression model. They want to minimize the root mean squared error (RMSE). The tuner has completed 20 jobs, and the RMSE has plateaued. Which action should the data scientist take to potentially improve the results?

Medium
68

A data scientist wants to fine-tune a Llama 2 7B model using SageMaker for a text summarization task. The dataset is 10 GB. The budget is limited, so cost efficiency is important. Which THREE steps should the data scientist take? (Choose THREE.)

Medium
69

A machine learning engineer is using SageMaker to train a model and wants to automatically stop a training job when the validation loss has not improved for 10 consecutive epochs, while still saving the best model artifacts. The engineer is using the SageMaker training toolkit in a custom container. Which combination of actions should the engineer take?

Medium
70

A data scientist is using SageMaker Experiments to track multiple training runs for a PyTorch model. They want to compare metrics across runs and identify the best hyperparameters. Which TWO capabilities should they use? (Choose TWO.)

Medium
71

A machine learning engineer is training a SageMaker job with the TensorFlow estimator and wants to automatically capture model training metadata such as loss curves and accuracy for later comparison, without writing any custom code. Which SageMaker feature should they enable?

Medium
72

A machine learning engineer trains a binary classifier in SageMaker and the model outputs class probabilities. The business requires that the model achieve at least 90% recall on the positive class, while keeping precision above 70%. The engineer uses the default threshold of 0.5 when deploying. Which approach should the engineer take to meet these requirements?

Medium
73

A data scientist is using SageMaker built-in XGBoost algorithm for a regression problem. Which metric is most appropriate as the objective metric for hyperparameter tuning?

Easy
74

A machine learning engineer is preparing a dataset for training a SageMaker model. The dataset contains missing values in several numerical features. The engineer wants to handle these missing values during the training pipeline. Which two methods are valid ways to handle missing values in SageMaker? (Choose two.)

Medium
75

A company is using SageMaker to train a model for image classification. The training dataset contains 100,000 labeled images. The team wants to use a pre-trained model to reduce training time. Which SageMaker feature should they use?

Medium
76

An ML engineer is fine-tuning a foundation model using RLHF on SageMaker. Which THREE components are essential for this workflow? (Select THREE.)

Hard
77

A team wants to fine-tune a pre-trained Hugging Face transformer model for text classification using SageMaker. They have a custom training script. Which SageMaker estimator should they use?

Easy
78

A machine learning engineer is preparing a training dataset stored in Amazon S3 for a SageMaker training job. The data is in CSV format, and the engineer wants to ensure that the training job can access the data efficiently and securely. The S3 bucket is in the same AWS Region as the SageMaker training job. The engineer needs to provide the training job with the necessary permissions to read the data. Which of the following is the MOST secure and appropriate way to grant the training job access to the S3 bucket?

Medium
79

A data scientist is using SageMaker built-in XGBoost algorithm for a binary classification task. Which objective metric is MOST appropriate for SageMaker Automatic Model Tuning to maximize?

Easy
80

A data scientist wants to track the training and validation accuracy of a SageMaker training job over time. They need to visualize these metrics in Amazon CloudWatch. Which action should they take?

Easy
81

A data scientist is using SageMaker Experiments to track multiple training runs. They want to compare runs based on the objective metric and visualize performance. Which THREE steps should they perform? (Choose THREE.)

Hard
82

A data scientist wants to quickly build a binary classification model without writing any code. Which SageMaker feature is MOST suitable?

Easy
83

A data scientist is training a model with SageMaker and needs to reduce the cost of a long-running training job that can tolerate interruptions. The job uses a custom training script and reads data from Amazon S3. The data scientist wants the job to resume from the last saved state if the underlying compute is reclaimed. (Choose two.)

Hard
84

A data scientist is training a model using SageMaker and wants to use spot instances to reduce costs. The training job is checkpointed every 5 minutes. However, the job gets interrupted frequently and never completes. What is the MOST likely cause?

Hard
85

A machine learning engineer is using SageMaker Autopilot for AutoML. Which TWO outputs does Autopilot produce?

Medium
86

A team is fine-tuning a Hugging Face BERT model for text classification using SageMaker. They want to use the Hugging Face estimator for convenience. Which parameter must be set to use a custom training script?

Medium
87

A machine learning engineer is training a TensorFlow model using SageMaker with distributed training. They need to implement data parallelism across multiple GPUs. Which SageMaker feature should they use to distribute the training?

Medium
88

A team is fine-tuning a large language model (LLM) using SageMaker and wants to reduce memory footprint during training. Which technique should they use?

Medium
89

A data scientist is using SageMaker Automatic Model Tuning to find the best hyperparameters for a model. They want to reduce the total tuning time for a given number of training jobs. Which tuning strategy should they choose?

Easy
90

A data scientist is using SageMaker Automatic Model Tuning with Hyperband. They want to stop poorly performing trials early to save resources. Which strategy does Hyperband use?

Medium
91

A machine learning team at a bank is training a binary classification model using SageMaker's built-in XGBoost algorithm on a dataset with 20 million rows and 300 features. They need to reduce training time while maintaining model accuracy. The data is stored in Amazon S3 as CSV files. Which approach should they take to speed up training?

Medium
92

A machine learning engineer is using SageMaker Debugger to monitor a training job and wants to detect issues early. The engineer wants to receive alerts when the training job is likely to fail due to vanishing gradients and when the loss is not decreasing. Which two actions should the engineer take to achieve this? (Choose two.)

Medium
93

A company wants to use SageMaker Autopilot for a regression problem. They require an explainability report that shows feature importance globally. Which Autopilot feature should they enable?

Medium
94

A machine learning engineer needs to reduce costs when training a large model on SageMaker. They are willing to accept potential interruptions and have checkpointing enabled. Which instance purchasing option should they use?

Easy
95

A machine learning engineer is evaluating a binary classification model that predicts customer churn. The model achieves 95% accuracy, but the engineer suspects class imbalance is causing a misleading metric. Which THREE evaluation steps should the engineer perform to properly assess the model? (Choose THREE.)

Hard
96

A data scientist is training an object detection model using SageMaker built-in Object Detection algorithm. They want to visualize the bounding boxes on validation images after training. Which approach should they use?

Medium
97

An ML engineer is fine-tuning a large language model using LoRA on SageMaker. The training is converging slowly, and GPU utilization is low. The engineer suspects the bottleneck is data loading. Which action should the engineer take to improve GPU utilization?

Hard
98

A team is training a large language model and needs to split the model layers across multiple GPUs due to memory constraints. Which distributed training strategy should they use?

Medium
99

A team is training a large deep learning model on SageMaker using a single ml.p3.16xlarge instance. Training is taking too long. They want to reduce time by distributing across multiple GPUs but are constrained by model size that does not fit in a single GPU memory. Which distributed training strategy should they use?

Medium
100

Which SageMaker feature provides AutoML capabilities, including automatic data preprocessing, model selection, and hyperparameter tuning?

Easy
101

Which SageMaker built-in algorithm is designed for time series forecasting?

Easy
102

A data scientist is evaluating a binary classification model. They have the confusion matrix and want to assess the model's performance comprehensively. Which THREE metrics should they consider? (Select THREE.)

Medium
103

A machine learning engineer is training a model using SageMaker and wants to set up monitoring to detect if gradients become too large, which could destabilize training. Which SageMaker Debugger built-in rule should they enable?

Medium
104

A data scientist needs to evaluate a binary classification model. The dataset is highly imbalanced (5% positive class). Which metric is MOST appropriate for assessing model performance?

Medium
105

A data scientist wants to use SageMaker Autopilot to automatically build a regression model. The dataset contains 200 features and 50,000 rows. Which output does SageMaker Autopilot provide?

Easy
106

A data scientist is using SageMaker to train an XGBoost model for a regression problem. After training, they evaluate the model on a test set and get an RMSE of 10 and an R² of 0.85. Which additional metric would give the MOST insight into the model's average prediction error magnitude?

Medium
107

A machine learning engineer is preparing a training script that must run on multiple GPU instances with SageMaker. The script currently reads the entire training dataset from local disk into memory, which fails on larger datasets. The engineer wants the script to stream training data from the SageMaker training channel path without loading everything into memory. Which approach should the engineer take?

Medium
108

A data scientist wants to fine-tune a large language model for a question-answering task. They want to reduce memory usage during training by using a low-rank approximation of the weight updates. Which technique should they use?

Medium

Frequently asked questions

What does the ML Model Development domain cover on the MLA-C01 exam?
Be able to pick the right SageMaker algorithm, metric, and job configuration for a described problem, then interpret results. The single most important thing: match the evaluation metric to the business objective, especially on imbalanced datasets.
How many questions are in this domain?
This page lists all 108 ML Model Development questions in the MLA-C01 question bank. The actual exam draws from this domain proportionally to its weighting in the official exam blueprint.
What is the best way to practise this domain?
Start with a short focused session (10 questions) to identify gaps, then work through explanations. Repeat with a longer session once the weak areas feel solid.
Can I practise only ML Model Development questions?
Yes — the session launcher on this page filters questions to this domain only. Choose any session length for inline explanations and scoring.
aws-ml-engineer-associate AWS-ML-ENGINEER-ASSOCIATE mla model development Practice Questions