Courseiva

AWS Certified Machine Learning Engineer Associate MLA-C01 (MLA-C01) — Questions 1–75

665 questions total · 9pages · All types, answers revealed

Page 1 of 9

Page 2
1
MCQmedium

A machine learning engineer is using Amazon SageMaker Feature Store to manage features for a fraud detection model. The engineer needs to ensure that the feature group can serve both batch and real-time predictions. The feature group is configured with an online store enabled. Which additional configuration is required to support batch predictions?

A.Increase the online store's read capacity to handle batch loads.
B.Set the feature group's event time to the ingestion time to allow batch queries.
C.Enable an offline store for the feature group and specify an S3 bucket for storage.
D.Configure the feature group to use a custom KMS key for encryption, which enables batch access.
AnswerC

To support batch predictions, the feature group must have an offline store, which stores historical feature data in Amazon S3. The online store is optimized for low-latency real-time serving, but batch predictions require access to historical data. Enabling the offline store and specifying an S3 bucket allows batch retrieval of features for training and batch inference.

Why this answer

Feature Store's online store is for real-time serving, while the offline store is for batch serving and training. To support batch predictions, the feature group must have an offline store enabled, which stores feature data in S3. This allows the engineer to retrieve historical features in bulk for batch inference jobs.

Without an offline store, only real-time serving is possible.

Exam trap

The trap here is assuming that the online store can handle batch predictions or that encryption or event time settings enable batch access, when in fact an offline store is required.

2
MCQhard

Refer to the exhibit. A data scientist used a SageMaker training job with a custom Scikit-learn script. The training job failed with the error shown. What is the most likely cause of this failure?

A.The training script is reading the CSV file incorrectly, causing a shape mismatch.
B.The InputDataConfig specifies ContentType text/csv but the actual file is not CSV.
C.The SageMaker training image is outdated and does not support Scikit-learn 1.0.
D.The training data contains missing values that need to be imputed.
AnswerA

The failure stems from the script reading the CSV without specifying that the first row contains column headers, so Scikit-learn treats header text as data and raises a value-conversion error. Supplying `header=0` (or `skiprows=1`) to pandas aligns the parsed shape with the model's expected feature count.

Why this answer

The error 'shape mismatch' typically occurs when the number of columns in the CSV file does not match the number of features expected by the Scikit-learn model. In SageMaker, the training script often loads data using pandas or numpy, and if the CSV has extra columns (e.g., an index column, header row misinterpreted, or trailing delimiter), the feature matrix shape will be inconsistent with the model's input dimensions, causing a ValueError during fit or transform.

Exam trap

In AWS SageMaker, a shape mismatch error often occurs when the CSV file contains extra columns (e.g., an index column from pandas during save) that increase the feature count beyond what the model expects. This is different from missing values or content type issues.

How to eliminate wrong answers

Option B is wrong because SageMaker's InputDataConfig ContentType is a metadata hint for the channel, not a validation mechanism; the training job will still attempt to read the file as CSV even if it is not, and the error would typically be a parsing error (e.g., pandas.errors.ParserError) rather than a shape mismatch. Option C is wrong because SageMaker training images are regularly updated and include compatible Scikit-learn versions; an outdated image would cause an import error or version mismatch, not a shape mismatch during data processing. Option D is wrong because missing values would cause a different error, such as a ValueError from NaN propagation or a fit failure, not a shape mismatch; imputation is a data preprocessing step, not a cause of dimension inconsistency.

3
Multi-Selecthard

A company is building a CI/CD pipeline for ML models using AWS CodePipeline and SageMaker. The pipeline should include steps to automatically retrain, evaluate, and deploy models. Which THREE components are essential for this pipeline? (Choose three.)

Select 3 answers
A.SageMaker Pipelines to orchestrate training and evaluation steps.
B.Amazon S3 bucket to store training data and model artifacts.
C.Amazon CloudWatch to log API calls.
D.SageMaker Model Registry to store and version models.
E.AWS Lambda function to trigger evaluation.
AnswersA, B, D

SageMaker Pipelines orchestrates the retrain, evaluate and conditional deploy steps as a directed acyclic graph, providing the workflow automation the CI/CD pipeline requires. It satisfies the essential-component requirement by coordinating training and evaluation before deployment within CodePipeline.

Why this answer

SageMaker Pipelines (A) is essential because it natively orchestrates the ML workflow steps—data preprocessing, training jobs, and evaluation processing jobs—and integrates directly with CodePipeline as the ML-specific orchestration layer. An Amazon S3 bucket (B) is required since SageMaker training jobs read training data from S3 and write model artifacts (model.tar.gz) back to S3, and CodePipeline itself uses S3 as its artifact store. SageMaker Model Registry (D) is essential to catalog, version, and track approval status of trained models, enabling controlled promotion from evaluation to deployment in the CI/CD flow.

Amazon CloudWatch (C) is not essential here because it provides monitoring and logging rather than being a required pipeline component for retrain/evaluate/deploy. AWS Lambda (E) is not essential because evaluation is handled by SageMaker Pipelines processing steps, and CodePipeline can invoke SageMaker actions directly without a custom Lambda trigger.

Exam trap

The trap here is that candidates often confuse monitoring services like CloudWatch with essential pipeline components, or assume that a serverless function like Lambda is required for evaluation when SageMaker Pipelines already provides native evaluation capabilities.

4
MCQhard

A team needs to deploy a model that has compliance requirements to log all inference requests and responses for auditing. The model will be served using a real-time endpoint. How can they achieve this without custom code?

A.Enable SageMaker Data Capture on the endpoint
B.Add a custom Lambda function using a container
C.Use SageMaker Debugger to monitor inference
D.Enable CloudTrail for the endpoint
AnswerA

SageMaker Data Capture on a real-time endpoint automatically logs inference requests and responses to Amazon S3, satisfying the audit requirement without custom code. It captures input and output payloads at the endpoint level, so compliance logging happens transparently for every invocation.

Why this answer

SageMaker Data Capture is the native, no-code feature that automatically logs inference requests and responses for real-time endpoints. It captures payload data to an S3 bucket without requiring any custom code, directly meeting the compliance requirement for audit logging.

Exam trap

The trap here is that candidates often confuse CloudTrail (which logs API calls) with Data Capture (which logs payloads), or they assume Debugger can be repurposed for inference logging, but Debugger only works during training.

How to eliminate wrong answers

Option B is wrong because adding a custom Lambda function using a container introduces custom code, which the question explicitly states should be avoided. Option C is wrong because SageMaker Debugger is designed for monitoring training jobs and debugging model performance, not for capturing inference request/response logs for auditing. Option D is wrong because AWS CloudTrail logs API calls to the SageMaker endpoint (e.g., InvokeEndpoint actions) but does not capture the actual inference request and response payloads.

5
MCQhard

A financial institution is training a fraud detection model using SageMaker. The dataset is highly imbalanced, with only 0.1% fraudulent transactions. The team wants to use SageMaker Automatic Model Tuning to find the best hyperparameters. They notice that the tuning job spends most of its time on configurations that predict all transactions as non-fraudulent. Which hyperparameter should they tune to directly address this issue?

A.max_depth
B.learning_rate
C.scale_pos_weight
D.subsample
AnswerC

In SageMaker's built-in XGBoost algorithm, scale_pos_weight controls the balance of positive and negative weights. Setting it to a higher value increases the weight of the positive class (fraudulent transactions), making the model pay more attention to them. This directly addresses the issue of the model predicting all transactions as non-fraudulent, as it penalizes misclassification of the minority class more heavily.

Why this answer

The scale_pos_weight hyperparameter in SageMaker's XGBoost algorithm adjusts the weight of the positive class, which is crucial for imbalanced datasets. By increasing this value, the model's loss function penalizes false negatives more, encouraging better detection of fraudulent transactions. This directly tackles the problem of the model predicting all instances as the majority class.

Exam trap

The trap here is focusing on general regularization or optimization hyperparameters, when the core issue is class imbalance that requires a weighting adjustment.

6
MCQhard

A machine learning engineer is using Amazon SageMaker Data Wrangler to prepare a dataset with a categorical feature that has over 5,000 distinct values (high cardinality). The engineer needs to transform this feature into a form suitable for a gradient boosting model while preserving as much information as possible. Which transform should be applied?

A.Label encoding
B.Target encoding with smoothing
C.One-hot encoding
D.Drop the feature
AnswerB

Target encoding with smoothing maps each of the 5,000 categories to a smoothed target statistic, producing one compact numeric column. Smoothing blends category means toward the global mean, preventing overfitting on rare categories and preserving information for the gradient boosting model.

Why this answer

Target encoding with smoothing is the best choice for high-cardinality categorical features in gradient boosting models. It replaces each category with a smoothed statistic (typically the mean of the target for that category), which preserves predictive information while avoiding the dimensionality explosion of one-hot encoding. Smoothing blends the category mean with the global mean to prevent overfitting on rare categories.

Exam trap

The trap is defaulting to one-hot encoding as the 'standard' categorical transform — candidates miss that with 5,000 categories it causes dimensionality explosion, and overlook target encoding's leakage risk if not done out-of-fold.

How to eliminate wrong answers

Option A is wrong because label encoding assigns arbitrary integers to categories, which gradient boosting trees may interpret as ordinal relationships that do not exist, harming model performance for high-cardinality features. Option C is wrong because one-hot encoding a 5,000-value feature creates 5,000 sparse binary columns, leading to extreme dimensionality, memory blowup, and poor tree split efficiency. Option D is wrong because dropping the feature discards potentially valuable predictive signal, which contradicts the requirement to preserve as much information as possible.

7
Multi-Selecthard

An MLOps team is designing a SageMaker Pipeline to automate model retraining. The pipeline must: (1) run training only if new training data is available, (2) register the model in SageMaker Model Registry only if evaluation metrics exceed a threshold, (3) deploy the approved model to a staging endpoint automatically. Which THREE steps should they include? (Choose THREE.)

Select 3 answers
A.ConditionStep to check if evaluation metrics exceed the threshold
B.RegisterModel step to register the model in the Model Registry
C.TuningStep to perform hyperparameter optimization
D.TransformStep to deploy the model to a staging endpoint
E.ProcessingStep to check for new training data availability
AnswersA, B, E

ConditionStep evaluates a property from the evaluation step against a threshold and branches accordingly, so RegisterModel executes only on the true branch. This enforces the stem's requirement that registration happens only when evaluation metrics exceed the threshold.

Why this answer

Option E (ProcessingStep to check for new training data availability) is correct because a ProcessingStep can run a script or job that inspects the input data source and outputs a boolean/property indicating whether new data exists, which is exactly what requirement (1) needs before training runs. Option A (ConditionStep to check if evaluation metrics exceed the threshold) is correct because ConditionStep is the pipeline construct that evaluates a condition (e.g., a metric from a previous evaluation step against a threshold) and gates downstream steps, satisfying requirement (2). Option B (RegisterModel step to register the model in the Model Registry) is correct because RegisterModel is the dedicated SageMaker Pipelines step that creates a model package in the Model Registry, which is required to register the model once the metric condition passes.

Option C (TuningStep) is not correct because hyperparameter optimization is not required by any of the stated requirements. Option D (TransformStep) is not correct because TransformStep runs batch transform inference jobs, not endpoint deployment; deploying to a staging endpoint would use a different mechanism such as a LambdaStep or a model deployment step, so it does not satisfy requirement (3).

Exam trap

MLA-C01 often tests the misconception that TransformStep deploys models to endpoints, when it actually runs batch transform jobs; endpoint deployment requires a separate deployment step or Lambda invocation.

8
MCQmedium

A company wants to run inference on a large dataset stored in S3 using a pre-trained model. The inference can tolerate latency from minutes to hours, and they want a fully managed solution that autoscales to handle large volumes. Which SageMaker inference option is most suitable?

A.Batch transform
B.Real-time endpoint
C.Asynchronous inference
D.Serverless inference
AnswerA

Batch transform runs asynchronous inference over an entire S3 dataset, writing predictions back to S3, so latency of minutes to hours is acceptable. It provisions and autoscales compute for the job, then terminates it, satisfying the fully managed requirement without a persistent endpoint.

Why this answer

Batch transform is the most suitable option because the company needs to run inference on a large dataset stored in S3 with latency tolerance from minutes to hours, and requires a fully managed, autoscaling solution. SageMaker Batch Transform processes the entire dataset as a single job, automatically provisions and scales compute resources, and writes results back to S3 without the need for persistent endpoints.

Exam trap

AWS often tests the distinction between latency tolerance and workload type, where candidates mistakenly choose asynchronous inference because they see 'tolerates latency' and 'autoscaling' without recognizing that batch transform is the only option designed for processing entire datasets stored in S3 as a single job, not individual requests.

How to eliminate wrong answers

Option B (Real-time endpoint) is wrong because it is designed for low-latency, synchronous inference (milliseconds to seconds) and requires a persistent endpoint that incurs ongoing costs, not suitable for large batch jobs with hour-long latency tolerance. Option C (Asynchronous inference) is wrong because it is intended for near-real-time requests with payloads up to 1 GB and queues requests for processing, but it still maintains a persistent endpoint and is optimized for workloads with latency in seconds to minutes, not hours-long batch processing of large datasets. Option D (Serverless inference) is wrong because it is designed for intermittent, short-lived inference requests with automatic scaling to zero, but it has a maximum timeout of 15 minutes per request and is not suitable for processing large datasets as a single job.

9
MCQeasy

A data scientist wants to compare the performance of two model versions (V1 and V2) in production by splitting traffic between them. They want to gradually increase the percentage of traffic to the new version while monitoring metrics. Which SageMaker feature enables this?

A.SageMaker production variants with traffic splitting
B.SageMaker shadow testing
C.SageMaker blue/green deployment
D.SageMaker canary deployment
AnswerA

SageMaker production variants let a single endpoint host multiple model versions with weighted traffic distribution, so the data scientist can shift percentages gradually while monitoring metrics. This satisfies the stem's requirement for controlled A/B comparison and progressive rollout between V1 and V2.

Why this answer

SageMaker production variants with traffic splitting let you host multiple model versions behind a single endpoint and assign a percentage of invocations to each variant. To compare V1 and V2 and gradually shift traffic, you create an endpoint configuration with two production variants and set the initial traffic distribution (e.g., 90/10), then update the weights as you gain confidence. This is the native SageMaker mechanism for A/B comparison and gradual rollout.

Exam trap

The trap is that 'canary deployment' sounds like a distinct SageMaker feature, but in SageMaker it is implemented via production variants and traffic splitting, so candidates who pick the canary-named option miss the actual configuration mechanism.

How to eliminate wrong answers

Option B is wrong because shadow testing mirrors production traffic to a new variant without returning its responses to callers, so it cannot be used to serve a percentage of real traffic to V2. Option C is wrong because blue/green deployment shifts all traffic from the old fleet to the new fleet at once (with a bake period), not a gradual percentage-based split. Option D is wrong because SageMaker does not expose a feature literally named 'canary deployment'; canary behavior is implemented using production variants and traffic splitting, so this option describes the outcome rather than the feature.

10
MCQmedium

A team is training a PyTorch model using SageMaker. They have a custom training script that requires specific Python packages not included in the SageMaker default PyTorch container. Which approach should they use?

A.Use the built-in PyTorch estimator and specify a requirements.txt in the source directory
B.Build a custom Docker container from scratch and push it to Amazon ECR
C.Use the SageMaker XGBoost estimator and modify the script to use PyTorch
D.Use SageMaker Autopilot to automatically handle dependencies
AnswerA

Specifying a requirements.txt in the source directory causes SageMaker to install those packages into the container before training begins. This satisfies the need for custom Python dependencies absent from the default PyTorch container, without building a custom image.

Why this answer

The SageMaker PyTorch estimator supports a requirements.txt file placed in the source_dir; the training toolkit automatically installs those packages into the container at the start of the job. This is the intended, low-effort way to add Python dependencies without rebuilding the image. It satisfies the requirement while keeping the managed PyTorch framework benefits (optimized libraries, GPU support, distributed training).

Exam trap

The trap here is assuming any dependency gap requires a custom container; MLA-C01 tests whether you know requirements.txt in source_dir is the supported lightweight mechanism for Python-only packages.

How to eliminate wrong answers

Option B is wrong because building a custom Docker container from scratch is unnecessary overhead when only Python packages are missing; it is the correct answer only when you need OS-level libraries, custom CUDA versions, or non-Python dependencies. Option C is wrong because the XGBoost estimator runs the XGBoost algorithm container, which cannot execute arbitrary PyTorch training code. Option D is wrong because SageMaker Autopilot is an AutoML service for tabular data that automates feature engineering and algorithm selection; it does not let you inject custom Python dependencies into a user-supplied PyTorch script.

11
MCQeasy

A data scientist needs to deploy a single ML model that will serve real-time predictions with low latency (under 10 ms) for a high-traffic web application. The model fits in memory and requires GPU acceleration. Which SageMaker inference option is MOST suitable?

A.Real-time endpoint on ml.m5 instances
B.Batch Transform
C.Real-time endpoint on ml.g4dn instances
D.Serverless Inference
AnswerC

A real-time endpoint on ml.g4dn instances provides GPU acceleration with persistent, low-latency inference, satisfying the sub-10 ms requirement for a high-traffic web application. Serverless inference lacks GPU support and cold starts, while batch transform cannot serve real-time requests.

Why this answer

Real-time endpoints on GPU instances (ml.g4dn) provide low latency and GPU acceleration, ideal for high-traffic, latency-sensitive workloads.

12
MCQhard

A machine learning team is preparing numerical features for a linear regression model. Feature 'A' ranges from 0 to 1000, feature 'B' ranges from 0 to 1, and feature 'C' ranges from -10000 to 10000. The team wants to ensure that feature scales do not affect the model's coefficients and that the features are bounded between 0 and 1. Which transformation should they apply?

A.RobustScaler (based on median and IQR)
B.Log transformation
C.StandardScaler (z-score normalization)
D.MinMaxScaler
AnswerD

MinMaxScaler applies x' = (x - min) / (max - min), mapping each feature linearly onto the [0, 1] range. This directly satisfies the bounded-scale constraint and prevents the wide-ranging features A and C from dominating coefficient magnitudes in the linear regression.

Why this answer

MinMaxScaler applies the formula (x - min) / (max - min), which linearly rescales each feature to the [0, 1] range regardless of its original bounds. This directly satisfies both stated requirements: eliminating scale-driven coefficient distortion in linear regression and bounding all features between 0 and 1. It is the only option that guarantees the [0, 1] output range.

Exam trap

MLA-C01 often tests the distinction between scalers by hiding the output-range requirement in the question; candidates who fixate on 'scale does not affect coefficients' alone may wrongly pick StandardScaler, missing that only MinMaxScaler guarantees the [0, 1] bound.

How to eliminate wrong answers

Option A is wrong because RobustScaler centers using the median and scales by the IQR, producing unbounded outputs that can be negative or exceed 1, so it does not satisfy the [0, 1] requirement. Option B is wrong because a log transformation compresses skew and changes the distribution shape but does not bound values to [0, 1] and is undefined for negative inputs like feature C (range -10000 to 10000). Option C is wrong because StandardScaler produces z-scores with mean 0 and standard deviation 1, which are unbounded and frequently negative, so features are not constrained to [0, 1].

13
MCQmedium

A company wants to implement a retraining pipeline that automatically triggers when SageMaker Model Monitor detects data drift. The retraining job should use the latest approved pipeline version in SageMaker Pipelines. Which approach meets these requirements?

A.Use a scheduled EventBridge rule to run the pipeline every day
B.Use SageMaker Model Monitor to update the model registry and trigger a deployment
C.Configure SageMaker Model Monitor to directly invoke a Lambda function on violation
D.Create an EventBridge rule that listens for SageMaker Model Monitor violation events and triggers a Lambda function that starts the pipeline
AnswerD

EventBridge natively consumes SageMaker Model Monitor violation events, and the rule's target Lambda starts the latest approved pipeline version through the SageMaker Pipelines API. This satisfies both constraints: automatic triggering on drift detection and use of the newest approved pipeline version.

Why this answer

It uses an EventBridge rule to listen for SageMaker Model Monitor violation events (e.g., `aws.sagemaker.model-monitoring-violation`), which then triggers a Lambda function that starts the latest approved pipeline version in SageMaker Pipelines. This creates an automated, event-driven retraining pipeline without manual intervention or scheduled polling.

Exam trap

The trap here is that candidates may think SageMaker Model Monitor can directly invoke Lambda or update the model registry, but in reality, it only emits events to EventBridge, and the integration requires an intermediate Lambda function to orchestrate the pipeline execution.

How to eliminate wrong answers

Option A is wrong because a scheduled EventBridge rule runs the pipeline daily regardless of whether data drift has occurred, leading to unnecessary retraining and resource waste. Option B is wrong because SageMaker Model Monitor does not directly update the model registry or trigger a deployment; it only publishes violation events and metrics. Option C is wrong because SageMaker Model Monitor cannot directly invoke a Lambda function; it emits events to EventBridge, which can then trigger Lambda, but the direct invocation is not supported.

14
MCQhard

A company uses AWS Glue ETL jobs to transform data for machine learning. They have a dataset with a column 'income' that is heavily right-skewed. Which transformation should be applied to make the distribution more Gaussian-like?

A.Log transformation (natural log)
B.Standardization (z-score)
C.Min-max scaling to [0,1]
D.Equal-width binning
AnswerA

A natural log transformation compresses the long right tail of a heavily skewed variable, pulling extreme high values closer to the bulk of the data. This reduces skewness and stabilises variance, making the 'income' distribution approximately Gaussian-like, which many ML algorithms assume.

Why this answer

A log transformation is appropriate for heavily right-skewed data because it compresses the long tail by applying a concave function, pulling extreme values closer to the mean and making the distribution more symmetric. In AWS Glue ETL, you can apply this using Spark SQL's `LOG` function or a Python UDF with `numpy.log`, which directly addresses the skewness to better approximate a Gaussian distribution for downstream ML models.

Exam trap

The trap here is that candidates confuse scaling (standardization or min-max) with shape-changing transformations, assuming any normalization makes data Gaussian, when in fact only non-linear transformations like log or Box-Cox address skewness.

How to eliminate wrong answers

Option B is wrong because standardization (z-score) centers and scales data to have mean 0 and standard deviation 1, but it does not change the shape of the distribution—it only rescales, so right-skewness remains. Option C is wrong because min-max scaling to [0,1] linearly compresses the data into a fixed range, which preserves the relative distances and does not alter skewness or make the distribution Gaussian-like. Option D is wrong because equal-width binning discretizes the continuous 'income' column into fixed intervals, which loses granularity and does not transform the distribution toward Gaussian—it creates a categorical or ordinal feature instead.

15
MCQmedium

A data scientist is using SageMaker Autopilot to automatically build a binary classification model on a balanced dataset. They want to understand the relationship between the input features and the model predictions. Which feature in SageMaker Autopilot should they use?

A.Explainability reports
B.Data visualizations
C.Model tuning results
D.Model candidate definitions
AnswerA

Explainability reports quantify each feature's contribution to predictions, satisfying the requirement to understand feature-prediction relationships. Autopilot generates these automatically for classification models, using SHAP values to attribute prediction outcomes to individual input features, revealing both global importance and per-instance effects without manual analysis.

Why this answer

SageMaker Autopilot's explainability reports provide feature importance and partial dependence plots (PDPs) that show how each input feature influences model predictions. For a binary classification model, these reports help data scientists understand the relationship between features and predictions. This is exactly what the question asks for.

Exam trap

MLA-C01 often tests the difference between data exploration features and model interpretability features — candidates may confuse 'data visualizations' (which describe the dataset) with 'explainability reports' (which describe the model).

How to eliminate wrong answers

Option B is wrong because data visualizations in Autopilot are for exploring the dataset, not for explaining model predictions. Option C is wrong because model tuning results show hyperparameter search outcomes, not feature-prediction relationships. Option D is wrong because model candidate definitions describe the algorithms and hyperparameters tried, not the interpretability of the final model.

16
MCQeasy

A company has a model that receives low traffic but needs to handle sudden spikes. Which deployment option is most cost-effective?

A.SageMaker Serverless Inference
B.SageMaker Real-Time Endpoint with Auto Scaling
C.SageMaker Multi-Model Endpoint
D.SageMaker Batch Transform
AnswerA

Serverless Inference scales to zero when idle and provisions capacity automatically during spikes, so the company pays only for actual usage. This suits low-traffic workloads with intermittent bursts, unlike always-on real-time endpoints that bill continuously.

Why this answer

SageMaker Serverless Inference is the most cost-effective option for low-traffic models with sudden spikes because it automatically scales to zero when not in use and scales up instantly to handle bursts, charging only for the compute time consumed per inference request. This eliminates the cost of idle provisioned infrastructure, making it ideal for unpredictable or intermittent traffic patterns.

Exam trap

AWS often tests the misconception that auto-scaling (Option B) is the most cost-effective for spikes, but the trap is that auto-scaling still requires a baseline of provisioned instances that incur cost even when idle, whereas serverless inference scales to zero and charges only for active compute time.

How to eliminate wrong answers

Option B (SageMaker Real-Time Endpoint with Auto Scaling) is wrong because it requires always-on provisioned instances, incurring costs even during idle periods, and auto-scaling has a lag that may not handle sudden spikes as quickly as serverless. Option C (SageMaker Multi-Model Endpoint) is wrong because it still uses provisioned instances that run continuously, and while it shares resources across models, it does not scale to zero or handle sudden spikes without pre-provisioned capacity. Option D (SageMaker Batch Transform) is wrong because it is designed for offline, asynchronous batch processing on a complete dataset, not for real-time inference with low latency or handling live traffic spikes.

17
MCQmedium

A data engineer needs to ingest streaming clickstream data from a website into an S3 data lake for ML training, with the ability to run real-time aggregations before storage. Which combination of AWS services meets these requirements?

A.Amazon Kinesis Data Streams → Kinesis Data Analytics → Kinesis Data Firehose → S3
B.Amazon Kinesis Data Streams → Amazon SageMaker Data Wrangler → S3
C.Amazon Kinesis Data Firehose → AWS Glue ETL → S3
D.Amazon SQS → AWS Lambda → S3
AnswerA

Kinesis Data Analytics performs the real-time aggregation the stem requires before storage, while Firehose handles reliable micro-batch delivery into S3. Data Streams ingests the clickstream continuously, satisfying both the streaming ingestion and pre-storage aggregation constraints in one pipeline.

Why this answer

It provides a complete pipeline for both real-time aggregation and durable storage. Kinesis Data Streams ingests the streaming clickstream data, Kinesis Data Analytics performs real-time SQL or Apache Flink-based aggregations on the stream, and Kinesis Data Firehose delivers the aggregated results to S3 with optional data transformation and buffering, meeting the requirements for ML training.

Exam trap

The trap here is that candidates may confuse Kinesis Data Firehose's ability to invoke Lambda for simple transformations with the need for real-time aggregations, overlooking that Kinesis Data Analytics is the only service that provides continuous, stateful stream processing required for real-time aggregations before storage.

How to eliminate wrong answers

Option B is wrong because Amazon SageMaker Data Wrangler is a tool for data preparation and feature engineering in batch mode, not for real-time streaming ingestion or aggregations; it cannot process a live Kinesis stream directly. Option C is wrong because Kinesis Data Firehose can ingest streaming data but lacks native real-time analytics capabilities—it can only buffer and deliver data, and AWS Glue ETL is a batch-oriented service, not suitable for real-time aggregations. Option D is wrong because Amazon SQS is a message queue for decoupling components, not designed for high-throughput streaming data, and AWS Lambda has a maximum execution time of 15 minutes and is not optimized for continuous real-time aggregations on clickstream data.

18
MCQhard

A machine learning team is training a large natural language processing model on Amazon SageMaker using the SageMaker Hugging Face container. The training job runs on multiple instances and uses Managed Spot Training to reduce costs. However, the job frequently gets interrupted by Spot interruptions, causing long training times. What should the team do to mitigate this issue?

A.Use a reserved capacity with Savings Plans
B.Use a larger instance type to finish faster
C.Enable checkpointing and increase the number of save intervals
D.Disable Managed Spot Training and use On-Demand instances
AnswerC

Checkpointing writes model state to Amazon S3 at each save interval, so when a Spot interruption reclaims an instance, the job resumes from the last checkpoint rather than restarting. Increasing save frequency reduces lost progress per interruption, directly mitigating the long training times caused by frequent Spot capacity reclamation.

Why this answer

Managed Spot Training on SageMaker can be interrupted when EC2 reclaims Spot capacity. Enabling checkpointing (saving model state to Amazon S3 at intervals) allows training to resume from the last checkpoint instead of restarting from scratch. Increasing the number of save intervals reduces the amount of lost work per interruption, directly mitigating long training times.

Exam trap

MLA-C01 often tests the trade-off between cost savings and reliability with Spot instances — candidates pick 'disable Spot and use On-Demand' as the safe answer, but the question asks to mitigate interruptions while presumably keeping cost benefits, making checkpointing the correct choice.

How to eliminate wrong answers

Option A is wrong because reserved capacity with Savings Plans applies to On-Demand instances and does not address Spot interruptions; it also increases cost, defeating the purpose. Option B is wrong because a larger instance type may finish faster per epoch but does not prevent Spot interruptions or reduce lost work when an interruption occurs. Option D is wrong because disabling Managed Spot Training and using On-Demand instances eliminates interruptions but significantly increases cost, which contradicts the team's goal of reducing costs; it is a valid fallback but not the best mitigation when checkpointing is available.

19
MCQeasy

A machine learning engineer has trained a model using SageMaker and wants to deploy it to a real-time endpoint. The engineer needs to specify the model artifacts, the inference code, and the environment. Which SageMaker resource should the engineer create first?

A.An endpoint configuration
B.A SageMaker pipeline
C.A SageMaker model
D.A SageMaker endpoint
AnswerC

A SageMaker model is the resource that encapsulates the model artifacts, inference code (as a Docker image), and environment variables. It is a prerequisite for creating an endpoint configuration and then an endpoint. Creating the model first is the correct initial step in the deployment process.

Why this answer

The deployment sequence in SageMaker begins with creating a model, which specifies the model artifacts and the inference container. This model is then referenced in an endpoint configuration, which defines the deployment settings. Finally, the endpoint is created from that configuration.

Therefore, the model is the first resource to create.

Exam trap

The trap here is thinking that an endpoint configuration or endpoint can be created without a model, but they both depend on the model resource.

20
MCQhard

A machine learning team is developing a deep learning model for image classification. They observe that the training loss decreases rapidly but the validation loss starts increasing after a few epochs. Which strategy should they implement to address this issue?

A.Increase the batch size
B.Add more convolutional layers
C.Increase the learning rate
D.Apply dropout regularization
AnswerD

Dropout randomly deactivates neurons during each training pass, which directly counteracts the overfitting causing validation loss to rise while training loss falls. By forcing the network to learn redundant, generalisable features rather than memorising the training images, it restores the validation-loss trajectory the stem requires.

Why this answer

The scenario describes overfitting, where the model memorizes training data but fails to generalize to validation data. Dropout regularization randomly deactivates a fraction of neurons during training, which prevents co-adaptation and forces the network to learn more robust features, thereby reducing the validation loss increase.

Exam trap

The AWS ML Engineer Associate exam often tests the distinction between underfitting and overfitting. The trap here is that candidates may confuse a rapidly decreasing training loss with successful learning, overlooking the validation loss divergence as the hallmark of overfitting that requires regularization rather than increased capacity or learning rate adjustments.

How to eliminate wrong answers

Option A is wrong because increasing the batch size typically stabilizes gradient estimates and can speed up training, but it does not directly address overfitting; in fact, larger batches may lead to sharper minima and worse generalization. Option B is wrong because adding more convolutional layers increases model capacity, which exacerbates overfitting by allowing the model to memorize more details from the training data. Option C is wrong because increasing the learning rate can cause the loss to oscillate or diverge, and it does not mitigate overfitting; it may even prevent convergence to a good solution.

21
MCQeasy

A data scientist notices that a SageMaker endpoint is returning HTTP 5XX errors under high load. The endpoint uses a single ml.m5.large instance. The team wants to reduce these errors without changing the instance type. What is the most cost-effective step?

A.Increase the endpoint's invocation timeout to 120 seconds
B.Deploy the model on a SageMaker batch transform job
C.Configure auto-scaling for the endpoint with a target tracking policy
D.Create a new endpoint with multiple instances and use weighted routing
AnswerC

Auto-scaling adds instances behind the endpoint when invocations rise, spreading load so no single ml.m5.large instance saturates and returns 5XX errors. A target tracking policy on a metric such as InvocationsPerInstance scales capacity automatically, satisfying the no-instance-type-change constraint while avoiding over-provisioning costs.

Why this answer

Configuring auto-scaling with a target tracking policy allows the endpoint to dynamically add more instances under high load, distributing the traffic and reducing HTTP 5XX errors. Since the team cannot change the instance type, scaling out is the most cost-effective way to handle increased demand, as it only adds capacity when needed and avoids over-provisioning.

Exam trap

The trap here is that candidates may think increasing the timeout (Option A) or using batch transform (Option B) can solve real-time load issues, but these options do not address the fundamental need for horizontal scaling under high concurrency.

How to eliminate wrong answers

Option A is wrong because increasing the invocation timeout to 120 seconds does not address the root cause of 5XX errors under high load; it merely extends the time the endpoint has to respond, which can lead to increased latency and potential timeouts, but does not prevent the endpoint from being overwhelmed. Option B is wrong because deploying the model on a SageMaker batch transform job is for offline, asynchronous inference on a static dataset, not for real-time serving; it cannot replace a real-time endpoint that needs to handle live traffic. Option D is wrong because creating a new endpoint with multiple instances and weighted routing adds cost by requiring manual management and does not automatically scale based on load; it is less cost-effective than auto-scaling, which adjusts capacity dynamically.

22
MCQhard

A company uses SageMaker Model Monitor for data quality. They notice that monitoring jobs are failing intermittently with constraint violations. Upon review, they see that some features have different data types in production compared to the baseline (e.g., string instead of integer). Which type of drift is this?

A.Schema drift
B.Concept drift
C.Statistical drift
D.Bias drift
AnswerA

Schema drift occurs when the structure or data types of incoming features diverge from the baseline, such as a string arriving where an integer was expected. This mismatch triggers the constraint violations the monitoring jobs report.

Why this answer

Schema drift occurs when the structure or data types of features in production data differ from the baseline used during model training. In this scenario, a feature that was an integer in the baseline is now a string in production, which is a classic example of schema drift. SageMaker Model Monitor detects this by comparing the inferred schema of production data against the baseline schema, flagging any type mismatches as constraint violations.

Exam trap

The trap here is that candidates may confuse schema drift with statistical drift, thinking any change in feature values qualifies as statistical drift, but the key differentiator is that schema drift specifically involves changes in data type or structure, not just distributional shifts.

How to eliminate wrong answers

Option B is wrong because concept drift refers to changes in the underlying relationship between features and the target variable, not changes in data types or schema. Option C is wrong because statistical drift (e.g., distribution shift) involves changes in the statistical properties of features (like mean or variance) while data types remain consistent. Option D is wrong because bias drift relates to changes in model fairness metrics over time, such as disparate impact, not to data type mismatches.

23
MCQmedium

A team is training a PyTorch model using SageMaker with a custom training script. They want to track hyperparameters and metrics across multiple experiments. Which service should they use?

A.SageMaker Clarify
B.SageMaker Experiments
C.SageMaker Model Monitor
D.SageMaker Debugger
AnswerB

SageMaker Experiments captures hyperparameters, metrics, and artefacts across training runs, letting the team compare and track multiple experiments from their custom PyTorch script. It satisfies the requirement to log and organise runs, which raw training jobs alone do not provide.

Why this answer

SageMaker Experiments is purpose-built for tracking ML experiment runs, automatically capturing hyperparameters, metrics, artifacts, and lineage across training jobs. It integrates natively with SageMaker training jobs and custom PyTorch scripts via the SageMaker SDK, letting teams compare runs in the Studio UI. This directly matches the requirement to track hyperparameters and metrics across multiple experiments.

Exam trap

MLA-C01 often tests the distinction between the four SageMaker 'specialty' services (Clarify, Experiments, Model Monitor, Debugger) — candidates confuse Debugger's training-time telemetry with Experiments' run-tracking purpose.

How to eliminate wrong answers

Option A is wrong because SageMaker Clarify is a bias detection and explainability tool (SHAP-based feature attribution), not an experiment tracking service. Option C is wrong because SageMaker Model Monitor detects data drift, model quality drift, and bias drift on deployed endpoints — it operates post-deployment, not during training. Option D is wrong because SageMaker Debugger captures tensors, gradients, and system resource utilization for debugging training convergence issues, not for organizing and comparing experiment runs.

24
Multi-Selectmedium

A machine learning engineer is using SageMaker Automatic Model Tuning (AMT) to optimize a model. They want to ensure the tuning job explores the hyperparameter search space efficiently and stops poorly performing trials early. Which two strategies should they use? (Choose two.)

Select 2 answers
A.Enable early stopping with the Hyperband strategy to terminate underperforming trials.
B.Define hyperparameters as categorical with a large number of discrete values to increase granularity.
C.Set the maximum number of training jobs to a very high value to ensure thorough exploration.
D.Use the Bayesian optimization strategy to model the objective function.
E.Use random search instead of Bayesian optimization to cover the search space uniformly.
AnswersA, D

Hyperband is a multi-fidelity optimization strategy that allocates resources to promising trials and stops those that perform poorly early. It is specifically designed to terminate underperforming trials, saving compute time and cost. Enabling early stopping with Hyperband directly addresses the requirement to stop poorly performing trials early.

Why this answer

Bayesian optimization efficiently models the objective function to select promising hyperparameters, while Hyperband early stopping terminates underperforming trials to save resources. Together, they maximize tuning efficiency. The other options either increase cost, use less efficient search, or do not address early stopping.

Exam trap

The trap here is thinking that more trials or random search will improve tuning efficiency, when in fact intelligent search and early stopping are the key strategies.

25
Multi-Selectmedium

A machine learning engineer is using SageMaker Pipelines to automate the training and deployment of a model. The pipeline includes a processing step for feature engineering, a training step, and a model registration step. The engineer wants to ensure that the pipeline is reproducible and that the model artifacts are versioned. Which two actions should be taken? (Choose two.)

Select 2 answers
A.Configure the training step to use the latest SageMaker training image without specifying a version.
B.Store the training dataset in Amazon S3 with versioning enabled and reference the specific version in the pipeline.
C.Register the trained model in the SageMaker Model Registry with a model package group and version.
D.Enable pipeline caching to reuse previous step outputs.
E.Use SageMaker Experiments to log the pipeline execution parameters and metrics.
AnswersB, C

Enabling S3 versioning and referencing a specific object version ensures that the exact dataset used for training is immutable and traceable. This supports reproducibility because rerunning the pipeline will use the same data version. Without versioning, the dataset could be overwritten, leading to inconsistent results. This action directly addresses the requirement for reproducibility and versioning of inputs.

Why this answer

Reproducibility requires pinning inputs, and versioning requires a registry for artifacts. Enabling S3 versioning and referencing a specific version ensures the training data is immutable and traceable. Registering the model in the SageMaker Model Registry creates a versioned entry that links to the artifacts and metadata.

Together, these actions provide a reproducible pipeline and a versioned model catalog.

Exam trap

The trap here is confusing tracking tools like SageMaker Experiments with actual versioning mechanisms, or assuming that caching provides reproducibility.

26
MCQeasy

A data scientist wants to track the lineage of models, datasets, and training jobs in SageMaker. Which SageMaker feature should they use to capture these relationships as artifacts and actions?

A.SageMaker Model Registry
B.SageMaker Experiments
C.SageMaker ML Lineage Tracking
D.SageMaker Feature Store
AnswerC

SageMaker ML Lineage Tracking automatically records relationships between datasets, training jobs and model artifacts as lineage entities, capturing both artifacts and actions. This directly satisfies the stem's requirement to track provenance across the machine learning workflow, which generic logging or experiment tracking alone would not provide.

Why this answer

SageMaker ML Lineage Tracking creates a graph of artifacts (datasets, models) and actions (training jobs, endpoints) to track the provenance of ML workflows.

27
MCQeasy

A data scientist needs to label a large dataset of product images for a classification model. They want to reduce labeling costs by prioritizing uncertain samples. Which Amazon SageMaker Ground Truth feature should they use?

A.Pre-built worker templates
B.Active learning
C.Consolidated labeling
D.Automated data labeling
AnswerB

Active learning selects the samples the model is least confident about and sends only those for human labelling, so annotation effort concentrates on uncertain images rather than the whole dataset. This directly satisfies the requirement to reduce labelling costs by prioritising uncertain samples.

Why this answer

Active learning in Amazon SageMaker Ground Truth automatically selects the most uncertain or informative samples from the unlabeled dataset and sends them to human annotators. This prioritization reduces labeling costs by focusing budget on the samples that will most improve model performance, rather than labeling all data indiscriminately.

Exam trap

The trap here is confusing 'active learning' with 'automated data labeling' — candidates often think automated labeling reduces costs by skipping humans entirely, but active learning specifically reduces costs by selectively using humans only on uncertain samples.

How to eliminate wrong answers

Option A is wrong because pre-built worker templates are UI frameworks for custom labeling workflows, not a mechanism to prioritize uncertain samples. Option C is wrong because consolidated labeling refers to merging multiple annotations for the same data point to produce a ground truth label, not selecting which samples to label. Option D is wrong because automated data labeling uses a trained model to label data without human review, which does not involve prioritizing uncertain samples for human annotation.

28
Multi-Selecteasy

A data science team wants to automate the retraining of a model when data drift is detected. Which TWO AWS services should they use in combination to achieve this? (Choose TWO)

Select 2 answers
A.AWS Cloud9
B.Amazon DynamoDB
C.SageMaker Model Monitor
D.AWS Lambda
E.Amazon Kinesis Data Analytics
AnswersC, D

SageMaker Model Monitor evaluates endpoint data against baselines and emits CloudWatch metrics when drift is detected, providing the detection half of the automation. Pairing it with an action service such as Lambda completes the retraining trigger.

Why this answer

SageMaker Model Monitor (C) is the correct service for detecting data drift: it continuously monitors a deployed model's endpoint, compares incoming inference data against a baseline, and can emit CloudWatch metrics/alerts when drift (or data quality, bias, or feature attribution issues) is detected. AWS Lambda (D) is the correct companion service because it can be triggered by those CloudWatch alarms/events to run the retraining logic — for example, invoking a SageMaker training job or pipeline to retrain and redeploy the model automatically. Together they form the detect-then-retrain automation loop the team needs.

AWS Cloud9 (A) is only a cloud IDE and provides no monitoring or automation capability. Amazon DynamoDB (B) is a NoSQL database and does not detect drift or orchestrate retraining. Amazon Kinesis Data Analytics (E) is for real-time stream processing with SQL/Flink, not for model drift detection or triggering retraining workflows.

29
Multi-Selectmedium

A team wants to deploy a new model using a canary deployment strategy on SageMaker. Which TWO configurations are necessary? (Choose two.)

Select 2 answers
A.Create a CloudWatch alarm to automatically rollback
B.Set the initial traffic distribution (e.g., 90% old, 10% new)
C.Enable data capture on the endpoint
D.Create a new endpoint configuration with two production variants, each pointing to a different model
E.Use SageMaker Model Registry to approve the new model
AnswersB, D

Setting the initial traffic distribution defines how much live inference traffic the new model variant receives from the outset. SageMaker's canary strategy requires this split to route a small percentage to the new variant while the remainder stays on the current one, satisfying the gradual, controlled rollout the stem demands.

Why this answer

Option B is correct because a SageMaker canary deployment is defined by specifying the initial traffic split between the existing (old) variant and the new variant, for example 90% to the old model and 10% to the new model, via the endpoint configuration's variant weights. Option D is correct because canary deployment requires a new endpoint configuration that contains two production variants, each referencing a different model, so traffic can be shifted between the old and new versions. Option A is not required, since CloudWatch alarms and automatic rollback are optional safeguards rather than mandatory canary configuration elements.

Option C is not required, as data capture is an optional monitoring feature and not part of the canary deployment definition. Option E is not required, because Model Registry approval is a governance step and not a necessary configuration for performing a canary deployment.

Exam trap

The trap is that several options are good practices (alarms, data capture, registry approval) but not necessary configurations, so candidates who conflate best practices with requirements select the wrong two.

30
MCQmedium

A company uses Amazon SageMaker Data Wrangler for data preparation. The data science team wants to automatically detect potential bias in their dataset before training a model. Which feature of Data Wrangler should they use?

A.Feature importance analysis
B.Amazon Rekognition
C.Model Monitor
D.Built-in bias detection with SageMaker Clarify
AnswerD

Data Wrangler's built-in bias detection invokes SageMaker Clarify to compute bias metrics such as class imbalance and disparate impact directly on the dataset before training. This satisfies the requirement to detect potential bias automatically during data preparation rather than post-training.

Why this answer

Amazon SageMaker Data Wrangler includes built-in bias detection powered by SageMaker Clarify. This feature allows data scientists to analyze datasets for potential bias before model training, directly within the Data Wrangler visual interface. Option D is correct because it is the only option that provides automated bias detection at the data preparation stage.

Exam trap

AWS often tests the distinction between pre-training bias detection (Clarify in Data Wrangler) and post-training monitoring (Model Monitor), causing candidates to confuse the two services.

How to eliminate wrong answers

Option A is wrong because feature importance analysis measures the contribution of each feature to a model's predictions, not bias detection in raw data. Option B is wrong because Amazon Rekognition is a computer vision service for image and video analysis, not a bias detection tool for tabular or structured data. Option C is wrong because Amazon SageMaker Model Monitor detects drift in deployed models' predictions and data quality over time, not bias in a static dataset before training.

31
MCQeasy

A company has 10 TB of log data in compressed JSON format stored in Amazon S3. The data needs to be processed and transformed into a structured format for machine learning. The processing requires complex transformations, including parsing nested JSON and joining with a reference table. The company wants to minimize infrastructure management. Which approach should the company use?

A.Use SageMaker Processing jobs to run custom scripts.
B.Use Amazon Athena to query and transform the data.
C.Use Amazon EMR with Apache Spark.
D.Use AWS Glue ETL with PySpark.
AnswerD

AWS Glue ETL with PySpark handles nested JSON parsing and reference-table joins at scale, while remaining serverless. This satisfies the 10 TB dataset and the requirement to minimise infrastructure management, since Glue provisions and scales the Spark environment automatically.

Why this answer

AWS Glue ETL with PySpark (Option D) is the best choice because it provides a fully serverless environment, minimizing infrastructure management. Glue can handle complex transformations like parsing nested JSON and joining with reference tables using PySpark, and it scales automatically for large datasets (10 TB). Amazon EMR (Option C) requires cluster management and provisioning, which contradicts the goal of minimizing management overhead.

Exam trap

Candidates may assume that large-scale data (10 TB) requires a provisioned cluster like EMR, but AWS Glue can scale to petabyte-scale workloads and is fully serverless, aligning with the goal of minimizing infrastructure management.

How to eliminate wrong answers

Option A is wrong because SageMaker Processing jobs are optimized for ML-specific tasks like training data preprocessing, not for general-purpose ETL on 10 TB of data; they lack native support for complex joins and nested JSON parsing at scale. Option B is wrong because Amazon Athena is a serverless query engine that excels at ad-hoc SQL queries but struggles with complex transformations like parsing deeply nested JSON and joining large reference tables due to its per-query pricing and lack of native procedural logic. Option D is wrong because AWS Glue ETL with PySpark is a valid alternative for ETL, but it is less performant and more expensive than EMR for large-scale (10 TB) data processing due to its auto-scaling overhead and limited tuning capabilities; EMR provides finer control over cluster configuration and cost optimization for batch jobs.

32
MCQeasy

A data scientist needs to split a time-series dataset into training and testing sets while avoiding data leakage from future values. Which splitting technique should the data scientist use?

A.Random shuffle followed by a 80/20 split
B.K-fold cross-validation with shuffled folds
C.Stratified sampling based on the target variable
D.Time-series split (walk-forward validation)
AnswerD

Time-series split (walk-forward validation) preserves chronological order, training only on past observations and testing on subsequent ones. This directly prevents future values leaking into training, satisfying the stem's no-leakage constraint. Random or stratified splitting would mix future data into training, invalidating the evaluation of a temporal model.

Why this answer

Time-series split (walk-forward validation) is the correct technique because it preserves the temporal order of observations, ensuring that training data always precedes test data. This prevents data leakage from future values, which is critical for time-series forecasting. Random shuffling or standard cross-validation would mix past and future data, leading to overly optimistic performance estimates.

Exam trap

MLA-C01 often tests the misconception that standard cross-validation or random splits are suitable for time-series data, but they cause data leakage; candidates must recognize that temporal order must be preserved.

How to eliminate wrong answers

Option A is wrong because random shuffle followed by an 80/20 split ignores temporal order, causing future data to leak into the training set and invalidating the model's predictive performance. Option B is wrong because K-fold cross-validation with shuffled folds also breaks temporal order, allowing future data to be used for training and leading to data leakage. Option C is wrong because stratified sampling based on the target variable is used for classification to maintain class distribution, but it does not respect time order and can still cause leakage in time-series data.

33
Multi-Selecteasy

A company is adopting Amazon SageMaker Pipelines to automate their ML workflow. They want to choose three key benefits that SageMaker Pipelines provides over traditional manual scripts and ad-hoc steps. Which THREE benefits are correct?

Select 3 answers
A.Model lineage tracking from raw data to trained model artifacts.
B.Automated deployment of models to endpoints upon pipeline completion.
C.Event-driven execution when new data arrives in S3.
D.Automatic scaling of compute resources based on data volume.
E.Reproducible execution through a directed acyclic graph (DAG) of steps with re-run capabilities.
AnswersA, C, E

SageMaker Pipelines records each step's inputs, outputs and parameters, producing automatic lineage from raw data through processing to trained model artefacts. This satisfies the stem's demand for benefits beyond manual scripts, where lineage must be reconstructed by hand and is rarely captured reliably.

Why this answer

SageMaker Pipelines provides model lineage tracking (Option A) because it automatically records the relationships between data, code, and model artifacts as they flow through pipeline steps, enabling end-to-end traceability from raw data to trained model artifacts. Option C is correct because Pipelines can be triggered by Amazon EventBridge events, such as when new data arrives in S3, enabling event-driven execution rather than manual invocation. Option E is correct because SageMaker Pipelines defines workflows as a directed acyclic graph (DAG) of steps, which makes executions reproducible and allows individual steps or entire pipelines to be re-run with consistent parameters.

Option B is not a guaranteed built-in benefit, since deployment to endpoints requires explicit steps such as a model registration or deployment step and is not automatic upon pipeline completion. Option D is also not a SageMaker Pipelines feature, as automatic scaling of compute resources based on data volume is handled by other services or configurations, not by the pipeline orchestration itself.

Exam trap

AWS often tests the distinction between orchestration features (like SageMaker Pipelines) and infrastructure management features (like auto-scaling), leading candidates to confuse pipeline benefits with SageMaker's broader managed service capabilities.

34
MCQmedium

A company is using SageMaker Model Registry to manage model versions. They want to automatically deploy the latest approved model to production after retraining. Which approach is best?

A.Manually deploy the approved model using the SageMaker console
B.Use AWS Lambda to update the endpoint whenever a new model version is created
C.Create a SageMaker Pipeline that includes a model approval step and deployment step
D.Schedule a CloudWatch Event to invoke a SageMaker update endpoint API daily
AnswerC

A SageMaker Pipeline can encode the approval step and subsequent deployment step as connected pipeline steps, so an approved model version automatically triggers production deployment after retraining. This automates the registry-to-endpoint promotion path the scenario requires.

Why this answer

A SageMaker Pipeline can orchestrate the entire workflow from retraining to deployment, including a model approval step that gates deployment to production only when the model is approved. This automates the process end-to-end, ensuring that only approved models are deployed, which aligns with the requirement to automatically deploy the latest approved model after retraining.

Exam trap

The trap here is that candidates may choose Option B because it sounds automated, but they overlook the critical requirement for model approval before deployment, which Lambda alone cannot enforce without additional logic.

How to eliminate wrong answers

Option A is wrong because manual deployment via the SageMaker console does not automate the process and violates the requirement for automatic deployment after retraining. Option B is wrong because using AWS Lambda to update the endpoint whenever a new model version is created would deploy models without waiting for approval, bypassing the model approval step and potentially deploying unapproved models. Option D is wrong because scheduling a CloudWatch Event to invoke a SageMaker update endpoint API daily does not tie deployment to model approval or retraining events; it deploys on a fixed schedule regardless of model status.

35
MCQmedium

A data scientist needs to ensure that the same train/test split is used across multiple experiments for reproducibility in SageMaker. Which approach should they take?

A.Use the same SageMaker instance type
B.Use the same hyperparameter values
C.Use the same dataset version
D.Set a random seed in the training script
AnswerD

Setting a fixed random seed makes the pseudo-random shuffling and splitting deterministic, so every experiment reproduces the identical train/test partition. This directly satisfies the reproducibility constraint across multiple SageMaker runs without altering the underlying data.

Why this answer

Setting a random seed in the training script ensures that the pseudo-random number generator used for splitting the dataset produces the same sequence of random indices across runs. This guarantees an identical train/test split regardless of instance type, hyperparameters, or dataset version, which is essential for reproducibility in SageMaker experiments.

Exam trap

The trap here is that candidates often confuse environmental consistency (instance type, dataset version) with algorithmic determinism, overlooking that reproducibility of data splits requires explicit control of the random seed in code.

How to eliminate wrong answers

Option A is wrong because the SageMaker instance type affects compute performance and memory, not the randomness of data splits; using the same instance type does not control the random seed. Option B is wrong because hyperparameter values influence model training behavior, not the deterministic splitting of data; they do not ensure the same train/test split. Option C is wrong because using the same dataset version ensures the data is identical, but without a fixed random seed, the split can still vary across runs due to different random number generator states.

36
MCQeasy

A company stores its model training data in Amazon S3. To meet compliance requirements, all data in transit between the S3 bucket and SageMaker must be encrypted. What should the company enforce?

A.Enable S3 versioning
B.Enable S3 access logging
C.Enforce HTTPS for all S3 access
D.Use S3 server-side encryption (SSE-S3)
AnswerC

Enforcing HTTPS via an S3 bucket policy with aws:SecureTransport denies any request not using TLS, so all traffic between S3 and SageMaker is encrypted in transit. This satisfies the compliance constraint, whereas SSE-KMS or SSE-S3 encrypt data at rest only.

Why this answer

Enforcing HTTPS for all S3 access ensures that data in transit between the S3 bucket and SageMaker is encrypted using TLS. This meets the compliance requirement for encrypting data in transit, as HTTPS uses TLS to protect data as it travels over the network.

Exam trap

The trap here is that candidates often confuse encryption at rest (SSE-S3) with encryption in transit (HTTPS/TLS), leading them to select Option D when the question explicitly asks about data in transit.

How to eliminate wrong answers

Option A is wrong because S3 versioning is a data protection feature that preserves, retrieves, and restores every version of an object stored in a bucket; it does not encrypt data in transit. Option B is wrong because S3 access logging provides detailed records of requests made to a bucket for auditing purposes, but it does not enforce or provide encryption for data in transit. Option D is wrong because S3 server-side encryption (SSE-S3) encrypts data at rest within S3, not data in transit between S3 and SageMaker.

37
MCQhard

A company is using Amazon SageMaker Data Wrangler to prepare a dataset with over 200 features. The dataset includes a categorical feature with more than 10,000 unique values (high cardinality). The ML engineer wants to transform this feature into a numeric representation suitable for a linear model without increasing dimensionality too much. Which built-in transform in Data Wrangler should the engineer use?

A.Ordinal encoding
B.Target encoding
C.One-hot encoding
D.Frequency encoding
AnswerB

Target encoding replaces each category with the mean of the target variable for that category, collapsing 10,000+ unique values into a single numeric column. This satisfies the stem's constraint of avoiding dimensionality growth, unlike one-hot encoding, while producing the numeric representation a linear model requires.

Why this answer

Target encoding is the correct choice because it replaces each category with the mean of the target variable for that category, producing a single numeric column that captures predictive signal without exploding dimensionality. This is ideal for high-cardinality features (e.g., >10,000 unique values) when used with linear models, as it avoids the sparsity and multicollinearity issues of one-hot encoding while retaining correlation with the target.

Exam trap

AWS often tests the misconception that high-cardinality categorical features must be one-hot encoded, but the trap here is that one-hot encoding explodes dimensionality, while target encoding provides a compact, target-informed numeric representation suitable for linear models.

How to eliminate wrong answers

Option A is wrong because ordinal encoding assigns arbitrary integer labels (e.g., 1, 2, 3) that imply an ordinal relationship, which can mislead linear models into assuming false ordering and degrade performance. Option C is wrong because one-hot encoding would create over 10,000 binary columns, drastically increasing dimensionality and causing memory/computation issues, defeating the goal of avoiding high dimensionality. Option D is wrong because frequency encoding replaces categories with their occurrence counts, which loses target-specific signal and can introduce bias if rare categories have low counts, making it less effective than target encoding for linear models.

38
MCQmedium

A data engineer is preparing a large training dataset stored in Amazon S3 as many small Parquet files, and a SageMaker training job that reads directly from S3 is spending most of its time on the input channel rather than on model computation. The engineer needs to improve the input throughput without changing the model code or the training algorithm. Which action should the engineer take?

A.Enable SageMaker Training Compiler and set the framework to accelerated mode.
B.Use the SageMaker File System Input with Amazon FSx for Lustre linked to the S3 bucket.
C.Increase the number of records per S3 GET request by enabling S3 Transfer Acceleration on the bucket.
D.Convert the output to TFRecord format and use Pipe mode with the SageMaker TensorFlow estimator.
AnswerB

FSx for Lustre linked to the S3 bucket exposes the dataset as a high-throughput, low-latency POSIX file system that SageMaker training jobs can mount, so the many small Parquet files are read far faster than repeated S3 GET requests. It improves input throughput without changing model code or the training algorithm, satisfying the requirement.

Why this answer

Feeding a training job from a high-performance shared file system removes the request-per-small-object overhead that dominates when many tiny Parquet files are read directly from S3. Amazon FSx for Lustre linked to the S3 bucket presents the data as files that the training container mounts and reads at high throughput, improving the input channel without altering model code or the training algorithm, which is exactly what the scenario requires.

Exam trap

The trap here is assuming that any SageMaker input-mode or acceleration feature will fix slow data loading, when the real issue is per-object overhead from many small files rather than raw bandwidth.

39
MCQhard

A company uses SageMaker Autopilot to build a binary classification model. The generated leaderboard shows an ensemble model as the best candidate. The team needs a model that can be deployed for real-time inference with latency < 10ms. What should they do?

A.Use SageMaker Inference Recommender to profile the ensemble model and optimize it
B.Deploy the ensemble model as a SageMaker endpoint; ensemble models are optimized for low latency
C.Retrain the ensemble model with fewer base estimators using a custom container
D.Select the best single model from the leaderboard (non-ensemble candidate) and deploy it
AnswerD

Ensemble models combine multiple learners, and their aggregated inference overhead typically exceeds the sub-10ms latency budget. A single non-ensemble candidate from the leaderboard has lower per-request compute, so selecting it satisfies the real-time latency constraint while retaining strong accuracy.

Why this answer

Ensemble models in SageMaker Autopilot combine multiple base models (e.g., stacking or voting), which increases inference latency because every base model must run and their outputs aggregated. For a strict <10ms real-time latency requirement, the best approach is to select the best-performing single (non-ensemble) model from the leaderboard, which has lower inference overhead. Autopilot's leaderboard explicitly lists both ensemble and individual model candidates, so the team can pick a single model that meets the latency SLA.

Exam trap

The trap is assuming the 'best' leaderboard model (highest accuracy) is always the right deployment choice — candidates forget that ensemble models trade latency for accuracy, and strict latency SLAs often force selection of a single model.

How to eliminate wrong answers

Option A is wrong because Inference Recommender profiles and recommends instance types for a given model — it does not reduce the inherent latency of an ensemble architecture enough to guarantee <10ms. Option B is wrong because ensemble models are not optimized for low latency; they are optimized for accuracy, and deploying one would likely violate the latency requirement. Option C is wrong because retraining with fewer base estimators requires a custom container and manual tuning, which is complex, not guaranteed to meet latency, and outside the standard Autopilot workflow.

40
MCQmedium

A machine learning engineer is training a tabular regression model using the SageMaker built-in XGBoost algorithm. They want to reduce overfitting and improve generalization without changing the algorithm. Which SageMaker hyperparameter should they tune to control the fraction of features randomly sampled per tree?

A.colsample_bytree
B.subsample
C.max_depth
D.eta
AnswerA

In the SageMaker built-in XGBoost algorithm, colsample_bytree specifies the subsample ratio of columns when constructing each tree. Lowering it introduces feature-level randomness, which reduces overfitting and often improves generalization on tabular regression tasks. It is a native XGBoost hyperparameter exposed by the SageMaker estimator, so tuning it directly addresses the scenario without changing the algorithm.

Why this answer

The SageMaker built-in XGBoost algorithm exposes colsample_bytree to control column subsampling per tree. Setting it below 1.0 introduces feature-level randomness that combats overfitting and can improve generalization on tabular data. Other hyperparameters such as subsample, eta, and max_depth affect different aspects of training and do not implement the requested feature-sampling behavior.

Exam trap

The trap here is confusing row subsampling (subsample) with column subsampling (colsample_bytree) when the scenario explicitly asks for feature sampling.

41
MCQmedium

A company deploys a real-time inference endpoint on SageMaker for a customer-facing application. Traffic patterns are unpredictable and sometimes spike. The endpoint must scale automatically to handle load while minimizing cost. Which approach should the company take?

A.Switch to batch transform for all inference requests.
B.Use a larger instance type to handle peak traffic.
C.Configure a target tracking scaling policy on the endpoint using Amazon CloudWatch metrics.
D.Deploy multiple models behind an Application Load Balancer.
AnswerC

Target tracking scaling adjusts instance count automatically against a CloudWatch metric such as InvocationsPerInstance, matching capacity to unpredictable spikes without manual intervention. It satisfies the stem's dual constraint: automatic scaling under variable load while minimising cost, since capacity shrinks during quiet periods rather than provisioning for peak.

Why this answer

SageMaker endpoints support automatic scaling through target tracking scaling policies based on Amazon CloudWatch metrics like InvocationsPerInstance. This allows the endpoint to dynamically adjust the number of instances in response to real-time traffic spikes, scaling out when demand increases and scaling in when it decreases, which optimizes cost by only paying for the capacity needed at any given time.

Exam trap

The trap here is that candidates often confuse scaling the SageMaker endpoint with scaling the instance size, thinking a larger instance (Option B) is the simplest solution, but the AWS exam tests understanding of dynamic, cost-optimized scaling using target tracking policies based on CloudWatch metrics rather than static over-provisioning.

How to eliminate wrong answers

Option A is wrong because batch transform is designed for offline, asynchronous inference on large datasets, not for real-time inference required by a customer-facing application; switching to batch would introduce unacceptable latency. Option B is wrong because using a larger instance type to handle peak traffic leads to over-provisioning and higher costs during normal or low traffic periods, as it does not scale dynamically with demand. Option D is wrong because deploying multiple models behind an Application Load Balancer (ALB) does not provide automatic scaling of the underlying compute resources; it only distributes traffic among existing instances, and SageMaker endpoints already have built-in load balancing without needing an external ALB.

42
MCQmedium

A machine learning team is building a fraud detection model. They have a dataset with a categorical feature 'merchant_id' that has over 10,000 unique values. Which feature engineering technique should they apply to 'merchant_id' to reduce dimensionality while retaining predictive power?

A.One-hot encoding
B.Target encoding
C.Ordinal encoding
D.Label encoding
AnswerB

Target encoding replaces each merchant_id with a statistic of the target computed from that category, collapsing 10,000 levels into one numeric column. This retains each merchant's fraud signal while eliminating the dimensionality one-hot encoding would create.

Why this answer

Target encoding replaces each merchant_id category with a statistic derived from the target variable (e.g., mean fraud rate per merchant), collapsing 10,000+ categories into a single numeric column while preserving the category's relationship to the label. This dramatically reduces dimensionality compared to one-hot encoding, which would create over 10,000 sparse columns. Because the encoded value carries predictive signal about fraud, target encoding retains useful information that label or ordinal encoding would lose.

Exam trap

MLA-C01 often tests the high-cardinality categorical scenario, and the trap is choosing one-hot encoding because it is the 'standard' approach, ignoring that 10,000+ categories make it computationally and statistically infeasible.

How to eliminate wrong answers

Option A (One-hot encoding) is wrong because it creates one binary column per unique merchant_id, exploding the feature space to 10,000+ sparse dimensions and causing the curse of dimensionality. Option C (Ordinal encoding) is wrong because assigning arbitrary integer ranks to merchant IDs implies a false ordering that has no business meaning and provides no target-related signal. Option D (Label encoding) is wrong because it also assigns arbitrary integers without any relationship to the target, so the model cannot exploit merchant-specific fraud patterns.

43
MCQmedium

A machine learning engineer is deploying a model to a SageMaker endpoint for real-time inference. The model must return predictions within 100 milliseconds for 95% of requests. The engineer wants to monitor the endpoint's latency and automatically roll back if latency exceeds the threshold. Which combination of SageMaker features should be used?

A.SageMaker Model Monitor with a custom monitoring schedule and AWS Lambda for rollback.
B.SageMaker deployment guardrails with a blue/green deployment and CloudWatch alarms for latency.
C.SageMaker Clarify for bias detection and AWS Step Functions for rollback orchestration.
D.SageMaker Inference Recommender to select the optimal instance type and automatic scaling.
AnswerB

SageMaker deployment guardrails support blue/green and linear deployments with automatic rollback triggered by CloudWatch alarms. You can create a CloudWatch alarm on the endpoint's model latency metric, and configure the guardrail to roll back if the alarm fires. This directly satisfies the need to monitor latency and automatically roll back when the threshold is breached.

Why this answer

SageMaker deployment guardrails are designed to safely update endpoints with automatic rollback. By associating a CloudWatch alarm that monitors model latency, the guardrail can trigger a rollback if the alarm state changes to ALARM. This provides the required automatic protection.

Other features like Model Monitor, Inference Recommender, or Clarify do not offer this latency-based rollback capability out of the box.

Exam trap

The trap here is assuming that SageMaker Model Monitor handles latency SLOs, but it focuses on data and model quality drift, not performance metrics like latency.

44
MCQeasy

A company wants to automate remediation when a SageMaker endpoint's latency exceeds a threshold for more than 5 minutes. The team needs to be notified and a Lambda function should be invoked to scale up the endpoint. Which combination of services should be used?

A.CloudWatch Alarm → SNS topic → Lambda function
B.EventBridge rule to trigger Lambda
C.CloudWatch Logs subscription filter → Lambda function
D.SageMaker Model Monitor → Lambda function
AnswerA

A CloudWatch alarm on the endpoint's latency metric detects the five-minute threshold breach, publishing to an SNS topic that both notifies subscribers and invokes the Lambda function to scale the endpoint, satisfying the combined notification and automated remediation requirement.

Why this answer

CloudWatch Alarms evaluate SageMaker endpoint metrics (e.g., ModelLatency) against thresholds over evaluation periods; when the alarm fires after 5 minutes of breach, it publishes to an SNS topic, which can fan out to both email/SMS subscribers for notification and a Lambda function for automated scaling. This is the canonical AWS pattern for metric-driven remediation.

Exam trap

MLA-C01 often tests whether candidates know that EventBridge handles event-driven triggers while CloudWatch Alarms handle metric-threshold triggers — mixing these up is the most common wrong-answer path.

How to eliminate wrong answers

Option B is wrong because EventBridge rules react to events/state changes, not to metric threshold breaches over time — EventBridge cannot natively evaluate 'latency > X for 5 minutes'. Option C is wrong because CloudWatch Logs subscription filters process log events, not CloudWatch metrics, and SageMaker endpoint latency is a metric, not a log stream. Option D is wrong because SageMaker Model Monitor detects data drift, bias, and quality issues in model predictions — it does not monitor endpoint latency or trigger scaling actions.

45
MCQhard

A company is using SageMaker to train a model with a custom container. The training script requires a specific version of a Python library that is not included in the default SageMaker containers. How should they provide this library?

A.Use SageMaker Script Mode and specify the library in a requirements.txt
B.Use SageMaker's lifecycle configuration to install the library on the training instance
C.Use pip install in the training script before model training
D.Extend a SageMaker framework container and install the library using a Dockerfile
AnswerD

Extending a SageMaker framework container with a Dockerfile allows bundling all required dependencies, including specific versions of Python libraries, into the Docker image that SageMaker will use for training. This ensures consistency and avoids runtime installation issues.

Why this answer

Using a custom container (BYOC) allows bundling all dependencies, including specific library versions, into a Docker image that SageMaker can run.

46
MCQhard

A data scientist is preparing a large dataset (50 GB) for training a TensorFlow model on SageMaker. The dataset consists of many small CSV files. Training is slow due to I/O bottlenecks. Which data preparation strategy most effectively accelerates training?

A.Convert the dataset to TFRecord format and use tf.data pipeline with prefetching
B.Convert the dataset to Parquet format and use Apache Arrow for loading
C.Compress the CSV files and decompress during data loading
D.Use a larger instance type with more vCPUs
AnswerA

TFRecord stores records in a compact binary format, eliminating per-file parsing overhead from thousands of small CSVs. The tf.data pipeline with prefetching overlaps data loading with GPU computation, directly resolving the I/O bottleneck constraining training throughput. This satisfies the scenario's requirement to accelerate training on a 50 GB dataset.

Why this answer

TFRecord format stores data in a binary, row-oriented format that TensorFlow's tf.data API can read efficiently, especially with prefetching to overlap data loading with model computation. This eliminates the per-file open/parse overhead of many small CSV files, which is the primary cause of I/O bottlenecks in this scenario.

Exam trap

The trap here is that candidates often choose larger instances (Option D) as a brute-force fix, failing to recognize that the root cause is the small-file I/O pattern, which requires a format change (TFRecord) rather than more compute resources.

How to eliminate wrong answers

Option B is wrong because Parquet is a columnar storage format optimized for analytical queries and selective column reads, not for sequential row-by-row training loops typical in deep learning; Apache Arrow adds overhead without solving the small-file problem. Option C is wrong because compressing CSV files reduces storage size but increases CPU load during decompression, often worsening I/O bottlenecks due to the many small files still requiring individual decompression. Option D is wrong because increasing vCPUs does not fix the fundamental I/O bottleneck caused by many small files; it may even exacerbate contention on shared storage without addressing the file access pattern.

47
Multi-Selectmedium

A data scientist is using SageMaker Autopilot for a regression problem. They want to see which data preprocessing steps Autopilot applied. Which TWO sources can they use to find this information?

Select 2 answers
A.Candidate definition notebook
B.Model leaderboard
C.Autopilot job description in AWS CloudTrail
D.Data exploration report
E.Explainability report
AnswersA, D

The candidate definition notebook documents the exact preprocessing pipeline Autopilot generated for that candidate, including transforms applied to features before training. It satisfies the requirement to inspect applied preprocessing by exposing the reproducible code rather than only summary statistics.

Why this answer

The candidate definition notebook (option A) is generated by SageMaker Autopilot for each candidate and contains the full ML pipeline code, including the exact data preprocessing and feature engineering transforms that were applied, so it directly answers the question. The data exploration report (option D) is produced during the Autopilot job's data exploration phase and documents the dataset's characteristics along with the preprocessing and feature engineering steps Autopilot selected, making it another valid source. The model leaderboard (option B) only ranks trained candidates by objective metric and does not describe preprocessing.

The Autopilot job description in AWS CloudTrail (option C) records API-level audit events, not the internal preprocessing steps. The explainability report (option E) covers feature attributions and model behavior, not the preprocessing pipeline.

48
MCQeasy

A data scientist wants to version and manage trained models, require approval before deployment, and enable cross-account deployment. Which SageMaker feature provides these capabilities?

A.SageMaker Neo
B.SageMaker Pipelines
C.Amazon Elastic Inference
D.SageMaker Model Registry
AnswerD

SageMaker Model Registry versions trained models through model package groups, enforces deployment approval via a manual approval status workflow, and supports cross-account deployment by sharing model packages with other AWS accounts. It directly satisfies all three stem requirements: versioning, approval gating, and cross-account deployment.

Why this answer

SageMaker Model Registry is the correct choice because it provides a centralized catalog for versioning trained models, supports approval workflows (e.g., pending, approved, rejected) to gate deployment, and enables cross-account deployment by sharing model package ARNs across AWS accounts via AWS Resource Access Manager (RAM) or cross-account IAM roles. This directly satisfies all three requirements: versioning, approval before deployment, and cross-account deployment.

Exam trap

The trap here is that candidates may confuse SageMaker Pipelines (which orchestrates the ML workflow) with Model Registry (which manages model versions and approvals), but Pipelines lacks native versioning and approval gatekeeping, while Model Registry is specifically designed for those governance tasks.

How to eliminate wrong answers

Option A is wrong because SageMaker Neo is a model optimization and compilation service that converts trained models into efficient runtime code for target hardware (e.g., ARM, Intel, NVIDIA), but it does not provide model versioning, approval workflows, or cross-account deployment capabilities. Option B is wrong because SageMaker Pipelines is a CI/CD orchestration service for building, training, and deploying ML pipelines, but it does not natively include a model registry with approval gates or cross-account deployment features; while it can integrate with Model Registry, Pipelines itself does not offer versioning or approval management. Option C is wrong because Amazon Elastic Inference is a service that attaches low-cost GPU-powered inference acceleration to SageMaker endpoints, but it has no role in model versioning, approval workflows, or cross-account deployment.

49
MCQhard

A company uses SageMaker Pipelines to orchestrate their ML workflow. They notice that if a pipeline step fails due to a transient error (e.g., a brief network issue), the entire pipeline fails and they must manually rerun from the beginning. They want to automatically retry failed steps a few times before failing. What should they do?

A.Use a Lambda function to catch step failures and re-invoke the step
B.Use the CreatePipelineExecution API with a flag to ignore failures
C.Configure a RetryPolicy in the pipeline step definition to specify the number of retry attempts and backoff
D.Use AWS Step Functions to orchestrate the workflow instead of SageMaker Pipelines
AnswerC

A RetryPolicy attached to the step definition lets SageMaker Pipelines automatically re-execute that step after a transient failure, using the configured max attempts and backoff interval, without restarting the whole pipeline. This directly satisfies the requirement to retry failed steps a few times before failing, removing manual reruns.

Why this answer

SageMaker Pipelines supports retry policies for steps. By setting a RetryPolicy in the step definition with a maximum number of retry attempts, the pipeline will automatically retry the step on failure. The other options do not achieve automatic retry within the pipeline: Step Functions would require rebuilding the pipeline, Lambda cannot retry pipeline steps, and CreatePipelineExecution does not handle retries.

50
MCQmedium

A data scientist is using SageMaker to train a deep learning model with the PyTorch estimator. They want to log custom scalar metrics such as validation accuracy and loss during training so they can monitor the job in SageMaker. Which approach should they use to emit these metrics from the training script?

A.Use the sagemaker_metrics API to write metrics to a local file and pass its path to the estimator's metric_definitions.
B.Print the metrics to stdout in a consistent format and define regex patterns in the estimator's metric_definitions parameter.
C.Call the SageMaker Metrics API directly from the training container to publish each metric.
D.Store metrics in Amazon CloudWatch Logs using the PutMetricData API and then reference them in the estimator.
AnswerB

SageMaker extracts training metrics by parsing the job's logs. When using the PyTorch estimator, the training script should print metric values to stdout in a predictable format, and the estimator's metric_definitions parameter provides regex patterns to capture them. This is the standard, supported method for custom scalar metrics and integrates with SageMaker monitoring and automatic model tuning.

Why this answer

For SageMaker training jobs, custom metrics are captured by parsing the job's logs. The PyTorch estimator supports metric_definitions, a list of name and regex pairs. Printing metrics to stdout in a consistent format lets SageMaker extract them for monitoring and tuning.

Direct API calls or file writes are not the supported mechanism for estimator metric capture.

Exam trap

The trap here is assuming there is a dedicated metrics API inside the container, when SageMaker actually parses stdout logs using regex patterns.

51
MCQmedium

A machine learning engineer is using a SageMaker training job with a custom training script. They need to save the trained model artifacts to Amazon S3 so that the model can be deployed later. Which parameter in the SageMaker estimator should they configure to specify the S3 location for model artifacts?

A.code_location
B.model_dir
C.output_path
D.dependencies
AnswerC

The output_path parameter in a SageMaker estimator specifies the S3 location where the training job stores model artifacts, such as model.tar.gz. It is the correct way to direct the trained model to a desired S3 bucket and prefix. Without setting it, artifacts go to a default SageMaker-managed bucket, which may not meet organizational requirements.

Why this answer

The output_path parameter in the SageMaker estimator defines the S3 URI where training job artifacts, including the final model, are stored. It allows control over the bucket and prefix, which is essential for governance and deployment. Other parameters like model_dir, code_location, and dependencies serve different purposes in the training lifecycle.

Exam trap

The trap here is confusing model_dir, which controls the local directory for saving the model inside the container, with output_path, which determines the final S3 location for artifacts.

52
MCQeasy

A machine learning engineer is training a model using SageMaker's built-in XGBoost algorithm. The training job fails with an error indicating insufficient memory. Which parameter should be adjusted to reduce memory usage?

A.subsample
B.num_round
C.max_depth
D.colsample_bytree
AnswerC

max_depth controls tree depth; deeper trees hold more split statistics and gradients in memory. Reducing it shrinks each tree's memory footprint, letting the SageMaker XGBoost training job fit within the instance's available memory and complete.

Why this answer

(max_depth) is correct because reducing the maximum depth of trees directly limits the number of splits and nodes per tree, which decreases the memory required to store the tree structure during training. In XGBoost, deeper trees exponentially increase the number of leaf nodes and intermediate splits, consuming more RAM for gradient statistics and tree data. Adjusting max_depth is the most direct way to reduce per-tree memory footprint without altering the dataset size or number of trees.

Exam trap

AWS exams often test the misconception that subsample or colsample_bytree are the primary knobs for memory reduction, when in fact max_depth has the most direct impact on per-tree memory consumption due to exponential node growth.

How to eliminate wrong answers

Option A is wrong because subsample controls the fraction of training data sampled per tree, which primarily reduces overfitting and training time, not memory usage per tree; it does not reduce the memory needed for the tree structure itself. Option B is wrong because num_round sets the number of boosting rounds (trees), which increases total training time and cumulative memory for model serialization, but does not reduce peak memory usage during a single round; reducing it may lower final model size but not the per-iteration memory spike. Option D is wrong because colsample_bytree controls the fraction of features used per tree, which can reduce memory for feature-related data structures but has a smaller impact on peak memory than tree depth, and is primarily a regularization technique.

53
MCQmedium

A team deploys a model with SageMaker and notices that the model returns inconsistent results during inference. They suspect a mismatch in feature transformation between the training pipeline and the inference pipeline. Which SageMaker feature can help compare the feature distributions?

A.Amazon SageMaker Model Monitor
B.Amazon SageMaker Autopilot
C.Amazon SageMaker Clarify
D.Amazon SageMaker Debugger
AnswerA

Amazon SageMaker Model Monitor captures baseline statistics from training data and continuously compares inference feature distributions against them, detecting drift and transformation mismatches. This directly satisfies the stem's need to compare training and inference feature distributions, surfacing the skew causing inconsistent results.

Why this answer

Amazon SageMaker Model Monitor is the correct choice because it continuously monitors the quality of deployed models by capturing inference data and comparing its distribution against the baseline training data distribution. When a mismatch in feature transformations occurs between training and inference pipelines, Model Monitor can detect data drift or feature attribution drift, alerting the team to the inconsistency. This allows them to identify and rectify the transformation discrepancy before it degrades model performance.

Exam trap

AWS often tests the distinction between monitoring (Model Monitor) and debugging (Debugger) — the trap here is that candidates confuse Debugger's training-time tensor analysis with the post-deployment data drift detection that Model Monitor provides.

How to eliminate wrong answers

Option B (Amazon SageMaker Autopilot) is wrong because it automates the process of building, training, and tuning machine learning models, but it does not provide ongoing monitoring or comparison of feature distributions between training and inference pipelines. Option C (Amazon SageMaker Clarify) is wrong because it focuses on bias detection and explainability of model predictions, not on detecting mismatches in feature transformations or data drift. Option D (Amazon SageMaker Debugger) is wrong because it is designed to debug training jobs by capturing tensors and metrics during training, not to compare inference data distributions against training baselines.

54
MCQhard

A financial services company is training a fraud detection model using SageMaker. The dataset is highly imbalanced, with only 0.2% fraudulent transactions. The team wants to optimize the model for recall at a fixed precision of 90%. They are using the SageMaker built-in XGBoost algorithm with binary:logistic objective. Which evaluation metric should they monitor during training and hyperparameter tuning?

A.F1 score
B.Precision-Recall AUC (PR-AUC)
C.Area Under the ROC Curve (AUC)
D.Matthews Correlation Coefficient (MCC)
AnswerB

PR-AUC summarizes the precision-recall curve, which is more informative than ROC for imbalanced datasets. It captures the trade-off between precision and recall, allowing the team to select a threshold that achieves 90% precision and then measure recall. By monitoring PR-AUC, they can compare models' ability to rank positive instances highly, which directly supports optimizing recall at a fixed precision.

Why this answer

For imbalanced classification with a precision constraint, the precision-recall curve is the right tool. PR-AUC summarizes performance across thresholds and emphasizes the minority class. The team can use the PR curve to find the threshold where precision is 90%, then read off recall.

This directly aligns with their goal of maximizing recall at that precision, unlike ROC-AUC or single-threshold metrics such as F1 or MCC.

Exam trap

The trap here is defaulting to ROC-AUC because it is common, but ROC-AUC can be overly optimistic on imbalanced data and does not enforce a precision constraint.

55
MCQeasy

A data engineer needs to prepare a large dataset for machine learning. The data is stored in an Amazon RDS MySQL database and needs to be transformed and moved to an S3 bucket in Parquet format for use with SageMaker. Which AWS service is most suitable for this extraction, transformation, and loading (ETL) task?

A.Use AWS Glue ETL jobs with PySpark to read from RDS, apply transformations, and write to S3 as Parquet.
B.Use Amazon Athena CTAS statements to copy data from RDS to S3.
C.Use SageMaker Data Wrangler to connect to RDS and export transformed data to S3.
D.Use Amazon EMR with Spark to read from RDS, transform, and write to S3.
AnswerA

AWS Glue ETL jobs with PySpark natively connect to Amazon RDS MySQL through JDBC, apply distributed transformations, and write Parquet directly to S3, satisfying the required format conversion. Glue's serverless Spark engine handles the large dataset scale without managing infrastructure, and its built-in SageMaker integration streamlines downstream machine learning consumption.

Why this answer

AWS Glue ETL jobs with PySpark are the most suitable service for this task because Glue is a fully managed, serverless ETL service that can natively connect to Amazon RDS MySQL via JDBC, apply transformations using PySpark, and write the output directly to S3 in Parquet format. This aligns perfectly with the requirement to extract, transform, and load a large dataset into a machine-learning-ready format without managing infrastructure.

Exam trap

The trap here is that candidates may confuse SageMaker Data Wrangler's ability to connect to RDS and export data with a full ETL capability, overlooking that it is an interactive tool for data preparation within SageMaker Studio rather than a serverless batch ETL service like AWS Glue.

How to eliminate wrong answers

Option B is wrong because Amazon Athena CTAS statements cannot read directly from Amazon RDS; Athena only queries data already in S3 or other data sources via federated queries, but CTAS itself requires the source to be in S3 or a cataloged table, not a live RDS database. Option C is wrong because SageMaker Data Wrangler is designed for interactive data preparation and feature engineering within SageMaker Studio, not for running serverless ETL jobs at scale; it can import data from RDS but lacks the native ability to schedule or run large-scale batch transformations and write to S3 as Parquet without additional infrastructure. Option D is wrong because while Amazon EMR with Spark can technically perform this task, it requires provisioning and managing a cluster, which adds operational overhead; AWS Glue is more suitable as a serverless, cost-effective alternative for this specific ETL workload without the need to manage EC2 instances or cluster lifecycle.

56
Multi-Selectmedium

A company is using SageMaker to serve a model for real-time predictions. They want to test a new model version by routing a small percentage of live traffic to it while the rest goes to the current model. They also need to compare performance metrics. Which TWO actions should they take? (Select TWO.)

Select 2 answers
A.Deploy the new model to a separate endpoint and use Route 53 to split traffic
B.Compile the new model with SageMaker Neo before deployment
C.Use SageMaker Batch Transform to evaluate the new model
D.Monitor the performance of both variants using SageMaker CloudWatch metrics
E.Configure a production variant with the new model and set initial traffic weight to a small percentage
AnswersD, E

SageMaker publishes per-variant invocation, latency and error metrics to CloudWatch, enabling direct comparison of the new model against the current one. This satisfies the requirement to compare performance metrics across both variants during the live traffic test.

Why this answer

Amazon CloudWatch provides built-in metrics for SageMaker endpoints, including latency, invocation counts, and error rates, which can be monitored per production variant. This allows the company to compare the performance of the new model version against the current model in real time. Option E is correct because SageMaker endpoints support multiple production variants, and you can set an initial traffic weight (e.g., 5%) to route a small percentage of live traffic to the new model while the rest goes to the existing variant.

Exam trap

A common misconception is that you need an external load balancer or DNS service (like Route 53) to split traffic between model versions, but SageMaker's built-in production variant feature handles this natively.

57
Multi-Selecteasy

A machine learning engineer wants to set up a retraining pipeline that triggers when model quality degrades. Which TWO components are essential for this automated retraining pipeline? (Select TWO)

Select 2 answers
A.CloudWatch Alarm on model quality metric
B.SNS topic to send notification to a Lambda function
C.SageMaker Ground Truth to collect new labels
D.SageMaker Data Wrangler to preprocess data
E.EventBridge rule to schedule retraining weekly
AnswersA, B

The alarm detects when model quality drops below a threshold.

58
MCQmedium

A company uses Amazon SageMaker Pipelines for automated retraining. The pipeline includes a processing step that runs a Python script. The script uses the boto3 library to call an AWS service, but the calls are being throttled. What is the MOST effective way to address this within the pipeline?

A.Increase the instance count for the processing step to distribute the API calls.
B.Modify the Python script to include retry with exponential backoff when receiving throttling exceptions.
C.Request a service quota increase for the throttling limit.
D.Add a wait step in the pipeline before the processing step.
AnswerB

Retry with exponential backoff directly addresses throttling by spacing repeated boto3 calls, letting the service recover between attempts. This satisfies the stem's constraint that throttling occurs inside the processing step's Python script, where the pipeline's own retry policy cannot intercept individual API calls. Increasing provisioned throughput or pipeline retries would not resolve per-call rate limiting.

Why this answer

Implementing retry with exponential backoff directly in the Python script is the most effective way to handle transient throttling exceptions from AWS service API calls. This approach is a best practice for managing service limits within a SageMaker Pipeline processing step, as it allows the script to automatically recover from throttling without modifying the pipeline structure or requiring manual intervention.

Exam trap

The trap here is that candidates often confuse infrastructure-level scaling (increasing instance count) with application-level retry logic, assuming more instances will reduce API call frequency, when in fact each instance independently makes the same number of calls and can still be throttled.

How to eliminate wrong answers

Option A is wrong because increasing the instance count for the processing step distributes the compute load, not the API calls made by the boto3 script; each instance still makes independent API calls to the same service endpoint, so throttling persists. Option C is wrong because requesting a service quota increase is a long-term, manual process that does not address immediate throttling within the pipeline and may not be necessary if the throttling is due to burst limits rather than hard quotas. Option D is wrong because adding a wait step before the processing step introduces a fixed delay that does not adapt to the actual throttling response; it cannot dynamically retry after a throttling exception occurs during the script execution.

59
MCQeasy

A machine learning engineer wants to monitor a deployed model for data drift. Which SageMaker feature should they use to automatically detect drift in the input data distribution compared to the training data baseline?

A.SageMaker Pipelines
B.SageMaker Model Monitor
C.SageMaker Debugger
D.SageMaker Clarify
AnswerB

SageMaker Model Monitor continuously evaluates endpoint input data against a training baseline, computing statistical distances to raise CloudWatch alerts when drift exceeds thresholds. This directly satisfies the requirement for automatic detection of input distribution shifts, unlike Model Registry or Clarify, which address governance and bias respectively.

Why this answer

SageMaker Model Monitor can be configured to run monitoring jobs that compare live inference data against a baseline created from training data to detect data drift.

60
MCQmedium

A machine learning engineer has trained a scikit-learn model and saved it as model.joblib in Amazon S3. The engineer wants SageMaker to host the model for real-time inference without writing a custom container or inference script, because the model uses only standard predict behavior. Which deployment approach should the engineer use?

A.Create a custom Docker image that installs scikit-learn, push it to Amazon ECR, and write an inference handler for the /invocations route.
B.Use SageMaker Batch Transform with the built-in Scikit-learn container and invoke the endpoint for each real-time request.
C.Deploy the model with the SageMaker Scikit-learn built-in framework container using the SageMaker Python SDK SKLearnModel class, pointing model_data to the S3 artifact.
D.Upload model.joblib to SageMaker Model Registry and rely on automatic endpoint creation when the model package is approved.
AnswerC

The Scikit-learn built-in framework container supports joblib and pickle artifacts and provides a default inference handler that loads the model and calls predict, so no custom script is needed. Supplying the S3 model artifact through SKLearnModel lets SageMaker extract the tarball and start a compliant real-time endpoint.

Why this answer

Because the model is a standard scikit-learn artifact and no custom inference logic is required, the built-in Scikit-learn framework container is the correct fit. Passing the S3 model artifact to the SKLearnModel class lets SageMaker load the joblib file and use the container's default predict handler, producing a real-time endpoint without custom code or image maintenance.

Exam trap

The trap here is assuming that any deployment requires a custom container, when SageMaker built-in framework containers already provide default inference handlers for supported libraries.

61
MCQhard

A team uses SageMaker Pipelines for CI/CD. The training step fails due to insufficient memory. How to fix without rewriting code?

A.Modify the training algorithm to use less memory
B.Reduce the batch size in the training script
C.Increase the instance type in the pipeline step configuration
D.Enable managed spot training
AnswerC

Increasing the instance type in the pipeline step configuration allocates more memory to the training container, resolving the out-of-memory failure. The training code and pipeline definition remain unchanged, satisfying the constraint of fixing the issue without rewriting code.

Why this answer

SageMaker Pipelines allows you to specify the instance type for each training step in the pipeline definition. By increasing the instance type (e.g., from ml.m5.large to ml.m5.xlarge or a memory-optimized instance like ml.r5.large), you allocate more memory to the training container without modifying the training script or algorithm. This directly resolves the out-of-memory error while preserving the existing code.

Exam trap

AWS often tests the distinction between infrastructure-level fixes (changing instance type in pipeline config) and code-level fixes (modifying script or algorithm), trapping candidates who think reducing batch size or enabling spot instances solves memory issues without considering the 'no code rewrite' constraint.

How to eliminate wrong answers

Option A is wrong because modifying the training algorithm to use less memory requires rewriting code, which violates the constraint of fixing the issue without rewriting code. Option B is wrong because reducing the batch size in the training script also requires modifying the training code, and while it may reduce memory usage, it does not meet the 'without rewriting code' condition. Option D is wrong because enabling managed spot training does not increase memory; it only reduces cost by using spare EC2 capacity and can cause interruptions, but it does not address insufficient memory for the training step.

62
Multi-Selecthard

A company is deploying a foundation model using SageMaker JumpStart. They want to minimize inference costs while maintaining low latency. Which TWO strategies should they consider? (Select TWO)

Select 2 answers
A.Enable data capture for all requests to analyze usage patterns
B.Use SageMaker Savings Plans for discounted compute rates
C.Deploy the model on a single large instance to maximize throughput
D.Enable auto-scaling with a target tracking policy based on Invocations per instance
E.Use SageMaker Inference Recommender to select the most cost-effective instance type
AnswersD, E

Auto-scaling adjusts capacity to match demand, avoiding over-provisioning.

Why this answer

Auto-scaling with a target tracking policy based on Invocations per instance dynamically adjusts the number of instances to match demand, ensuring you only pay for the compute capacity you need while maintaining low latency. This avoids over-provisioning and reduces idle costs, directly addressing the goal of minimizing inference costs.

Exam trap

The AWS exam often tests the misconception that cost minimization is achieved solely through discount plans (Savings Plans) or instance size, rather than through dynamic scaling and right-sizing based on actual workload patterns.

63
Multi-Selecthard

A company operates multiple AWS accounts with SageMaker workloads. They need to implement governance and security controls for model monitoring and maintenance. Which THREE actions should they take to meet compliance requirements?

Select 3 answers
A.Deploy a SageMaker model registry in a centralized account.
B.Use AWS CloudTrail to log all API calls to SageMaker and S3.
C.Enable VPC Flow Logs for SageMaker notebooks.
D.Use IAM roles with cross-account trust policies for all SageMaker endpoints.
E.Use AWS Config rules to enforce encryption of model artifacts.
AnswersA, B, E

A centralised SageMaker Model Registry in a dedicated account provides a single authoritative catalogue of model versions, approval statuses and metadata across all accounts. This satisfies the governance constraint by enabling consistent approval workflows and audit trails, while cross-account registry access lets teams publish and consume models without duplicating registries per account.

Why this answer

Option A is correct because a SageMaker Model Registry deployed in a centralized account provides a single governance point for cataloging, versioning, and approving models across multiple AWS accounts, which is essential for compliance and maintenance oversight. Option B is correct because AWS CloudTrail records all API activity, including calls to SageMaker and S3, giving the audit trail required to demonstrate who accessed or modified models and data. Option E is correct because AWS Config rules can continuously evaluate model artifacts in S3 for required encryption settings and flag or remediate noncompliant resources, directly enforcing a compliance control.

Option C is not required because VPC Flow Logs capture network traffic metadata for notebooks, which is useful for network troubleshooting but does not address model monitoring or maintenance governance. Option D is not required because cross-account IAM trust policies for endpoints are an access mechanism, not a governance or compliance control, and the question asks for monitoring and maintenance actions.

Exam trap

The trap here is that candidates often confuse VPC Flow Logs (network-level logging) with CloudTrail (API-level logging) or assume that cross-account IAM roles are a governance control rather than an access mechanism, leading them to select options that do not directly address compliance requirements for model monitoring and maintenance.

64
MCQmedium

A fraud-detection team runs a SageMaker real-time endpoint. Compliance requires that every inference request be logged with its full request and response payloads, and that a security engineer be able to prove later which requests were captured. The team enables SageMaker Model Monitor data capture with a capture percentage of 100. Where are the captured records stored, and what must be configured so the records are encrypted with a customer-managed key rather than an AWS-managed key?

A.Records are written to the S3 bucket in the DataCaptureConfig, and the customer-managed key is passed as the VolumeKmsKeyId property on the endpoint configuration.
B.Records are written to Amazon CloudWatch Logs under the endpoint's log group, and the customer-managed key is supplied through the KmsKeyId parameter of the DataCaptureConfig object.
C.Records are written to the S3 bucket specified in the DataCaptureConfig, and encryption with the customer-managed KMS key is applied by setting the KmsKeyId on that bucket's default encryption or by using an S3 bucket policy requiring the key.
D.Records are written to an Amazon Kinesis Data Firehose delivery stream created automatically by SageMaker, and the customer-managed key is set on the Firehose stream's SSE configuration.
AnswerC

Data capture writes JSON lines to the S3 destination given in DataCaptureConfig, so the storage location is the customer's own bucket. Because SageMaker delivers with S3 PutObject, the SSE-KMS setting comes from the bucket's default encryption or a policy that rejects uploads not using the customer-managed key. This gives the audit trail and key control the scenario demands.

Why this answer

Data capture stores request and response records as objects in the S3 bucket named in the endpoint's DataCaptureConfig, so key control is exercised at the S3 layer through default bucket encryption with a customer-managed KMS key or a bucket policy that denies uploads lacking that key. Endpoint volume keys and CloudWatch log groups do not govern those capture objects, making the S3-based approach the one that satisfies the audit and encryption requirement.

Exam trap

The trap here is assuming that a KMS key parameter on the endpoint configuration governs captured payloads, when payload encryption is actually controlled by the destination S3 bucket's encryption settings.

65
MCQeasy

A company uses an Amazon SageMaker endpoint for real-time inference. The security team requires that all traffic between the endpoint and the client application be encrypted in transit. Which configuration ensures this?

A.Deploy the endpoint in a VPC and use VPC Endpoints.
B.Use AWS Key Management Service (KMS) to encrypt the data in transit.
C.The endpoint is automatically served over HTTPS; no additional configuration is needed.
D.Attach an AWS Certificate Manager (ACM) certificate to the endpoint.
AnswerC

SageMaker real-time endpoints expose an HTTPS inference URL by default, with TLS terminating at the endpoint and enforced for every InvokeEndpoint call. Because the stem's constraint is encryption in transit between client and endpoint, this built-in TLS satisfies it without extra configuration, unlike custom inference code or VPC settings.

Why this answer

Amazon SageMaker endpoints are automatically served over HTTPS, which encrypts all data in transit between the client application and the endpoint. This is a default behavior of SageMaker real-time inference endpoints, so no additional configuration is required to meet the encryption-in-transit requirement.

Exam trap

The trap here is that candidates often overthink security requirements and assume additional configuration (like VPC endpoints or ACM certificates) is needed, when in fact SageMaker endpoints are inherently encrypted in transit via HTTPS by default.

How to eliminate wrong answers

Option A is wrong because deploying the endpoint in a VPC and using VPC Endpoints controls network traffic routing and provides private connectivity, but does not inherently enforce encryption in transit; the traffic could still be unencrypted if not using HTTPS. Option B is wrong because AWS KMS is used for encrypting data at rest (e.g., model artifacts, endpoint storage), not for encrypting data in transit; TLS/SSL handles in-transit encryption. Option D is wrong because attaching an ACM certificate to the endpoint is not a supported configuration for SageMaker endpoints; SageMaker automatically provisions and manages the TLS certificate for HTTPS, so manual certificate attachment is unnecessary and not possible.

66
Multi-Selectmedium

A company is deploying a machine learning model using SageMaker hosting. They need to support multiple versions of the model for A/B testing. Which TWO actions are required to set up the A/B test? (Choose two.)

Select 2 answers
A.Enable shadow variants to capture traffic for the new model without affecting users
B.Set up a batch transform job to compare performance offline
C.Configure the endpoint to route a percentage of traffic to each variant using initial variant weight
D.Register both models in SageMaker Model Registry
E.Create an endpoint with two production variants, each serving a different model version
AnswersC, E

SageMaker A/B testing requires traffic distribution across variants, configured through initial variant weight on the endpoint configuration. Setting weights routes a defined percentage of inference requests to each model version, satisfying the requirement to compare live performance between the two deployed versions.

Why this answer

Option E is correct because a SageMaker A/B test requires a single real-time endpoint that hosts two (or more) production variants, each pointing to a different model version, so live traffic can be split between them. Option C is correct because the traffic split is controlled by assigning each production variant an initial variant weight (a percentage of invocations), which the endpoint uses to route requests proportionally between the variants. Together, creating the multi-variant endpoint and configuring the initial variant weights are the required steps to run the A/B test.

Option A is not required because shadow variants only mirror traffic to a variant without returning its response to users, which is for testing without impacting users rather than A/B comparison. Option B is not required because a batch transform job performs offline inference and does not route live endpoint traffic. Option D is not required because registering models in the Model Registry is a governance/versioning practice, not a prerequisite for configuring endpoint traffic splitting.

Exam trap

The trap here is that candidates confuse shadow variants (which are for passive monitoring) with production variants (which are for active traffic splitting), leading them to select Option A instead of understanding that A/B testing requires explicit traffic routing via variant weights.

67
MCQeasy

A company is using Amazon SageMaker Ground Truth to create a labeled dataset for object detection in images. The team wants to minimize labeling costs while maintaining high accuracy. Which feature should they use to achieve this?

A.Use a private workforce of internal employees
B.Enable automated data labeling with active learning
C.Use a larger initial training set with pre-labeled public datasets
D.Reduce the number of label categories
AnswerB

Automated data labeling with active learning uses a model to label high-confidence objects and routes only uncertain images to human annotators, reducing the volume of manual labelling required while sustaining accuracy, thereby lowering overall Ground Truth labelling costs.

Why this answer

Amazon SageMaker Ground Truth's automated data labeling uses active learning to train a model on the labels already produced by humans, then automatically labels the high-confidence data and only sends low-confidence samples to human labelers. This dramatically reduces the number of human-labeled items required, cutting costs while preserving accuracy because humans still validate uncertain cases. This is the specific cost-optimization feature designed for exactly this scenario.

Exam trap

MLA-C01 often tests the misconception that cost reduction in labeling comes from using more pre-labeled public data or fewer categories, when the intended answer is the active-learning-based automated labeling feature that specifically minimizes human labeling volume.

How to eliminate wrong answers

Option A is wrong because a private workforce of internal employees still requires paying for every single label and does not reduce the volume of human labeling work — it only changes who does it. Option C is wrong because pre-labeled public datasets rarely match the specific object classes and image domain of a proprietary dataset, and using them does not reduce the cost of labeling the company's own images. Option D is wrong because reducing label categories changes the task definition and may hurt model accuracy; it does not address the cost of labeling effort per image.

68
MCQeasy

A data scientist is building a model to predict customer churn based on historical data. The dataset has 10 features and 100,000 records, and the target is binary. Which algorithm is most appropriate for this binary classification problem?

A.Principal component analysis
B.K-means clustering
C.Linear regression
D.Logistic regression
AnswerD

Logistic regression outputs a probability between 0 and 1 via the sigmoid function, making it suited to the binary churn target. With 100,000 records and only 10 features, it trains quickly, resists overfitting, and yields interpretable coefficients, satisfying the classification constraint.

Why this answer

Logistic regression is the most appropriate algorithm for this binary classification problem because it directly models the probability of the binary target variable using a logistic (sigmoid) function, making it a natural fit for predicting customer churn (yes/no). It is efficient with 100,000 records and 10 features, providing interpretable coefficients that indicate feature importance, which is crucial for understanding churn drivers.

Exam trap

AWS often tests the distinction between supervised and unsupervised learning, and the trap here is that candidates may confuse logistic regression with linear regression due to the similar name, or incorrectly choose PCA or K-means because they are familiar with them for feature reduction or segmentation, ignoring that the task is explicitly binary classification.

How to eliminate wrong answers

Option A is wrong because Principal Component Analysis (PCA) is an unsupervised dimensionality reduction technique, not a classification algorithm, and it cannot predict a binary target. Option B is wrong because K-means clustering is an unsupervised learning method used for grouping unlabeled data, not for predicting a known binary outcome. Option C is wrong because linear regression predicts a continuous numeric value, not a binary class, and applying it to a binary target would violate the assumption of normally distributed errors and produce probabilities outside [0,1].

69
Multi-Selectmedium

A data scientist wants to use SageMaker Clarify to analyze bias during training of a binary classification model. Which TWO types of bias metrics can SageMaker Clarify compute? (Select TWO.)

Select 2 answers
A.Feature importance
B.Post-training bias metrics (e.g., Difference in Positive Proportions, AD)
C.SHAP values
D.Pre-training bias metrics (e.g., Class Imbalance, DPL)
E.Confusion matrix
AnswersB, D

Post-training bias metrics are computed from model predictions on a dataset, comparing predicted labels across facets; measures such as Difference in Positive Proportions in Predicted Labels and Accuracy Difference quantify bias introduced by the trained model itself.

Why this answer

SageMaker Clarify computes bias metrics in two distinct phases of the ML lifecycle, and both are valid answers here. Option B is correct because post-training bias metrics evaluate the model's predictions after training, using measures such as Difference in Positive Proportions in Predicted Labels (DPPL), Accuracy Difference (AD), and Disparate Impact, which quantify whether outcomes differ across groups. Option D is correct because pre-training bias metrics analyze the training data before any model is built, using measures such as Class Imbalance (CI), Difference in Proportions of Labels (DPL), and Label Imbalance to detect skew in the dataset itself.

Option A is not a bias metric category — feature importance (e.g., via SHAP) explains which features drive predictions, not whether outcomes are biased. Option C is likewise an explainability technique, not a bias metric, and Clarify reports SHAP values under its explainability analysis. Option E is a standard model evaluation artifact for classification performance, not a bias metric computed by Clarify.

Exam trap

MLA-C01 often tests the confusion between bias metrics and explainability metrics, leading candidates to select feature importance or SHAP values as bias metrics.

70
MCQhard

A healthcare company is deploying a model for predicting patient outcomes. The model must be deployed across multiple AWS accounts to meet compliance requirements. Each account has its own Amazon SageMaker endpoint. The company wants to centralize monitoring of model performance without exposing data across accounts. Which solution should the company use?

A.Establish VPC peering between accounts and call the endpoints from a central monitoring service.
B.Replicate the inference data to a central S3 bucket in the management account using cross-account replication, then run Model Monitor centrally.
C.Use SageMaker Model Monitor in each account and publish custom metrics to a central CloudWatch account using cross-account observability.
D.Create a shared SageMaker Model Registry across accounts and aggregate monitoring.
AnswerC

Cross-account observability lets each account's Model Monitor publish metrics into a central CloudWatch account, so performance is aggregated without moving patient data between accounts. This satisfies the compliance constraint requiring per-account endpoints while centralising monitoring.

Why this answer

It uses SageMaker Model Monitor in each account to detect data drift and model degradation locally, then publishes custom metrics to a central CloudWatch account via cross-account observability. This approach centralizes monitoring without moving raw inference data across accounts, satisfying the compliance requirement of not exposing data.

Exam trap

The trap here is confusing data replication (which exposes raw data) with metric aggregation (which exposes only statistical summaries), leading candidates to pick Option B despite its compliance violation.

How to eliminate wrong answers

Option A is wrong because VPC peering enables network connectivity but does not provide a mechanism to centralize monitoring metrics or avoid exposing inference data across accounts; it would require direct data transfer to a central service, violating compliance. Option B is wrong because replicating inference data to a central S3 bucket using cross-account replication exposes raw data across accounts, which directly violates the requirement to not expose data. Option D is wrong because a shared SageMaker Model Registry aggregates model metadata and versions, not real-time monitoring metrics or data drift detection; it does not provide centralized performance monitoring.

71
MCQhard

A machine learning engineer is setting up automated retraining for a model using SageMaker Pipelines. The pipeline should trigger when a data drift alert is received from Model Monitor. Which event source should the engineer use to initiate the pipeline?

A.Amazon CloudWatch Events (Amazon EventBridge) rule that captures Model Monitor outcome.
B.AWS Lambda function that polls CloudWatch logs.
C.S3 event notification on the monitoring output bucket.
D.SageMaker model monitor webhook.
AnswerA

Model Monitor publishes drift outcomes to Amazon CloudWatch, which EventBridge captures as events. An EventBridge rule matching those Model Monitor outcome events can invoke the SageMaker Pipeline, satisfying the requirement to trigger retraining automatically when drift is detected.

Why this answer

Amazon EventBridge (formerly CloudWatch Events) is the native AWS service for reacting to state changes in AWS resources. SageMaker Model Monitor publishes data drift alerts as events to EventBridge, so a rule can be configured to match those specific events and trigger the SageMaker Pipeline execution as a target. This provides a fully managed, event-driven architecture without polling or custom integrations.

Exam trap

The trap here is that candidates confuse S3 event notifications (which are for object-level events) with the structured, high-level alerts emitted by Model Monitor, leading them to choose Option C instead of the correct EventBridge integration.

How to eliminate wrong answers

Option B is wrong because polling CloudWatch Logs with a Lambda function introduces unnecessary latency, complexity, and cost; AWS best practice is to use EventBridge for event-driven triggers rather than polling. Option C is wrong because S3 event notifications on the monitoring output bucket would fire on every object write, not specifically on a data drift alert, leading to false triggers and wasted compute. Option D is wrong because SageMaker Model Monitor does not expose a webhook; it integrates with EventBridge for event delivery, not with external HTTP callbacks.

72
MCQmedium

An ML engineer needs to orchestrate a multi-step workflow that includes data preprocessing on Spark, model training on SageMaker, and deployment to a production endpoint. They require tight integration with other AWS services and the ability to add custom logic. Which AWS service should they use alongside SageMaker?

A.AWS Step Functions
B.AWS CloudFormation
C.SageMaker Pipelines
D.Amazon EventBridge
AnswerA

AWS Step Functions orchestrates the workflow as a state machine, invoking Spark preprocessing, SageMaker training jobs and endpoint deployment as discrete steps. Its native AWS service integrations plus Lambda and Activity tasks satisfy the requirement for custom logic, while retries, error handling and visual tracking coordinate the multi-step pipeline that SageMaker alone cannot sequence.

Why this answer

AWS Step Functions is the correct choice because it provides a serverless workflow orchestration service that can coordinate multi-step ML pipelines involving Spark on AWS Glue or EMR, SageMaker training jobs, and endpoint deployments. It offers tight integration with over 200 AWS services via direct SDK integrations, supports custom logic through Lambda functions, and includes built-in error handling, retries, and parallel execution — making it ideal for complex, heterogeneous ML workflows that extend beyond SageMaker's native capabilities.

Exam trap

The trap here is that candidates confuse SageMaker Pipelines (a SageMaker-native orchestrator) with a general-purpose orchestrator, overlooking the requirement for tight integration with non-SageMaker services like Spark and custom logic — Step Functions is the correct choice for heterogeneous, multi-service ML workflows.

How to eliminate wrong answers

Option B (AWS CloudFormation) is wrong because it is an Infrastructure as Code (IaC) service for provisioning and managing AWS resources declaratively, not a workflow orchestrator — it cannot sequence steps like 'run Spark job, then train model, then deploy endpoint' with conditional logic or dynamic state management. Option C (SageMaker Pipelines) is wrong because while it can orchestrate SageMaker-native steps (training, tuning, batch transform), it lacks direct integration with external services like Spark on EMR or Glue and cannot easily incorporate custom logic outside the SageMaker ecosystem — the question explicitly requires tight integration with other AWS services and custom logic beyond SageMaker. Option D (Amazon EventBridge) is wrong because it is an event bus service for routing events between services based on rules, not a workflow orchestrator — it cannot manage sequential dependencies, retries, or stateful execution of a multi-step pipeline.

73
Multi-Selecthard

A company is training a deep learning model for object detection using SageMaker. The training is very slow and the GPU memory is insufficient for the batch size. The team wants to scale across multiple GPUs efficiently. Which THREE actions should they take? (Choose THREE.)

Select 3 answers
A.Use SageMaker distributed model parallelism
B.Use SageMaker distributed data parallelism
C.Use managed spot instances
D.Use a SageMaker distributed training configuration with the SageMaker SDK
E.Enable SageMaker Debugger to identify bottlenecks
AnswersA, B, D

Model parallelism shards a single large model's layers across multiple GPUs, distributing parameters and activations so each device holds only a fraction. This directly addresses insufficient per-GPU memory for the batch size while scaling training across GPUs.

Why this answer

Option A is correct because SageMaker distributed model parallelism shards the model itself across multiple GPUs, which directly addresses the insufficient GPU memory problem by allowing a model too large for a single GPU to be split across devices, and it also speeds up training of large models. Option B is correct because SageMaker distributed data parallelism splits each mini-batch across GPUs and uses AllReduce for efficient gradient synchronization, enabling the team to scale the effective batch size and throughput across multiple GPUs efficiently. Option D is correct because the SageMaker distributed training configuration via the SageMaker SDK (e.g., the Distribution parameter with smdistributed settings in the estimator) is the required mechanism to actually enable and launch model or data parallelism on the training cluster.

Option C is not correct because managed spot instances reduce cost, not training time or GPU memory constraints, and can even introduce interruptions. Option E is not correct because SageMaker Debugger only monitors and reports training bottlenecks; it does not itself scale training across multiple GPUs or resolve memory limitations.

Exam trap

MLA-C01 often tests the distinction between cost-optimization features (spot instances), observability features (Debugger), and actual distributed-training mechanisms — candidates pick spot instances or Debugger thinking they 'help with scaling,' but only the distributed libraries and their SDK configuration address the memory and multi-GPU scaling requirement.

74
Multi-Selecthard

A healthcare company has deployed a SageMaker model that predicts patient risk scores. The security team requires that all access to the model's endpoint be authenticated and authorized, and that every invocation be traceable to a specific user or application for audit purposes. The team also wants to enforce least privilege so that only specific applications can invoke the endpoint. Which TWO actions should the machine learning engineer take to meet these requirements? (Choose two.)

Select 2 answers
A.Enable network isolation on the endpoint to prevent unauthorized outbound calls.
B.Attach an IAM policy to each calling application's IAM role that allows the sagemaker:InvokeEndpoint action on the specific endpoint ARN.
C.Enable AWS CloudTrail data events for the SageMaker endpoint to log every InvokeEndpoint API call.
D.Store the endpoint's invocation URL in AWS Secrets Manager and require applications to retrieve it before calling.
E.Configure the endpoint to use a resource-based policy that allows anonymous invocation from the VPC.
AnswersB, C

IAM policies with the sagemaker:InvokeEndpoint action scoped to the specific endpoint ARN enforce least privilege and ensure that only authorized applications can invoke the endpoint. This also provides authentication and authorization because every call is signed with AWS Signature Version 4 using the caller's IAM credentials, which can be audited.

Why this answer

IAM policies scoped to the specific endpoint ARN enforce authentication and least privilege, ensuring only authorized applications can invoke the endpoint. CloudTrail data events for SageMaker endpoints record each invocation with caller identity and request details, providing the required audit trail. Together they satisfy authentication, authorization, and traceability.

Exam trap

The trap here is thinking that network-level controls like network isolation or storing the URL in Secrets Manager provide authentication and auditability, when they do not identify or authorize callers.

75
MCQhard

A team runs a SageMaker Pipeline that trains a model and registers it in the Model Registry. Compliance requires that the pipeline run automatically every time new labeled data lands in S3, and that each run record the exact S3 data prefix, the training image URI, and the git commit hash as lineage metadata. The engineer wants the least operational overhead. Which approach meets these requirements?

A.Enable SageMaker Data Wrangler scheduled jobs to export data and manually start the pipeline after reviewing the export in SageMaker Studio.
B.Configure an S3 event notification to invoke a Lambda function that calls StartPipelineExecution with the data prefix and commit hash passed as pipeline parameters.
C.Create an EventBridge scheduled rule that runs the pipeline every 15 minutes and relies on the pipeline's caching to skip runs when data is unchanged.
D.Use AWS Step Functions with a Wait state polling S3 every minute, then call StartPipelineExecution once new objects are detected.
AnswerB

S3 event notifications can trigger Lambda on object creation, and the Lambda handler can call StartPipelineExecution with parameter overrides carrying the data prefix and commit hash. Pipeline parameters flow into processing and training steps, where they are recorded as lineage metadata via the Model Registry. This is event-driven, requires no polling infrastructure, and keeps operational overhead low while satisfying automatic triggering and full lineage capture.

Why this answer

S3 event notifications driving a Lambda that invokes StartPipelineExecution provides immediate, event-driven execution with parameter overrides for the data prefix and commit hash. Those parameters propagate into pipeline steps and are captured as lineage metadata in the Model Registry, satisfying compliance. It avoids polling, scheduling waste, and manual gates, giving the lowest operational overhead of the listed designs.

Exam trap

The trap here is treating a scheduled poll or cached pipeline as equivalent to event-driven execution, which misses both trigger immediacy and the lineage metadata requirements.

Page 1 of 9

Page 2

All pages