Courseiva

AWS Certified Machine Learning Engineer Associate MLA-C01 (MLA-C01) — Questions 151–225

665 questions total · 9pages · All types, answers revealed

Page 2

Page 3 of 9

Page 4
151
Multi-Selectmedium

A data engineer is preparing a dataset in Amazon SageMaker Data Wrangler for a binary classification model. The dataset contains missing values in several numeric columns, and the engineer wants a reusable, reproducible transformation that can be applied identically to the training data and to future inference data. The engineer plans to export the transformation and integrate it into a SageMaker Pipeline. Which TWO actions should the engineer take to ensure the imputation is consistent between training and inference? (Choose two.)

Select 2 answers
A.Configure the imputation transform in the Data Wrangler flow so it learns the fill statistic (such as the mean) from the training data during the flow run.
B.Replace missing numeric values with zero in both the training and inference datasets using a custom Python script.
C.Export the Data Wrangler flow to a SageMaker Pipeline and include the generated processing step so it runs on both training and inference inputs.
D.Compute the mean of each numeric column on the full dataset including inference data, then hard-code those values into the pipeline.
E.Delete all rows that contain missing values from the training dataset before training.
AnswersA, C

Learning the fill statistic from the training data ensures the imputation parameters are derived from the correct distribution and not from inference data, which would leak information. Data Wrangler captures these learned parameters as part of the flow, making the transformation reproducible. This is the foundation for applying the same imputation consistently at inference time.

Why this answer

Consistency between training and inference requires that the imputation statistic be learned from training data and that the same transformation be reused at inference. Configuring Data Wrangler to learn the fill statistic from training data captures the parameters, and exporting the flow into a SageMaker Pipeline processing step ensures those exact parameters and logic are applied to both training and inference inputs.

Exam trap

The trap here is treating imputation as a one-time data cleanup rather than a parameterized transformation whose learned statistics must be persisted and reapplied at inference.

152
MCQeasy

A company has a trained machine learning model that needs to be deployed as a real-time inference endpoint on Amazon SageMaker. The endpoint must automatically scale based on incoming traffic. Which SageMaker feature should be used?

A.SageMaker Endpoint Auto Scaling
B.SageMaker Elastic Inference
C.SageMaker Batch Transform
D.SageMaker Model Monitor
AnswerA

SageMaker Endpoint Auto Scaling dynamically adjusts the number of instances behind a real-time endpoint using target-tracking or step-scaling policies, driven by metrics such as InvocationsPerInstance. This directly satisfies the stem's requirement to scale automatically with incoming traffic, unlike batch transform or asynchronous inference, which cannot serve continuous real-time requests.

Why this answer

Amazon SageMaker Endpoint Auto Scaling is the correct feature because it automatically adjusts the number of instances serving a real-time inference endpoint based on the incoming traffic load. It uses Application Auto Scaling policies, which monitor CloudWatch metrics (e.g., InvocationsPerInstance) to scale in or out, ensuring low latency and cost efficiency without manual intervention.

Exam trap

The trap here is that candidates confuse SageMaker Elastic Inference (which accelerates inference) with auto scaling, or they assume Batch Transform can be used for real-time endpoints, but only Endpoint Auto Scaling directly manages dynamic instance count based on traffic.

How to eliminate wrong answers

Option B is wrong because SageMaker Elastic Inference attaches GPU acceleration to an endpoint for low-cost deep learning inference, but it does not handle automatic scaling of the endpoint itself. Option C is wrong because SageMaker Batch Transform is designed for offline, asynchronous batch predictions on entire datasets, not for real-time inference endpoints that require automatic scaling. Option D is wrong because SageMaker Model Monitor tracks data quality, bias, and drift for deployed models, but it does not manage scaling of the endpoint infrastructure.

153
MCQeasy

A company wants to detect anomalies in login events from a large user base, focusing on unusual patterns that may indicate compromised accounts. Which SageMaker built-in algorithm is most suitable for this task?

A.IP Insights
B.K-Means
C.DeepAR
D.Factorisation Machines
AnswerA

IP Insights learns normal patterns of entity-to-IP associations, flagging unusual login behaviour such as an account authenticating from an atypical address. This directly targets compromised-account detection across a large user base, unlike classification or forecasting algorithms that require labelled anomaly data.

Why this answer

IP Insights is a SageMaker built-in unsupervised algorithm designed to learn the relationship between user entities and IP addresses, making it ideal for detecting anomalous login events such as a user logging in from an unusual IP or a compromised account accessing from a new location. It is specifically built for the login-anomaly use case described.

Exam trap

The trap is confusing general anomaly detection algorithms with IP Insights, which is purpose-built for user-IP login anomaly detection and is the only option that directly addresses the scenario.

How to eliminate wrong answers

Option B is wrong because K-Means is a clustering algorithm for grouping similar data points, not for detecting anomalous login patterns tied to user-IP relationships. Option C is wrong because DeepAR is a time-series forecasting algorithm, not an anomaly detection algorithm for login events. Option D is wrong because Factorization Machines are used for recommendation and classification tasks with sparse data, not for user-IP login anomaly detection.

154
MCQeasy

A data scientist has a 200 GB Parquet dataset in Amazon S3 that will be used to train a SageMaker model. The training script reads the data with the SageMaker training toolkit's File mode, and the job currently spends a long time downloading before training begins. The team wants to reduce startup time without changing the training algorithm. Which change should the data scientist make?

A.Switch the input channel to FastFile mode so the training container streams data instead of downloading the full dataset first.
B.Enable Pipe mode and reformat the Parquet files into the recordIO-protobuf format.
C.Add more instances to the training cluster and enable distributed data parallel.
D.Increase the volume size of the training instance so the full dataset fits in local storage.
AnswerA

SageMaker FastFile mode exposes S3 objects to the training container through a POSIX-like interface that streams bytes on demand, so the job can begin reading immediately without a full upfront download. For large Parquet datasets this substantially reduces the time before training starts while leaving the algorithm unchanged.

Why this answer

FastFile mode removes the blocking full download by presenting S3 objects as a streamed filesystem inside the container, so the training process starts as soon as it reads the first bytes. This directly shortens the pre-training wait for a large Parquet dataset and requires no change to the training algorithm or data format.

Exam trap

The trap here is conflating capacity with latency, adding disk space when the real delay is waiting for a complete copy of the data.

155
MCQhard

A company wants to use a pre-trained NLP model from SageMaker JumpStart for sentiment analysis. Which step is required to make predictions?

A.Label the dataset for fine-tuning
B.Train the model from scratch on the company's data
C.Convert the model to ONNX format
D.Deploy the model to an endpoint
AnswerD

Deploying the JumpStart model to a SageMaker endpoint provisions a hosted inference container that serves real-time prediction requests. Without this deployment step, the pre-trained NLP model cannot be invoked for sentiment analysis, so it is required to obtain predictions.

Why this answer

D is correct because SageMaker JumpStart provides pre-trained models that are ready for inference without additional training. To make predictions, you must deploy the model to a SageMaker endpoint, which creates a hosted inference endpoint that can accept input data and return sentiment analysis results.

Exam trap

AWS often tests the misconception that pre-trained models require fine-tuning or additional data preparation before inference, when in fact they can be used directly for predictions after deployment to an endpoint.

How to eliminate wrong answers

Option A is wrong because labeling the dataset for fine-tuning is only necessary if you want to adapt the pre-trained model to a specific domain or task, but it is not required for making predictions with the pre-trained model as-is. Option B is wrong because training from scratch defeats the purpose of using a pre-trained model from JumpStart, which is designed to avoid the cost and time of training from scratch. Option C is wrong because converting the model to ONNX format is an optimization step for cross-platform deployment or performance, but it is not a prerequisite for making predictions with SageMaker JumpStart models, which natively support SageMaker inference.

156
MCQeasy

A company wants to version and track ML models, with an approval workflow for promoting models from staging to production. Which SageMaker feature should they use?

A.SageMaker Model Monitor
B.SageMaker Experiments
C.SageMaker Pipelines
D.SageMaker Model Registry
AnswerD

SageMaker Model Registry stores versioned model groups with metadata and approval status, letting you gate promotion from staging to production through an explicit approval workflow. It directly satisfies the versioning and approval constraint, unlike raw S3 artefacts or plain endpoints, which lack governance states.

Why this answer

SageMaker Model Registry is the correct choice because it provides a centralized repository to catalog, version, and manage ML models, and it supports approval workflows (e.g., PendingApproval, Approved, Rejected) to promote models from staging to production. This directly addresses the requirement for version tracking and an approval gate for model promotion.

Exam trap

The trap here is that candidates confuse SageMaker Pipelines (which orchestrates the workflow) with SageMaker Model Registry (which manages the model versions and approvals), but the question specifically asks for the feature that handles versioning and approval workflow, not the orchestration of the pipeline itself.

How to eliminate wrong answers

Option A is wrong because SageMaker Model Monitor is designed for detecting data and model quality drift in production, not for versioning or approval workflows. Option B is wrong because SageMaker Experiments is used for tracking and comparing training runs (e.g., hyperparameters, metrics), not for managing model versions or approval states. Option C is wrong because SageMaker Pipelines orchestrates end-to-end ML workflows (e.g., data processing, training, deployment) but does not natively provide a model version registry or approval workflow; it can integrate with Model Registry for that purpose.

157
MCQmedium

A financial services company is deploying a model for loan approval. They must ensure that the model's predictions do not show bias against protected groups. They plan to monitor for bias drift after deployment. Which SageMaker feature should they use?

A.SageMaker Model Monitor with data quality monitoring.
B.SageMaker Debugger to capture tensors.
C.SageMaker Ground Truth for fairness labels.
D.SageMaker Clarify with bias drift detection.
AnswerD

SageMaker Clarify computes bias metrics against protected attributes and, with bias drift detection, continuously compares live inference data to the baseline, alerting when disparity shifts. This directly satisfies the requirement to monitor bias drift post-deployment.

Why this answer

SageMaker Clarify is the correct choice because it provides built-in bias detection and monitoring capabilities, including the ability to detect bias drift over time after deployment. It can analyze predictions for protected groups and generate reports on metrics like disparate impact and conditional demographic disparity, which directly addresses the requirement to monitor for bias drift post-deployment.

Exam trap

AWS often tests the distinction between data quality monitoring (which tracks input data drift) and bias drift monitoring (which tracks fairness in predictions), leading candidates to mistakenly choose SageMaker Model Monitor when the question specifically asks about bias.

How to eliminate wrong answers

Option A is wrong because SageMaker Model Monitor with data quality monitoring focuses on detecting changes in the input data distribution (e.g., feature drift), not on bias or fairness metrics for predictions. Option B is wrong because SageMaker Debugger is designed to capture tensors and debug training jobs (e.g., gradient issues, overfitting), not to monitor bias drift in deployed models. Option C is wrong because SageMaker Ground Truth is a data labeling service used to create training datasets with human annotations, not for monitoring bias drift in production predictions.

158
MCQmedium

A machine learning engineer runs a training job and notices the loss is NaN after a few steps. Which SageMaker Debugger rule can help identify this issue?

A.Overfit
B.Exploding gradients
C.Dead ReLU
D.Class imbalance
AnswerB

Exploding gradients detects rapidly increasing gradient magnitudes that cause numerical overflow, producing NaN loss during training. SageMaker Debugger's built-in rule monitors gradient tensors across steps and raises an alert when values exceed thresholds, directly diagnosing the NaN loss the engineer observed.

Why this answer

The SageMaker Debugger built-in 'ExplodingGradients' rule monitors gradient tensors during training and detects when gradient values grow excessively large, which is a classic cause of NaN loss values. When gradients explode, weight updates become enormous, causing numerical overflow and NaN in the loss. The ExplodingGradients rule flags this condition so the engineer can apply gradient clipping or reduce the learning rate.

Exam trap

MLA-C01 often tests the mapping between training symptoms and the specific Debugger rule name — candidates confuse 'Overfit' with 'ExplodingGradients' because both relate to loss behavior, but only ExplodingGradients addresses NaN loss from gradient magnitude.

How to eliminate wrong answers

Option A is wrong because the Overfit rule detects when training loss decreases while validation loss increases, indicating overfitting — it does not detect NaN loss or gradient magnitude issues. Option C is wrong because the DeadReLU rule detects when ReLU activations output zero for most inputs, causing dead neurons and stalled learning, not NaN loss from exploding gradients. Option D is wrong because ClassImbalance is not a SageMaker Debugger built-in rule for detecting NaN loss; class imbalance is a data distribution issue addressed through sampling or loss weighting, not a gradient-monitoring rule.

159
MCQhard

A machine learning engineer is training a deep learning model on SageMaker using the PyTorch estimator. The training job fails with an error indicating that the GPU memory is exhausted. The engineer wants to reduce memory usage without changing the model architecture. Which SageMaker feature should the engineer use?

A.SageMaker data parallelism
B.SageMaker Debugger
C.SageMaker Automatic Model Tuning
D.SageMaker model parallelism
AnswerD

SageMaker model parallelism allows training of large models by partitioning the model across multiple GPUs. This reduces the memory footprint on each GPU, enabling training of models that would otherwise not fit. It is specifically designed to address memory constraints without altering the model architecture, making it the correct choice for this scenario.

Why this answer

SageMaker model parallelism partitions a large model across multiple GPUs, reducing the memory required on each GPU. This directly addresses GPU memory exhaustion without changing the model architecture. Data parallelism replicates the model, which does not help with memory constraints, and the other services are for monitoring or tuning, not memory reduction.

Exam trap

The trap here is confusing data parallelism with model parallelism; data parallelism replicates the model and does not reduce per-GPU memory usage.

160
MCQmedium

A financial institution uses SageMaker to train and deploy models. They need to track every experiment, model version, and deployment step for audit purposes. Which SageMaker feature should they use to capture the full lineage of artifacts, actions, and contexts?

A.SageMaker Clarify
B.SageMaker Model Registry
C.SageMaker Experiments
D.SageMaker ML Lineage Tracking
AnswerD

SageMaker ML Lineage Tracking automatically records relationships among artifacts, actions and contexts across training and deployment, giving the auditable end-to-end history the institution requires. Experiment tracking alone does not capture deployment steps or cross-resource lineage.

Why this answer

SageMaker ML Lineage Tracking is the feature designed to capture the full lineage of artifacts, actions, and contexts, providing an end-to-end audit trail of the machine learning workflow. It automatically records relationships between data, models, and experiments.

Exam trap

MLA-C01 often tests the confusion between SageMaker features, leading candidates to select Model Registry or Experiments when full lineage tracking across all artifacts is required.

How to eliminate wrong answers

Option A is wrong because SageMaker Clarify is used for bias detection and explainability, not lineage tracking. Option B is wrong because SageMaker Model Registry manages model versions and approval status but does not capture the full lineage of all artifacts and actions. Option C is wrong because SageMaker Experiments tracks experiment runs and metrics but does not provide comprehensive lineage across all entities like data and endpoints.

161
MCQhard

A team is deploying a real-time inference endpoint in SageMaker. The model requires access to an S3 bucket containing customer data, which is encrypted with SSE-KMS. The team needs to ensure that the endpoint can decrypt the data. Which IAM role configuration is necessary?

A.Add kms:GenerateDataKey permission to the SageMaker execution role.
B.Attach a policy to the S3 bucket granting s3:GetObject to the KMS key.
C.Add kms:Decrypt permission to the SageMaker execution role for the specific KMS key.
D.Configure the endpoint to assume the S3 bucket's IAM role.
AnswerC

Granting kms:Decrypt on the specific customer-managed key to the SageMaker execution role lets the endpoint's containers call the KMS Decrypt API, satisfying the SSE-KMS constraint. Without this key-policy or IAM allowance, S3 returns AccessDenied on GetObject, since the role, not the bucket, must hold decrypt rights.

Why this answer

The SageMaker execution role must have the kms:Decrypt permission for the specific KMS key that encrypted the S3 objects. When the endpoint reads data from the S3 bucket, SageMaker uses its execution role to call KMS to decrypt the data. Without this permission, the endpoint will fail with an access denied error, even if the S3 bucket policy allows s3:GetObject.

Exam trap

The trap here is that candidates confuse the permissions needed for encryption (kms:GenerateDataKey) with those needed for decryption (kms:Decrypt), or incorrectly think that S3 bucket policies can grant permissions to KMS keys.

How to eliminate wrong answers

Option A is wrong because kms:GenerateDataKey is used to create new data keys for encryption, not to decrypt existing data; the endpoint needs to decrypt, not encrypt. Option B is wrong because attaching a policy to the S3 bucket granting s3:GetObject to the KMS key is syntactically incorrect—KMS keys are not IAM principals, and bucket policies grant actions to principals, not to keys. Option D is wrong because the endpoint cannot assume the S3 bucket's IAM role; IAM roles are assumed by principals (users, services), not by buckets, and SageMaker endpoints use their own execution role for S3 access.

162
Multi-Selecthard

A data science team is using SageMaker Experiments to track hyperparameters and metrics for a model training project. They need to compare multiple trials and identify the best model. Which THREE actions are part of a typical workflow? (Select THREE.)

Select 3 answers
A.Log hyperparameters and metrics using the SageMaker SDK
B.Generate confusion matrices for each trial automatically
C.Use the SageMaker SDK to list trials and compare metrics
D.Create an experiment in SageMaker Experiments
E.Automatically deploy the best trial to an endpoint
AnswersA, C, D

Logging hyperparameters and metrics through the SageMaker SDK writes each trial's parameters and metric values into the experiment run, satisfying the need to capture comparable data across trials. Without this instrumentation, no run records exist for the SDK's analytics and visualisation tools to rank or compare, so identifying the best model becomes impossible.

Why this answer

Option D is correct because a SageMaker Experiments workflow begins by creating an experiment (via the SageMaker SDK, e.g., Experiment.create) that serves as the top-level container grouping runs and trials for the project. Option A is correct because during training the team logs hyperparameters and metrics to the experiment using the SageMaker SDK (for example with Run/Tracker and log_parameter/log_metric calls), which is what makes trials comparable. Option C is correct because the SDK provides APIs such as Experiment.list_runs or Trial/TrialComponent lookups and metric retrieval so the team can list trials and compare their metrics to identify the best model.

Option B does not belong because SageMaker Experiments does not automatically generate confusion matrices for each trial; any such artifact must be computed and logged explicitly by the training code. Option E does not belong because automatic deployment of the best trial to an endpoint is not part of the Experiments tracking workflow; deployment is a separate step typically handled via the SageMaker model registry, pipelines, or manual deployment.

Exam trap

MLA-C01 often tests what SageMaker Experiments does versus what adjacent services do — the trap is assuming Experiments auto-generates evaluation artifacts or auto-deploys models, which are responsibilities of other tools.

163
MCQeasy

A data scientist wants to train a binary classification model using Amazon SageMaker with a built-in algorithm that performs well on tabular data. Which algorithm should they choose?

A.Image Classification
B.DeepAR
C.XGBoost
D.BlazingText
AnswerC

XGBoost is a gradient-boosted decision tree algorithm built into Amazon SageMaker, and it excels on tabular datasets for binary classification through boosted ensemble learning. It directly satisfies the stem's requirement for a built-in algorithm performing well on tabular data, unlike linear or deep learning alternatives.

Why this answer

XGBoost is a built-in Amazon SageMaker algorithm optimized for tabular and structured data, and it consistently performs well on binary classification tasks. It implements a gradient-boosted decision tree framework that handles missing values, categorical features, and class imbalance reasonably well out of the box. For a binary classification problem on tabular data, XGBoost is the standard choice among SageMaker built-in algorithms.

Exam trap

MLA-C01 often tests algorithm-to-task mapping — candidates may pick DeepAR or BlazingText because they sound sophisticated, but the exam expects recognition that XGBoost is the go-to built-in for tabular binary classification.

How to eliminate wrong answers

Option A is wrong because Image Classification is designed for computer vision tasks on image data, not tabular binary classification. Option B is wrong because DeepAR is a time-series forecasting algorithm, not a classifier for tabular data. Option D is wrong because BlazingText is a text classification and word embedding algorithm for natural language data, not tabular binary classification.

164
MCQmedium

A company is deploying a large NLP model on SageMaker for real-time inference. They want to reduce inference latency and cost by optimizing the model for the target hardware. The model is trained in PyTorch. Which SageMaker feature should they use to compile the model for best performance on the chosen instance?

A.SageMaker Neo
B.AWS Step Functions
C.Amazon Elastic Inference
D.SageMaker Triton Inference Server
AnswerA

SageMaker Neo compiles PyTorch models into optimised executables tuned to the target instance's specific processor architecture, cutting inference latency and cost. It satisfies the stem's requirement to compile the trained model for best performance on the chosen hardware, unlike generic deployment or autoscaling features that leave the model graph unoptimised.

Why this answer

SageMaker Neo is the correct choice because it is specifically designed to compile trained models (including PyTorch models) into an optimized binary for a target hardware instance, reducing inference latency and improving throughput. Neo applies hardware-specific optimizations such as operator fusion, memory layout tuning, and quantization, which directly address the need for best performance on the chosen SageMaker instance.

Exam trap

The trap here is that candidates confuse model compilation (Neo) with inference serving (Triton) or hardware acceleration (Elastic Inference), leading them to pick a service that addresses a different part of the inference pipeline.

How to eliminate wrong answers

Option B is wrong because AWS Step Functions is a serverless workflow orchestration service, not a model compilation tool; it cannot optimize model performance for hardware. Option C is wrong because Amazon Elastic Inference attaches a separate accelerator to an instance for cost-effective inference, but it does not compile or optimize the model itself; it only provides additional compute resources. Option D is wrong because SageMaker Triton Inference Server is a high-performance inference server that supports multiple frameworks and model formats, but it does not compile the model for the target hardware; it serves models as-is, relying on the underlying framework's runtime.

165
Multi-Selectmedium

A machine learning engineer must grant a data scientist the least-privilege permissions needed to invoke one specific SageMaker real-time endpoint from their own AWS account, and to view that endpoint's CloudWatch metrics without being able to modify the endpoint. The endpoint ARN is known. Which TWO IAM policy statements should the engineer include? (Choose two.)

Select 2 answers
A.Allow cloudwatch:PutMetricAlarm on all resources so the data scientist can create alarms on endpoint metrics.
B.Allow sagemaker:CreateEndpointConfig so the data scientist can tune the endpoint's instance type.
C.Allow sagemaker:InvokeEndpoint on the specific endpoint ARN.
D.Allow sagemaker:UpdateEndpoint on the specific endpoint ARN.
E.Allow cloudwatch:GetMetricData on the endpoint's metrics and allow cloudwatch:GetMetricStatistics for the relevant namespace.
AnswersC, E

InvokeEndpoint is the runtime action used to send inference requests to a real-time endpoint, and scoping the resource to the exact endpoint ARN grants access only to that endpoint rather than all endpoints in the account. This is the minimum permission required for the data scientist to call the model, and it does not confer any ability to change the endpoint's configuration.

Why this answer

Least privilege here means two read/invoke capabilities and nothing that changes the endpoint. InvokeEndpoint scoped to the endpoint ARN provides inference access, and the CloudWatch read actions GetMetricData and GetMetricStatistics provide visibility into endpoint metrics. Management actions such as UpdateEndpoint, PutMetricAlarm, and CreateEndpointConfig are excluded because they either modify the endpoint or exceed the requested permissions.

Exam trap

The trap here is bundling a write-oriented monitoring permission such as PutMetricAlarm with the read-only metrics permissions, when viewing metrics only requires the CloudWatch read actions.

166
MCQmedium

A healthcare analytics team stores model artifacts and training datasets in Amazon S3 and uses SageMaker. An internal audit finds that some S3 buckets containing protected health information are missing encryption and that access is granted broadly. The team must remediate quickly and prevent future misconfiguration. Which combination of actions should the ML engineer take FIRST?

A.Apply default bucket encryption with a customer-managed KMS key, enable S3 Block Public Access, and tighten bucket policies to least privilege.
B.Enable S3 server access logging and CloudTrail data events, then review the logs to identify who accessed the unencrypted data.
C.Enable AWS Config rules to detect unencrypted buckets and use AWS Security Hub to aggregate findings for remediation.
D.Migrate all datasets and artifacts to a new bucket and delete the original buckets to eliminate the misconfiguration.
AnswerA

Default bucket encryption applies SSE-KMS with a customer-managed key to all new objects, addressing the encryption gap. S3 Block Public Access and tightened bucket policies remove broad access. Together these directly remediate the audit findings and establish secure defaults that prevent recurrence, which is the appropriate first action.

Why this answer

Applying default SSE-KMS encryption with a customer-managed key secures new objects, while S3 Block Public Access and least-privilege bucket policies eliminate broad access. These actions directly fix the audit findings and set secure defaults to prevent recurrence. Logging, detection services, and bucket migration address visibility or workaround concerns rather than correcting the misconfiguration itself.

Exam trap

The trap here is choosing monitoring or logging services as the fix, when detection does not remediate existing unencrypted or broadly accessible buckets.

167
MCQeasy

Refer to the exhibit. A data scientist reviews the CloudWatch Logs from an Amazon SageMaker real-time endpoint. What is the MOST likely root cause of the NaN output?

A.The model weights became corrupted due to a disk write error.
B.The input data contains out-of-range values not seen during training, causing the model to output NaN.
C.The endpoint is overloaded and returning a default NaN response.
D.The model artifact failed to load correctly, resulting in NaN weights.
AnswerB

Out-of-range inputs violate the feature distribution the model learned, so activations or gradients overflow and produce NaN. This directly satisfies the stem's real-time endpoint scenario: CloudWatch Logs capture inference-time anomalies, and unseen extreme values are the most likely trigger rather than training or infrastructure faults.

Why this answer

The NaN (Not a Number) output from a SageMaker real-time endpoint is most commonly caused by input data containing values outside the range seen during training. This can lead to numerical instability in the model's forward pass, such as division by zero, log of zero, or exponent overflow, which propagates through layers and results in NaN predictions.

Exam trap

AWS often tests the misconception that NaN outputs are caused by infrastructure issues like overload or file corruption, when the actual root cause is almost always data-related numerical instability in the model's inference logic.

How to eliminate wrong answers

Option A is wrong because disk write errors would corrupt the model artifact file, causing a failure to load the model or produce a different error (e.g., model loading failure), not a NaN output during inference. Option C is wrong because an overloaded endpoint would typically return HTTP 429 (Too Many Requests) or 503 (Service Unavailable) errors, not a default NaN response; SageMaker does not have a built-in mechanism to return NaN for overload conditions. Option D is wrong because if the model artifact failed to load correctly, the endpoint would fail to deploy or return an error during invocation, not produce NaN outputs; a successful deployment implies the weights loaded correctly.

168
MCQhard

A machine learning engineer runs a SageMaker Processing job that must load a 200 GB dataset from S3, compute statistics, and write a small summary to S3. The job repeatedly fails with an out-of-disk-space error on the processing instance. Which change is MOST likely to resolve the failure?

A.Add a lifecycle configuration script that deletes files from the input directory after each epoch.
B.Raise the max_runtime_in_seconds parameter and enable network isolation on the processing job.
C.Increase the processing instance's volume size and use ShardedByS3Key so each instance downloads only its share of objects.
D.Switch the input channel to Pipe mode so the container streams records instead of downloading files.
AnswerC

The failure is caused by the local volume filling with downloaded input data. Enlarging the attached EBS volume provides headroom, and ShardedByS3Key distributes the S3 objects across instances so no single instance downloads the entire dataset. Together these directly address the root cause rather than merely retrying the job.

Why this answer

Processing jobs download input channels to local storage, so a 200 GB dataset can exhaust the default volume. Enlarging the volume gives the container room, and ShardedByS3Key spreads objects across multiple instances so each holds only a fraction. Pipe mode is not applicable to Processing jobs, and the other options adjust runtime or cleanup behavior without changing disk consumption.

Exam trap

The trap here is assuming Pipe mode applies to every SageMaker job type, when streaming input is a training-job feature and Processing jobs always mount files locally.

169
MCQmedium

A machine learning engineer uses Amazon SageMaker Data Wrangler to preprocess a dataset. After applying a transform, the engineer wants to export the data to a feature group in Amazon SageMaker Feature Store for reuse in training and inference. Which export option should they choose?

A.Export to Amazon DynamoDB
B.Export to Amazon SageMaker Feature Store
C.Export to Amazon S3 as CSV
D.Export to Amazon SageMaker Pipelines
AnswerB

Exporting directly to Amazon SageMaker Feature Store writes the transformed data into a feature group, satisfying the requirement for reuse across both training and inference. This native Data Wrangler destination handles the offline and online store writes, avoiding intermediate Amazon S3 staging and manual ingestion jobs that other export options would demand.

Why this answer

SageMaker Data Wrangler can export directly to a feature group in Feature Store, making the features available for both training (offline) and inference (online).

170
Multi-Selectmedium

A data scientist is preparing text data for a sentiment analysis model using Amazon SageMaker. Which two data preprocessing techniques are commonly used when working with text data for natural language processing? (Choose two.)

Select 2 answers
A.One-hot encoding of all words
B.Image resizing
C.Tokenization
D.Principal component analysis (PCA)
E.Stop word removal
AnswersC, E

Tokenization splits raw text into individual words or subword units, converting unstructured strings into discrete tokens that models can vectorise. This satisfies the text-preprocessing requirement, forming the essential first step before embedding or vectorisation in NLP pipelines.

Why this answer

Tokenization (C) is correct because NLP models require raw text to be split into individual tokens (words, subwords, or characters) so they can be mapped to numerical vectors before being fed into a sentiment analysis model in SageMaker. Stop word removal (E) is correct because eliminating high-frequency, low-information words such as 'the', 'is', and 'and' reduces noise and dimensionality, helping the model focus on sentiment-bearing terms. One-hot encoding of all words (A) is not a common standalone preprocessing technique for text at scale, since it produces extremely sparse, high-dimensional vectors and ignores word order and semantics; embeddings are typically preferred.

Image resizing (B) applies to computer vision data, not text. Principal component analysis (D) is a dimensionality reduction technique for numerical feature matrices, not a standard text preprocessing step for NLP.

Exam trap

The trap here is that candidates may confuse one-hot encoding as a preprocessing technique for raw text, when it is actually a feature engineering step applied after tokenization, and they may overlook that stop word removal is a standard preprocessing step despite its potential to remove sentiment-bearing words in certain contexts.

171
MCQhard

A company is training a deep learning model using SageMaker and wants to reduce the time spent on data loading from Amazon S3 during training. The training dataset consists of many small files. Which approach is MOST effective to accelerate data loading?

A.Use SageMaker Pipe mode to stream data directly from S3.
B.Increase the number of training instances in the cluster.
C.Enable SageMaker Debugger to monitor data loading metrics.
D.Package the small files into a few large files (e.g., TFRecord or RecordIO) and use File mode.
AnswerD

Combining many small files into a few large files reduces the number of S3 requests and improves I/O throughput. Using File mode then downloads these larger files quickly to the training instance's storage, allowing the training script to read them efficiently. This approach is the most effective for accelerating data loading when dealing with many small files.

Why this answer

Many small files cause high S3 request overhead. Consolidating them into a few large files in a format like TFRecord or RecordIO reduces the number of requests and allows efficient sequential reads. File mode then downloads these large files quickly to local storage, significantly speeding up data loading compared to streaming many small files or adding instances.

Exam trap

The trap here is assuming that Pipe mode always accelerates data loading, but for many small files its per-file overhead can make it slower than consolidating files and using File mode.

172
MCQeasy

A data scientist wants to track feature definitions, share them across teams, and serve features for both training and real-time inference. Which AWS service provides these capabilities?

A.Amazon DynamoDB
B.Amazon S3
C.AWS Glue Data Catalog
D.Amazon SageMaker Feature Store
AnswerD

Amazon SageMaker Feature Store provides a centralised repository that stores feature definitions with metadata, enabling discovery and reuse across teams. It supports both online and offline stores, satisfying the stem's dual requirement: low-latency retrieval for real-time inference and bulk access for training datasets.

Why this answer

Amazon SageMaker Feature Store is purpose-built for ML workflows, providing a centralized repository to define, share, and serve features for both training (batch) and real-time inference (low-latency retrieval). It supports offline and online stores, enabling consistent feature definitions across teams and automatic feature ingestion via SageMaker Pipelines or custom code.

Exam trap

The trap here is that candidates may confuse a general-purpose storage or catalog service (like S3 or Glue Data Catalog) with a purpose-built ML feature store, overlooking the need for both offline and online serving with feature-specific management.

How to eliminate wrong answers

Option A is wrong because Amazon DynamoDB is a NoSQL key-value and document database optimized for transactional workloads, not for managing feature definitions, versioning, or serving features specifically for ML training and inference. Option B is wrong because Amazon S3 is an object storage service that can store feature data but lacks built-in capabilities for feature definition management, sharing across teams, or low-latency online serving for real-time inference. Option C is wrong because AWS Glue Data Catalog is a metadata repository for data sources and ETL jobs, not a feature store; it does not provide online serving or feature-specific versioning and sharing for ML.

173
Multi-Selectmedium

Which TWO of the following are best practices for deploying machine learning models on SageMaker? (Select TWO.)

Select 2 answers
A.Store model artifacts in Amazon EBS volumes attached to the endpoint instances
B.Use separate production and staging endpoints to test new models before full rollout
C.Manually track model versions using tags because SageMaker Model Registry is not available for deployment
D.Disable CloudWatch Logs to reduce costs during inference
E.Enable data capture on endpoints to log predictions for auditing and model monitoring
AnswersB, E

Staging endpoints let you validate a new model version against test traffic before shifting production traffic, isolating deployment risk. This satisfies the best-practise requirement for safe rollout, since production remains unaffected until the staged model is verified.

Why this answer

Option B is correct because using separate staging and production endpoints lets you validate a new model version against real traffic patterns and roll back safely before promoting it to the production endpoint, which is a core SageMaker deployment best practice. Option E is correct because SageMaker Data Capture on endpoints logs request/response payloads to Amazon S3, enabling auditing, model monitoring, and drift detection via Model Monitor. Option A is wrong because model artifacts should be stored in Amazon S3 and loaded by the endpoint, not on EBS volumes attached to instances.

Option C is wrong because SageMaker Model Registry does exist and is the recommended way to version and track models, not manual tags. Option D is wrong because disabling CloudWatch Logs removes observability and is not a best practice for production inference.

Exam trap

The trap is selecting cost-saving or manual options (disabling logs, manual tagging) that sound pragmatic but violate observability and governance best practices — the exam rewards automated, auditable, and versioned deployment patterns.

174
MCQhard

A machine learning engineer is performing feature selection for a regression model with 200 features. The dataset has 10,000 samples. The engineer wants to remove irrelevant features while keeping those that have a strong non-linear relationship with the target. Which feature selection method is best suited for this requirement?

A.Lasso regularization (L1)
B.Recursive feature elimination (RFE) with a linear model
C.Pearson correlation coefficient
D.Mutual information
AnswerD

Mutual information captures arbitrary non-linear dependence between each feature and the target, unlike correlation-based filters that detect only linear relationships. With 10,000 samples across 200 features it remains computationally practical, satisfying the non-linear relevance requirement.

Why this answer

Mutual information captures any kind of dependency between a feature and the target, including non-linear relationships, making it the best choice when the engineer wants to detect strong non-linear associations. Unlike correlation-based methods, mutual information does not assume a linear relationship and can identify features that Pearson or linear-model-based methods would miss. It is well-suited for the 200-feature, 10,000-sample regression scenario.

Exam trap

The trap is assuming that correlation or linear-model-based selection is sufficient for all relationships — candidates may pick Pearson or Lasso without noticing the question explicitly requires capturing non-linear relationships, which only mutual information handles among the options.

How to eliminate wrong answers

Option A is wrong because Lasso (L1) regularization performs feature selection implicitly through linear coefficients, so it can only capture linear relationships and may discard features with non-linear but strong predictive power. Option B is wrong because RFE with a linear model also relies on linear coefficients to rank features, so it inherits the same limitation of missing non-linear dependencies. Option C is wrong because the Pearson correlation coefficient measures only linear correlation between two continuous variables and will assign near-zero scores to features with non-linear relationships, causing them to be incorrectly removed.

175
MCQeasy

Refer to the exhibit. A data scientist ran a training job using a custom algorithm container. The job failed with the error shown. What is the most likely cause?

A.The S3 output path is incorrect
B.The algorithm script references an undefined variable or metric named 'loss'
C.The training image is not accessible
D.The instance type is insufficient
AnswerB

The container's training script references a metric named 'loss' that was never defined, so the job aborts when the logging call executes. This satisfies the stem's constraint: the custom algorithm container itself is faulty, not the compute configuration or data pipeline.

Why this answer

The error shown in the exhibit indicates that the training job failed because the algorithm script references an undefined variable or metric named 'loss'. In SageMaker custom algorithm containers, the training script must define and log metrics explicitly; if the script tries to use a metric like 'loss' without defining it or if it's misspelled, the job fails with a NameError or similar. The other options relate to infrastructure or configuration issues that would produce different error messages.

Exam trap

MLA-C01 often tests the ability to distinguish between infrastructure errors and script-level errors, tricking candidates into blaming S3 paths or instance types when the error is a simple undefined variable.

How to eliminate wrong answers

Option A is wrong because an incorrect S3 output path would typically result in an Access Denied or NoSuchBucket error, not a reference to an undefined variable. Option C is wrong because an inaccessible training image would cause an image pull error or a failure to start the container, not a script-level variable error. Option D is wrong because an insufficient instance type would cause resource exhaustion errors (e.g., out of memory) or job timeouts, not a variable reference error.

176
MCQmedium

A team is using AWS Step Functions to orchestrate a machine learning workflow that includes data preprocessing, training, and model evaluation. The team wants to run the workflow whenever new data arrives in an S3 bucket. Which approach should they use to trigger the Step Functions workflow?

A.Configure the S3 bucket to send an event notification directly to the Step Functions state machine.
B.Use S3 event notifications to send a message to an Amazon SQS queue, and have a Lambda function poll the queue to start the execution.
C.Use a CloudWatch Logs metric filter to trigger the Step Functions execution.
D.Configure the S3 bucket to send events to Amazon EventBridge, and create an EventBridge rule that targets the Step Functions state machine.
AnswerD

S3 event notifications route to EventBridge, whose rules can target a Step Functions state machine directly, giving event-driven, filtered triggering as objects arrive. This satisfies the new-data trigger constraint without polling, unlike Lambda-based polling or scheduled executions that add latency and cost.

Why this answer

Amazon S3 can send event notifications directly to Amazon EventBridge, and EventBridge rules can target AWS Step Functions state machines as a target. This provides a fully managed, serverless integration that allows the Step Functions workflow to be triggered automatically whenever new data arrives in the S3 bucket, without needing intermediate polling or custom code.

Exam trap

The trap here is that candidates may assume S3 can directly invoke Step Functions (Option A) because they know S3 can trigger Lambda, but they overlook that Step Functions is not a supported direct destination for S3 event notifications.

How to eliminate wrong answers

Option A is wrong because S3 event notifications cannot directly target a Step Functions state machine; S3 event notifications support only Lambda, SQS, SNS, and EventBridge as destinations. Option B is wrong because while it would work, it introduces unnecessary complexity and latency by requiring a Lambda function to poll an SQS queue, which is not the simplest or most efficient approach when EventBridge provides direct integration. Option C is wrong because CloudWatch Logs metric filters are designed to monitor log data and trigger alarms or metrics, not to trigger Step Functions executions; they cannot directly invoke a state machine.

177
MCQmedium

A data engineer is building a data pipeline for a machine learning model that requires both structured and unstructured data. The structured data (customer demographics) is in Amazon RDS, and the unstructured data (customer support chat logs) is in Amazon S3 as JSON files. The engineer needs to combine these datasets into a single training dataset stored in S3 in Parquet format. They must also perform feature engineering such as text vectorization on the chat logs. The pipeline should be serverless and cost-effective. Which approach should they use?

A.Use a SageMaker Processing job with a custom Python script that reads from both sources and writes to S3.
B.Use Amazon Athena to join the data from RDS and S3, then export the results as Parquet.
C.Use AWS Glue ETL with a Spark script that reads from RDS (via JDBC) and S3, performs transformations, and writes Parquet.
D.Use Amazon Kinesis Data Analytics to read from RDS and S3 and produce a continuous stream of processed data.
AnswerC

AWS Glue ETL provides a serverless Spark environment that reads RDS through JDBC and S3 JSON concurrently, then applies ML Transform for text vectorisation. It writes Parquet directly to S3, satisfying the serverless, cost-effective constraint without managing clusters or provisioning infrastructure.

Why this answer

AWS Glue ETL with a Spark script is the correct choice because it natively supports reading from both Amazon RDS (via JDBC) and Amazon S3 (JSON), performing complex transformations like text vectorization, and writing the output as Parquet. Glue is serverless, cost-effective (pay per DPU-hour), and fully managed, making it ideal for batch ETL pipelines that combine structured and unstructured data for ML training.

Exam trap

The trap here is that candidates often choose SageMaker Processing (Option A) because it is associated with ML, but they overlook that Glue ETL is the designated AWS service for serverless data preparation and transformation, especially when combining disparate data sources like RDS and S3.

How to eliminate wrong answers

Option A is wrong because SageMaker Processing jobs are designed for ML-specific tasks like training or inference, not general-purpose ETL; they lack native JDBC connectors for RDS and require custom networking setup, increasing complexity and cost. Option B is wrong because Amazon Athena cannot perform feature engineering like text vectorization; it is an interactive query service for SQL-on-data, not a transformation engine, and cannot write Parquet with custom logic. Option D is wrong because Kinesis Data Analytics is for real-time stream processing, not batch ETL; it would introduce unnecessary latency and cost for a one-time or scheduled training dataset generation, and it cannot directly write Parquet to S3 without additional sinks.

178
Multi-Selectmedium

A machine learning engineer is preparing a training job on SageMaker with a custom Docker container. Which TWO actions are required to use the container with SageMaker? (Choose TWO.)

Select 2 answers
A.Push the container image to Amazon ECR
B.Use a SageMaker Estimator with image_uri parameter pointing to the ECR image
C.Upload the container image to Amazon S3
D.Enable SageMaker Debugger to monitor the custom container
E.Register the container in SageMaker Model Registry
AnswersA, B

SageMaker pulls custom training images from Amazon ECR, so the image must be pushed there and its URI supplied to the estimator. This satisfies the requirement to host the container where SageMaker can access it.

Why this answer

Option A is correct because SageMaker can only pull custom training container images from Amazon ECR, so the image must be built and pushed to an ECR repository that the SageMaker execution role can access. Option B is correct because the SageMaker Estimator must be configured with the image_uri parameter set to the ECR image URI (for example, <account>.dkr.ecr.<region>.amazonaws.com/<repo>:<tag>) so SageMaker knows which container to run for training. Option C is incorrect because container images cannot be stored or executed from Amazon S3; S3 is used for training data and model artifacts, not Docker images.

Option D is incorrect because SageMaker Debugger is an optional monitoring and profiling feature, not a requirement for using a custom container. Option E is incorrect because the SageMaker Model Registry is used to catalog trained models for governance and deployment, and is not needed to run a custom training container.

Exam trap

The trap here is confusing the storage location for container images (ECR) with other AWS storage services like S3, and assuming that optional monitoring or registry features are required for custom container usage.

179
MCQmedium

A data scientist is building a binary classification model on a highly imbalanced dataset where the positive class represents only 1% of the data. The scientist needs to train the model using Amazon SageMaker's built-in XGBoost algorithm. Which strategy should be used to address the class imbalance?

A.Undersample the majority class until the dataset is balanced
B.Set the `scale_pos_weight` hyperparameter to `sum(negative cases) / sum(positive cases)`
C.Use the `max_delta_step` hyperparameter to increase the learning rate for the majority class
D.Use SMOTE to oversample the minority class before passing the data to XGBoost
AnswerB

Setting `scale_pos_weight` to the negative-to-positive ratio (roughly 99) reweights the positive class's gradient contributions during boosting, so the 1% minority class influences tree splits proportionally to its cost. This directly counteracts the imbalance in SageMaker's built-in XGBoost without resampling the training data.

Why this answer

XGBoost's `scale_pos_weight` hyperparameter is specifically designed to handle class imbalance by adjusting the weight of the positive class during training. Setting it to `sum(negative cases) / sum(positive cases)` (i.e., 99/1 = 99) tells the algorithm to penalize misclassifications of the minority class more heavily, effectively balancing the gradient updates. This is the recommended approach for built-in XGBoost in SageMaker, as it directly modifies the loss function without altering the dataset.

Exam trap

The trap here is that candidates often confuse `scale_pos_weight` with resampling techniques (like SMOTE or undersampling) or with other hyperparameters like `max_delta_step`, assuming any imbalance-handling method is equally valid, but the exam expects knowledge of the specific built-in mechanism for SageMaker's XGBoost.

How to eliminate wrong answers

Option A is wrong because undersampling the majority class discards valuable data, which can lead to loss of important patterns and reduced model performance, especially when the dataset is large and the imbalance is extreme (1% positive). Option C is wrong because `max_delta_step` controls the step size in tree boosting to prevent overfitting on imbalanced data, but it does not increase the learning rate for the majority class; it caps the update magnitude for all classes, and is not a direct imbalance correction mechanism. Option D is wrong because while SMOTE can be used to oversample the minority class, it is not a built-in feature of SageMaker's XGBoost algorithm and requires preprocessing outside the training job; moreover, SMOTE can introduce synthetic noise and is less efficient than using `scale_pos_weight` directly.

180
MCQmedium

An ML team uses AWS Step Functions to orchestrate a multi-step inference pipeline: data preprocessing, model inference, and postprocessing. The pipeline runs on demand for single records. The team notices that the pipeline occasionally fails due to timeouts in the preprocessing step. They want to implement retries with exponential backoff and a maximum retry count of 3 for that step. How should they configure this?

A.Implement retry logic inside the preprocessing Lambda function code.
B.Modify the Step Functions state machine definition to add a Retry field on the preprocessing state with a maximum retry count of 3 and an exponential backoff rate of 2.0.
C.Wrap the preprocessing step in a SageMaker Pipeline step with retry policy.
D.Add a Catch in the state machine to rerun the entire pipeline if preprocessing fails.
AnswerB

Step Functions handles transient failures declaratively via the Retry field on a state. Setting MaxAttempts to 3 with BackoffRate 2.0 applies exponential backoff to the preprocessing state, satisfying the required retry count without custom error-handling code.

Why this answer

AWS Step Functions natively supports retry logic with exponential backoff directly in the state machine definition. By adding a `Retry` field on the preprocessing state with `MaxAttempts: 3` and `BackoffRate: 2.0`, the service automatically retries the step on specified errors (e.g., `States.Timeout` or `Lambda.ServiceException`) with exponentially increasing wait times, without requiring custom code or external orchestration.

Exam trap

The trap here is that candidates often assume retry logic must be coded inside the Lambda function (Option A) or that a Catch block (Option D) is the correct way to handle failures, but Step Functions provides a declarative Retry mechanism that is more robust and easier to maintain for orchestrated workflows.

How to eliminate wrong answers

Option A is wrong because implementing retry logic inside the Lambda function code would not leverage Step Functions' built-in exponential backoff and would require custom sleep logic, increasing complexity and violating the separation of concerns between orchestration and business logic. Option C is wrong because SageMaker Pipeline steps are designed for batch training and model building workflows, not for orchestrating a lightweight inference pipeline with single-record processing; wrapping a preprocessing Lambda in a SageMaker Pipeline step adds unnecessary overhead and does not natively support the simple retry policy needed here. Option D is wrong because adding a `Catch` to rerun the entire pipeline on preprocessing failure would restart all steps (including inference and postprocessing), wasting compute time and resources, whereas a targeted retry on only the preprocessing step is more efficient and aligns with the requirement.

181
Multi-Selectmedium

A data scientist is preparing a dataset for a linear regression model. The features have different scales: one feature ranges from 0 to 1000, another from 0 to 1, and a third from -5 to 5. The scientist wants to ensure that all features contribute equally to the model. Which TWO scaling techniques should the scientist consider? (Select TWO.)

Select 2 answers
A.MinMaxScaler (min-max normalization)
B.Principal Component Analysis (PCA)
C.One-hot encoding
D.Log transformation
E.StandardScaler (z-score normalization)
AnswersA, E

MinMaxScaler rescales each feature to a fixed 0–1 range by subtracting the minimum and dividing by the range, so the 0–1000 feature no longer dominates the 0–1 and -5–5 features. This equalises their contribution to the linear regression, satisfying the equal-contribution constraint in the stem.

Why this answer

MinMaxScaler (A) is correct because it rescales each feature to a fixed range, typically [0, 1], using the formula (x - min) / (max - min), which directly addresses the differing ranges (0-1000, 0-1, -5-5) and puts all features on a comparable scale so they contribute equally to the linear regression. StandardScaler (E) is also correct because it standardizes features by removing the mean and scaling to unit variance using z = (x - μ) / σ, which is a standard and effective way to equalize feature contributions when scales differ. PCA (B) is not a scaling technique; it is a dimensionality-reduction method and does not by itself normalize feature ranges.

One-hot encoding (C) is for converting categorical variables into binary vectors, not for rescaling numeric features. Log transformation (D) is a nonlinear transform that can reduce skew but does not guarantee equal scaling across features with different ranges.

Exam trap

MLA-C01 often tests the misconception that any preprocessing step (PCA, one-hot encoding, log transform) counts as scaling, when only MinMaxScaler and StandardScaler directly normalize feature ranges.

182
MCQmedium

An engineer runs: aws sagemaker describe-endpoint --endpoint-name my-endpoint and receives the exhibit output. The engineer wants to update the endpoint to use a new model version stored in ECR with tag ':2'. Which step is necessary to perform the update?

A.Create a new endpoint configuration (my-endpoint-config-v2) referencing the new image, then call update-endpoint with the new config name.
B.Modify the existing endpoint configuration (my-endpoint-config-v1) to use the new image, then update the endpoint.
C.Use the update-endpoint command directly with the new image ARN.
D.Delete the endpoint and recreate it with the new model image.
AnswerA

SageMaker endpoints are immutable; they run a fixed endpoint configuration. Updating the model image requires creating a new configuration referencing the ':2' image, then calling update-endpoint with that configuration name, which triggers a rolling deployment.

Why this answer

SageMaker endpoints are immutable with respect to their configuration; you cannot modify an existing endpoint configuration in place. To update an endpoint to use a new model version, you must create a new endpoint configuration (e.g., my-endpoint-config-v2) that points to the new ECR image tag ':2', then call update-endpoint with the new configuration name. This triggers a zero-downtime deployment where SageMaker gradually shifts traffic to the new variant.

Exam trap

The trap here is that candidates assume endpoint configurations are mutable like a text file, but AWS SageMaker enforces immutability — you must create a new configuration for any change, even a simple image tag update.

How to eliminate wrong answers

Option B is wrong because SageMaker endpoint configurations are immutable after creation; you cannot modify an existing configuration (my-endpoint-config-v1) to reference a new image — you must create a new configuration. Option C is wrong because the update-endpoint command does not accept a direct image ARN; it only accepts an endpoint configuration name, and the model image is specified within that configuration. Option D is wrong because deleting and recreating the endpoint would cause downtime and is unnecessary; SageMaker supports rolling updates via update-endpoint with a new configuration, which avoids service interruption.

183
Multi-Selectmedium

A team wants to evaluate a binary classification model for credit risk. They need to understand the trade-off between false positives and false negatives. Which TWO metrics should they use? (Select TWO.)

Select 2 answers
A.Recall
B.Precision
C.NDCG
D.AUC-ROC
E.RMSE
AnswersA, B

Recall measures the proportion of actual positives correctly identified, directly exposing false negatives — missed credit risks. Pairing it with precision, which exposes false positives, quantifies the trade-off the team needs. Recall alone satisfies the false-negative half of the stem's requirement, making it one of the two correct metrics.

Why this answer

Recall (A) is correct because it measures the proportion of actual positives (e.g., true defaults) that the model correctly identifies, directly quantifying the cost of false negatives, which is critical in credit risk where missing a defaulter is costly. Precision (B) is correct because it measures the proportion of predicted positives that are actually positive, directly quantifying the cost of false positives, such as rejecting good customers; together, recall and precision expose the false-positive/false-negative trade-off. NDCG (C) is a ranking-quality metric for graded relevance in search/recommendation, not binary classification.

AUC-ROC (D) summarizes ranking performance across thresholds but does not separately expose the precision/recall trade-off the team wants to evaluate. RMSE (E) is a regression error metric and is inappropriate for binary classification.

Exam trap

The trap here is selecting AUC-ROC as a metric for understanding the trade-off between false positives and false negatives, when AUC-ROC provides a threshold-independent summary rather than the direct trade-off that precision and recall offer.

184
MCQeasy

A data engineer needs to catalog metadata from multiple data sources across the organization for use in ML workflows. Which AWS Glue component should be used to store and manage this metadata?

A.AWS Glue Crawler
B.AWS Glue Studio
C.AWS Glue Data Catalog
D.AWS Glue ETL
AnswerC

The Glue Data Catalog is the central, Hive-compatible metastore that stores table definitions, schemas and partition metadata from crawlers across sources, giving ML workflows a single queryable catalogue. It satisfies the stem's requirement to store and manage metadata rather than transform or move data.

Why this answer

The AWS Glue Data Catalog is a persistent, centralized metadata repository that stores table definitions, schemas, and partition information for data sources across the organization. It integrates natively with Athena, Redshift Spectrum, EMR, and Glue ETL jobs, making it the correct component for cataloging metadata used in ML workflows. Crawlers populate the Data Catalog, but the Catalog itself is the storage and management layer.

Exam trap

MLA-C01 often tests the confusion between Glue Crawler (which populates the catalog) and Glue Data Catalog (which stores the metadata), causing candidates to select the crawler when the question asks where metadata is stored.

How to eliminate wrong answers

Option A is wrong because AWS Glue Crawler is a component that scans data sources and infers schemas to populate the Data Catalog, but it does not store or manage the metadata itself. Option B is wrong because AWS Glue Studio is a visual authoring tool for building ETL jobs, not a metadata repository. Option D is wrong because AWS Glue ETL refers to the job execution engine that transforms data, not the catalog that stores metadata.

185
MCQmedium

A data scientist needs to split a dataset into training, validation, and test sets. The dataset has a categorical target variable with imbalanced class distribution. Which splitting technique ensures that each subset has a similar proportion of each class?

A.K-fold cross-validation split
B.Chronological split
C.Stratified split
D.Random split
AnswerC

Stratified splitting preserves the target class proportions within each training, validation, and test subset by sampling per class rather than randomly across the whole dataset. This directly satisfies the stem's imbalanced categorical target constraint, preventing minority classes from being underrepresented or absent in any split.

Why this answer

Stratified splitting preserves the original class proportions in each subset (training, validation, test) by sampling each class independently. This is critical for imbalanced datasets to avoid skewed distributions that could bias model evaluation or training.

Exam trap

AWS often tests the distinction between data splitting techniques and model evaluation methods, so the trap here is that candidates confuse k-fold cross-validation (a validation strategy) with a static split technique, leading them to select option A.

How to eliminate wrong answers

Option A is wrong because k-fold cross-validation is a resampling technique for model evaluation, not a method for creating a single static split into training, validation, and test sets. Option B is wrong because chronological split orders data by time, which is irrelevant for a categorical target with imbalanced classes and does not guarantee proportional class representation. Option D is wrong because random split does not account for class distribution; with imbalanced data, it can produce subsets with significantly different class proportions, especially for rare classes.

186
MCQmedium

A data scientist is preparing a dataset for a binary classification model. The dataset has 10,000 samples, but the positive class represents only 2% of the data. The data scientist needs to train a model that will be evaluated on a hold-out test set that preserves the original class distribution. Which data preparation strategy is MOST appropriate?

A.Undersample the majority class in the training set to match the minority class size.
B.Oversample the minority class in the training set using SMOTE, and keep the test set as is.
C.Randomly oversample the minority class in the entire dataset and then split.
D.Apply SMOTE to the entire dataset before splitting into training and test sets.
AnswerB

SMOTE on the training set only addresses class imbalance during training; the test set preserves the original distribution for a realistic assessment.

Why this answer

With a 2% positive class, the training set needs class balancing to help the model learn the minority signal, but the hold-out test set must reflect the true 2% prevalence so evaluation metrics (precision, recall, PR-AUC) are realistic. Applying SMOTE only to the training set and leaving the test set untouched satisfies both requirements.

Exam trap

MLA-C01 often tests data leakage from resampling before the train/test split — candidates pick 'apply SMOTE to the whole dataset' thinking more data is always better.

How to eliminate wrong answers

Option A is wrong because undersampling the majority class to 200 samples discards ~98% of legitimate negative examples, losing information and biasing the model. Option C is wrong because oversampling before splitting leaks synthetic minority samples into the test set and destroys the original class distribution the question requires preserving. Option D is wrong for the same reason — applying SMOTE to the entire dataset before splitting causes data leakage (synthetic neighbors of test points appear in training) and inflates test performance, and it also violates the requirement to preserve the original distribution in the hold-out set.

187
MCQmedium

A company uses SageMaker Pipelines to automate model retraining. The pipeline runs daily but sometimes fails due to data quality issues. What is the best design to handle this?

A.Add a data quality check step with Conditional to skip training if data fails.
B.Use SageMaker Debugger to monitor training.
C.Use SageMaker Model Registry to track model versions.
D.Increase the instance size for the training step.
AnswerA

A quality-check step followed by a Condition step evaluates the data against thresholds and branches, skipping training when checks fail. This satisfies the stem's daily-failure scenario by preventing wasted training runs on poor data while still allowing the pipeline to complete.

Why this answer

SageMaker Pipelines supports a data quality check step that can be integrated with a ConditionStep. If the data quality check fails, the ConditionStep can skip the training step entirely, preventing the pipeline from failing due to bad data. This design ensures the pipeline completes successfully (or exits gracefully) without wasting compute resources on training with invalid data.

Exam trap

The trap here is that candidates may confuse monitoring tools (Debugger) or model management (Model Registry) with pipeline orchestration and conditional logic, failing to recognize that a ConditionStep is the correct mechanism to gate execution based on data quality.

How to eliminate wrong answers

Option B is wrong because SageMaker Debugger is designed to monitor training jobs for issues like overfitting, vanishing gradients, or hardware bottlenecks, not to prevent pipeline failures caused by data quality issues before training starts. Option C is wrong because SageMaker Model Registry is used for cataloging, versioning, and approving model artifacts, not for handling data quality checks or pipeline failure prevention. Option D is wrong because increasing the instance size for the training step addresses performance or memory constraints, not data quality issues; it would not prevent the pipeline from failing if the input data is invalid.

188
MCQmedium

A data engineer is preparing a dataset for a SageMaker training job. The dataset contains a timestamp column and is stored in Amazon S3 as CSV files. The engineer needs to ensure that the training job reads the data efficiently and that the data is partitioned by date to improve query performance in Amazon Athena. Which action should the engineer take?

A.Convert the CSV files to Parquet format and store them in a single prefix without partitioning.
B.Convert the CSV files to Parquet format and partition them by date in S3 using a Hive-style partition scheme (e.g., year=2023/month=01/day=01/).
C.Keep the CSV format but partition the files by date using a flat prefix structure (e.g., 2023-01-01/).
D.Use AWS Glue to crawl the CSV files and create a table in the AWS Glue Data Catalog, then use Athena to query the data directly without converting or partitioning.
AnswerB

Converting to Parquet improves compression and columnar read performance, which benefits SageMaker training jobs. Partitioning by date using Hive-style prefixes allows Athena to prune partitions and reduce query scan, improving performance. This meets both efficiency and query performance requirements.

Why this answer

Converting CSV to Parquet reduces storage size and enables columnar reads, which speeds up SageMaker training jobs. Partitioning by date with Hive-style prefixes allows Athena to prune partitions, reducing query cost and latency. Together, these actions satisfy both the training efficiency and Athena query performance requirements.

Exam trap

The trap here is thinking that simply cataloging CSV data or partitioning without converting format is sufficient for both SageMaker efficiency and Athena performance.

189
MCQmedium

A company wants to use SageMaker Autopilot to automatically build a binary classification model. Which output does Autopilot provide to help understand model decisions?

A.A leaderboard of models with only accuracy metrics
B.An explainability report with feature importance
C.A confusion matrix for each candidate model
D.A SHAP values summary plot for each trial
AnswerB

Autopilot generates an explainability report quantifying each feature's contribution to predictions, typically using SHAP values. This reveals which input variables drove the binary classification decisions, giving the transparency the company requires without manual model interrogation.

Why this answer

SageMaker Autopilot automatically generates an explainability report that includes feature importance, which helps users understand which features influenced the model's predictions. This report is produced for the best model and provides insights into model decisions, aligning with the requirement to understand model decisions. Autopilot's explainability report is a built-in feature that leverages SHAP values to compute feature attributions, but it is presented as a report, not just a raw plot.

Exam trap

MLA-C01 often tests the misconception that Autopilot automatically provides detailed diagnostic outputs like confusion matrices or SHAP plots for every candidate model, when in fact it focuses on a leaderboard and an explainability report for the best model.

How to eliminate wrong answers

Option A is wrong because Autopilot's leaderboard includes multiple metrics (e.g., accuracy, F1, AUC) for each candidate model, not only accuracy. Option C is wrong because while a confusion matrix can be generated for a model, Autopilot does not automatically provide a confusion matrix for each candidate model as a standard output; it focuses on the leaderboard and explainability report. Option D is wrong because Autopilot does not provide a SHAP values summary plot for each trial; it provides an explainability report for the best model, which includes feature importance based on SHAP, but not a separate plot for every trial.

190
MCQmedium

A machine learning engineer is using AWS Glue ETL to transform a large dataset stored in Amazon S3. The transformation involves joining two tables on a high-cardinality column and aggregating results. The job is running slowly and the engineer needs to improve performance. Which optimization technique should the engineer apply?

A.Increase the number of workers and memory allocation
B.Use bucketing on the join key for both tables
C.Use Spark SQL instead of the DynamicFrame API
D.Convert the job to use Python Shell
AnswerB

Bucketing partitions both tables by hash of the join key into a fixed number of buckets, so matching keys land in the same bucket. This lets Glue process each bucket pair independently, avoiding a full shuffle of the high-cardinality column and cutting join and aggregation time.

Why this answer

Bucketing on the join key for both tables co-locates rows with the same key into the same bucket files, which allows Spark to perform a bucketed join that avoids a full shuffle of the large dataset. This dramatically reduces network I/O and speeds up joins on high-cardinality columns.

Exam trap

MLA-C01 often tests the misconception that throwing more workers at a slow Glue job solves performance problems, when the real issue is shuffle-heavy joins that require bucketing or partition pruning.

How to eliminate wrong answers

Option A is wrong because simply adding more workers and memory increases cost and may not address the underlying shuffle bottleneck caused by the high-cardinality join. Option C is wrong because switching from DynamicFrame to Spark SQL does not change the physical join strategy and will not resolve shuffle-heavy joins. Option D is wrong because Python Shell jobs are single-node and cannot handle large distributed datasets, making performance far worse.

191
MCQmedium

A data scientist runs this pipeline but the Train step fails with "ResourceLimitExceeded". What is the most likely cause?

A.The account has a limit of 0 for ml.p3.2xlarge instances.
B.The volume size is too small for training.
C.The Preprocess step did not complete successfully.
D.The training image is not accessible.
AnswerA

The Train step requests an ml.p3.2xlarge instance, but the account's service quota for that instance type is zero, so Amazon SageMaker cannot provision it and returns ResourceLimitExceeded. Raising the quota for ml.p3.2xlarge compute, or switching to an instance type with available capacity, resolves the failure.

Why this answer

The 'ResourceLimitExceeded' error indicates that the requested instance type (ml.p3.2xlarge) exceeds the account's service quota for that specific instance family. In AWS SageMaker, each account has a default limit of 0 for certain GPU instance types like ml.p3.2xlarge unless a quota increase has been requested and approved. This error occurs at the Train step because SageMaker attempts to launch the training job with an instance type that is not allowed by the current quota.

Exam trap

AWS often tests the distinction between resource limits (quotas) and other failure modes; the trap here is that candidates may confuse 'ResourceLimitExceeded' with a generic 'insufficient capacity' error, but the error specifically refers to account-level service quotas, not AWS resource availability.

How to eliminate wrong answers

Option B is wrong because volume size limits (e.g., EBS volume size) do not cause a 'ResourceLimitExceeded' error; they would result in an 'InsufficientVolumeCapacity' or 'VolumeLimitExceeded' error. Option C is wrong because if the Preprocess step had failed, the pipeline would stop at that step and the Train step would not be attempted, so the error would be a different one (e.g., 'StepFailure'). Option D is wrong because an inaccessible training image would produce an 'ImageNotFoundException' or 'AccessDeniedException', not a 'ResourceLimitExceeded' error.

192
MCQmedium

A data scientist has trained a model that achieves 95% accuracy on the training set but only 70% on the test set. Which of the following is the most likely cause?

A.Data leakage
B.Overfitting
C.Convergence to local minimum
D.Underfitting
AnswerB

Overfitting occurs when the model memorises training noise and patterns, producing high training accuracy but poor generalisation to unseen data. The 95% versus 70% gap directly reflects this memorisation, satisfying the stem's symptom of a large train-test performance disparity.

Why this answer

A large gap between high training accuracy (95%) and much lower test accuracy (70%) is the classic signature of overfitting: the model has memorized the training data, including its noise, and fails to generalize to unseen data. The model has learned patterns specific to the training set rather than the underlying distribution. Regularization, more data, or a simpler model would typically help.

Exam trap

The trap is that candidates confuse overfitting with data leakage — both can produce misleadingly high training metrics, but only overfitting shows the sharp train-high/test-low gap, while leakage typically inflates test performance too.

How to eliminate wrong answers

Option A is wrong because data leakage occurs when information from outside the training set (e.g., the target variable) inadvertently leaks into training features, which usually inflates both training and test performance — it does not produce the train-high/test-low pattern seen here. Option C is wrong because convergence to a local minimum is an optimization issue that would typically result in poor performance on both training and test sets, not a large train-test gap. Option D is wrong because underfitting produces low accuracy on both training and test sets (e.g., 70% train and 68% test) — the model is too simple to capture the pattern, which is the opposite of the scenario described.

193
MCQhard

A company deploys a model using SageMaker and enables data capture for monitoring. After a week, they notice that the captured data is not being written to the specified S3 bucket. The endpoint is running and invocations are successful. What is the most likely cause?

A.The IAM role used for the endpoint does not have s3:PutObject permission for the capture bucket.
B.The capture bucket is in a different region.
C.The endpoint is using a multi-model endpoint which does not support data capture.
D.The DataCaptureConfig parameter in the endpoint configuration is missing the "CaptureOptions" field.
AnswerA

Data capture writes inference request and response records to S3 using the endpoint's execution role. Successful invocations confirm the model works, so the missing writes point to the role lacking s3:PutObject on the capture bucket, which silently blocks capture delivery.

Why this answer

The most likely cause is that the IAM role associated with the SageMaker endpoint lacks the `s3:PutObject` permission for the target S3 bucket. Without this permission, the endpoint cannot write the captured inference data to S3, even though invocations succeed because the model itself does not require S3 write access to serve predictions.

Exam trap

The trap here is that candidates often assume data capture fails due to endpoint misconfiguration (like missing CaptureOptions) or regional restrictions, when in fact the root cause is almost always an IAM permissions issue with the S3 bucket.

How to eliminate wrong answers

Option B is wrong because SageMaker data capture supports cross-region S3 buckets; the bucket can be in a different region as long as the endpoint has network access and proper permissions. Option C is wrong because multi-model endpoints fully support data capture; there is no restriction that prevents capture on multi-model endpoints. Option D is wrong because the `CaptureOptions` field is optional; if omitted, SageMaker uses default capture options (e.g., capturing both input and output).

The missing field would not prevent data from being written to S3.

194
MCQhard

Refer to the exhibit. An IAM policy is attached to a user to allow invoking a SageMaker endpoint. A developer tries to call the endpoint from a laptop with IP 203.0.113.5 and receives an access denied error. What is the most likely reason?

A.The resource ARN is incorrect.
B.The condition restricts the IP address to the 10.0.0.0/8 range.
C.The user does not have permission to assume the SageMaker role.
D.The policy does not include access to the API action.
AnswerB

The policy's IpAddress condition permits only the 10.0.0.0/8 private range, so the request from 203.0.113.5 fails the condition and is denied. The endpoint itself is reachable; the source IP simply falls outside the allowed range.

Why this answer

The policy includes a condition that restricts the source IP address to the 10.0.0.0/8 private range. The developer's laptop has a public IP of 203.0.113.5, which does not fall within that range, so the condition fails and access is denied. This is the most likely reason for the error because the condition explicitly blocks requests from outside the specified private network.

Exam trap

The trap here is that candidates may overlook the condition element and assume the error is due to a missing action or incorrect ARN, when in fact the condition is the restrictive factor that denies access based on the source IP.

How to eliminate wrong answers

Option A is wrong because if the resource ARN were incorrect, the error would typically indicate an invalid ARN or a mismatch, not an access denied due to IP restriction; the ARN in the policy appears correctly formatted for a SageMaker endpoint. Option C is wrong because the policy does not involve assuming a role; it directly grants invoke permissions to the user, and the error is not related to role assumption. Option D is wrong because the policy explicitly includes the 'sagemaker:InvokeEndpoint' action, so the user does have permission to the API action; the denial is caused by the condition, not a missing action.

195
MCQmedium

A data scientist is working with a dataset containing a categorical feature 'country' with 200 unique values. They plan to use a linear regression model. Which encoding method is most suitable to avoid the dummy variable trap while maintaining interpretability?

A.Binary encoding
B.Target encoding
C.Label encoding
D.One-hot encoding with dropping the first category
AnswerD

One-hot encoding creates a binary column per country, and dropping one category avoids perfect multicollinearity — the dummy variable trap — which would make linear regression coefficients unstable. Retaining 199 columns preserves interpretability, unlike target or ordinal encoding, which impose false ordering on nominal country values.

Why this answer

One-hot encoding with dropping the first category is the most suitable method because it creates binary columns for each category except one, avoiding perfect multicollinearity (the dummy variable trap) while preserving interpretability. Each coefficient directly represents the effect of that category relative to the dropped reference category, which is intuitive for linear regression.

Exam trap

Candidates often mistakenly think one-hot encoding must include all categories, overlooking the dummy variable trap, or assume label encoding is acceptable for nominal data in linear models due to its simplicity.

How to eliminate wrong answers

Option A is wrong because binary encoding represents categories as binary numbers, which introduces arbitrary ordinal relationships and makes coefficients uninterpretable for linear regression. Option B is wrong because target encoding replaces categories with the mean of the target variable, which can cause data leakage and overfitting, and the encoded values lose direct categorical interpretability. Option C is wrong because label encoding assigns arbitrary integer labels (e.g., 1, 2, 3) that imply an ordinal relationship, which is inappropriate for a nominal categorical feature in linear regression.

196
Multi-Selectmedium

An ML engineer is setting up monitoring for a SageMaker endpoint. Which THREE metrics should be monitored to detect performance issues? (Select THREE.)

Select 3 answers
A.Model latency
B.Invocations per second
C.CPUUtilization
D.MemoryUtilization
E.DiskWriteBytes
AnswersA, C, D

Model latency measures the time the endpoint takes to return a prediction, so rising values directly reveal inference performance degradation. It satisfies the requirement to detect endpoint performance issues, exposing slowdowns before they breach application response-time expectations.

Why this answer

Model latency (A) is correct because it directly measures the time the endpoint takes to respond to inference requests, and rising latency is a primary indicator of performance degradation. CPUUtilization (C) is correct because high CPU usage on the hosting instances can throttle inference processing and cause slower responses or timeouts. MemoryUtilization (D) is correct because excessive memory consumption can lead to swapping, out-of-memory errors, or container restarts that degrade endpoint performance.

Invocations per second (B) reflects traffic volume rather than endpoint health, so it does not by itself indicate a performance issue. DiskWriteBytes (E) is a storage I/O metric that is largely irrelevant to the in-memory inference workload of a SageMaker endpoint.

Exam trap

The trap here is that candidates often confuse throughput metrics (like invocations per second) with performance health indicators, but the question specifically asks for metrics that detect performance issues, not just operational statistics.

197
MCQmedium

A SageMaker endpoint is logging an error when processing inference requests that require database access. What is the most likely cause?

A.Data capture is not enabled
B.The endpoint instance type is too small
C.The model is not compatible with the instance
D.The endpoint lacks a VPC configuration with proper security groups
AnswerD

SageMaker endpoints run outside your VPC by default, so they cannot reach private database resources. Attaching a VPC configuration with appropriate security groups and subnets grants the endpoint network access, directly resolving the database connectivity failure during inference.

Why this answer

When a SageMaker endpoint needs to access an external database (e.g., Amazon RDS) during inference, it must be launched within a VPC that has proper security group and subnet configurations. If the endpoint is not in a VPC or the security groups do not allow outbound traffic to the database, the endpoint will be unable to connect, resulting in errors. The other options are less likely: data capture is for logging requests (A), instance size affects performance not connectivity (B), model compatibility affects deployment but not specifically database access (C).

Exam trap

Some candidates might think that database connectivity issues are due to endpoint instance size or missing data capture configuration. However, the key for accessing external resources from SageMaker is VPC configuration.

How to eliminate wrong answers

Option A is wrong because data capture is a feature for logging inference request/response payloads, not for enabling network connectivity to a database. Option B is wrong because an undersized instance type would cause performance issues like latency or out-of-memory errors, not a network connectivity failure to a database. Option C is wrong because model-instance compatibility issues typically manifest as runtime errors (e.g., 'CUDA error' or 'model loading failed'), not as a failure to establish a database connection.

198
Multi-Selecteasy

A data engineer is using AWS Glue to prepare a dataset for ML. The engineer wants to split the dataset into training and testing sets while preserving the distribution of the target variable. Which TWO methods achieve this goal? (Select TWO)

Select 2 answers
A.Use Amazon Athena to create views with random sampling
B.Use the `train_test_split` function from scikit-learn in a SageMaker notebook
C.Use AWS Glue's built-in random split transform
D.Use a custom Spark script with stratified sampling
E.Use Amazon SageMaker's built-in SplitType parameter in a Processing Job
AnswersB, D

The stratify parameter maintains class proportions.

Why this answer

The `train_test_split` function from scikit-learn supports the `stratify` parameter, which preserves the distribution of the target variable when splitting a dataset into training and testing sets. This is a standard, reliable method for stratified splitting in Python-based ML workflows, and it can be used directly in a SageMaker notebook.

Exam trap

The trap here is that candidates often confuse random splitting (which is available in many tools like Glue and Athena) with stratified splitting, assuming that any 'random' operation preserves distribution, but only stratified methods explicitly maintain class proportions.

199
Multi-Selectmedium

A data scientist is using SageMaker Data Wrangler to prepare a dataset for a binary classification model. The dataset contains a mix of numerical and categorical features. The scientist wants to perform feature engineering to improve model performance. Which TWO actions are appropriate for handling categorical features in Data Wrangler? (Choose two.)

Select 2 answers
A.Apply one-hot encoding to categorical features with low cardinality.
B.Apply target encoding to categorical features with high cardinality.
C.Apply PCA to categorical features for dimensionality reduction.
D.Apply log transformation to categorical features.
E.Apply min-max scaling to categorical features.
AnswersA, B

One-hot encoding is a standard technique for converting categorical variables into a numerical format suitable for many machine learning algorithms. It creates binary columns for each category. For low-cardinality features (few unique values), this is efficient and avoids the curse of dimensionality. Data Wrangler provides a 'One-Hot Encoding' transform that can be applied directly.

Why this answer

Categorical features need to be encoded into numerical representations before training. One-hot encoding is effective for low-cardinality features, creating binary indicators. Target encoding is suitable for high-cardinality features, replacing categories with target statistics and reducing dimensionality.

Both are available as transforms in SageMaker Data Wrangler. Scaling, PCA, and log transformations are designed for numerical features and are not applicable to categorical data.

Exam trap

The trap here is assuming that numerical transforms like scaling or PCA can be applied to categorical features without encoding; they cannot and will cause errors.

200
MCQeasy

Which SageMaker built-in algorithm should be used for forecasting time series data with seasonal patterns?

A.IP Insights
B.BlazingText
C.DeepAR
D.Factorization Machines
AnswerC

DeepAR is a supervised recurrent neural network algorithm built into SageMaker that learns from multiple related time series and models seasonality and uncertainty, producing probabilistic forecasts, which matches the requirement for forecasting data exhibiting seasonal patterns.

Why this answer

DeepAR is a SageMaker built-in algorithm specifically designed for time series forecasting using recurrent neural networks (RNNs). It excels at handling complex seasonal patterns and can incorporate additional features, making it ideal for forecasting time series data with seasonality. It is the only built-in algorithm among the options that is purpose-built for time series forecasting.

Exam trap

MLA-C01 often tests the confusion between algorithms for different tasks, where candidates might choose BlazingText for time series because it sounds like it handles sequences, but it is actually for text.

How to eliminate wrong answers

Option A is wrong because IP Insights is used for identifying anomalous IP address usage patterns, not for time series forecasting. Option B is wrong because BlazingText is a natural language processing algorithm for text classification and word embeddings, unrelated to time series. Option D is wrong because Factorization Machines are used for recommendation systems and classification tasks, not for sequential time series forecasting.

201
MCQmedium

A data scientist is using Amazon SageMaker Data Wrangler to prepare a dataset for classification. They want to detect potential bias in the data before training. Which SageMaker service should they use in conjunction with Data Wrangler to detect bias?

A.Amazon SageMaker Clarify
B.Amazon SageMaker Debugger
C.Amazon SageMaker Pipelines
D.Amazon SageMaker Model Monitor
AnswerA

SageMaker Clarify computes bias metrics such as class imbalance and disparate impact on datasets and trained models. Running it alongside Data Wrangler lets the team quantify pre-training bias in the prepared features, which is exactly the detection the stem requires.

Why this answer

Amazon SageMaker Clarify is designed to detect bias in datasets and models. It integrates with Data Wrangler to provide bias analysis during the data preparation phase.

202
MCQhard

A retail company uses a SageMaker Model Monitor data quality monitor on a real-time endpoint. The monitor's baseline was generated from a training dataset in which the "promo_code" feature was often null. In production the feature is now populated for nearly every record, and the monitor reports violations even though model accuracy has not degraded. The team wants the monitor to stop flagging this expected change without disabling monitoring entirely. What should they do?

A.Regenerate the baseline statistics and constraints from a recent, representative production dataset and update the monitoring schedule to use the new baseline.
B.Delete the monitoring schedule and create a new one with a longer monitoring interval so that fewer violations accumulate over time.
C.Edit the constraints JSON file in Amazon S3 to remove the promo_code entry, then leave the schedule unchanged.
D.Enable explainability monitoring on the schedule so that feature attribution drift accounts for the change in promo_code.
AnswerA

Model Monitor compares incoming data against the statistics and constraints stored in the baseline. Because the production distribution of promo_code legitimately differs from the training data, the old constraints encode a stale expectation. Rebuilding the baseline from recent representative production data realigns the constraint thresholds with current behavior, so violations stop without turning monitoring off.

Why this answer

Data quality monitors flag when observed statistics fall outside the constraints derived from the baseline. When a feature's production distribution legitimately diverges from training, the correct fix is to re-establish the baseline from representative recent data so constraints reflect the new normal, rather than muting or deleting the check.

Exam trap

The trap here is treating monitor violations as a monitoring-configuration bug to be silenced, rather than as a stale-baseline problem to be corrected with fresh representative data.

203
MCQeasy

A machine learning engineer trains a binary classifier and obtains an accuracy of 95% on the test set. The dataset is imbalanced with 95% positive class. What is the most important metric to evaluate the model's performance?

A.R-squared
B.F1 score
C.Accuracy
D.RMSE
AnswerB

With 95% positives, a model predicting only the majority class scores 95% accuracy yet detects no negatives. F1 score balances precision and recall on the minority class, exposing that failure, whereas accuracy and raw recall remain misleadingly high.

Why this answer

With a 95% positive class imbalance, a model that always predicts the majority class achieves 95% accuracy, making accuracy a misleading metric. The F1 score (option B) is the harmonic mean of precision and recall, providing a balanced evaluation of the model's ability to correctly identify the minority class while penalizing false positives and false negatives. This makes it the most important metric for imbalanced binary classification.

Exam trap

The trap here is that candidates see 95% accuracy and assume the model is performing well, failing to recognize that accuracy is inflated by the class imbalance and that the F1 score is the correct metric to evaluate minority class performance.

How to eliminate wrong answers

Option A is wrong because R-squared is a metric for regression models, measuring the proportion of variance explained by the independent variables, and is not applicable to binary classification. Option C is wrong because accuracy is misleading in imbalanced datasets; a model that predicts only the majority class (positive) would achieve 95% accuracy without learning any meaningful patterns, so it does not reflect true performance on the minority class. Option D is wrong because RMSE (Root Mean Square Error) is a regression metric that measures the square root of the average squared differences between predicted and actual values, and it is not designed for evaluating binary classification outcomes.

204
MCQhard

A financial services company must ensure that a SageMaker model deployed to a real-time endpoint only produces predictions consistent with a fairness constraint on a protected attribute, and that any violation is detected within minutes and triggers an alert to the compliance team. The model is already deployed and monitored for data quality. Which approach should the machine learning engineer implement?

A.Enable SageMaker Model Monitor with a bias drift baseline created by SageMaker Clarify, schedule the monitoring job to run every few minutes, and configure CloudWatch alarms on the bias metrics.
B.Use SageMaker Model Registry to gate the model on a fairness condition and configure EventBridge to notify the compliance team when the model version changes.
C.Attach a Clarify explainability job to the endpoint and configure a CloudWatch alarm on the endpoint's ModelLatency metric.
D.Create a SageMaker Model Monitor data quality job with a custom metric that flags predictions where the protected attribute equals a specific value, and alarm on that metric.
AnswerA

Model Monitor supports bias drift monitoring using a Clarify-generated baseline, which defines the fairness constraint and computes bias metrics on captured data. Scheduling the job at a short interval detects violations within minutes, and emitting the metrics to CloudWatch lets alarms notify the compliance team. This directly ties the fairness constraint to automated detection and alerting.

Why this answer

Bias drift monitoring in Model Monitor compares live predictions against a Clarify-generated bias baseline that encodes the fairness constraint, computing metrics on captured data. Running the job frequently yields detection within minutes, and publishing those metrics to CloudWatch enables alarms that notify compliance. Explainability, data quality, and registry gating do not evaluate outcome fairness on a protected attribute.

Exam trap

The trap here is confusing Clarify explainability or data quality statistics with bias monitoring, when only a Clarify bias baseline evaluated by Model Monitor measures fairness constraints on a protected attribute over time.

205
Multi-Selectmedium

A data science team detects that a deployed model's prediction accuracy is degrading over time due to concept drift. They need to implement a retraining strategy. Which THREE actions are recommended best practices for handling concept drift?

Select 3 answers
A.Automatically roll back to a previous model version upon drift detection.
B.Monitor prediction quality using ground truth labels when available.
C.Retrain the model on a fixed schedule regardless of performance.
D.Incrementally update the model with new data using SageMaker Pipelines.
E.Use SageMaker Model Monitor to detect drift and trigger retraining.
AnswersB, D, E

Monitoring prediction quality against ground truth labels detects actual accuracy degradation, not just input drift. This satisfies the stem's concept drift scenario by measuring whether the model's outputs still match reality as the relationship between features and target changes.

Why this answer

Option B is correct because continuously monitoring prediction quality against ground truth labels (e.g., via SageMaker Model Monitor's ModelQuality baseline and CloudWatch metrics) is the only way to quantify real accuracy degradation and confirm concept drift rather than mere data drift. Option D is correct because SageMaker Pipelines can orchestrate incremental retraining workflows that incorporate newly labeled data, allowing the model to adapt to the changed concept without rebuilding the entire pipeline manually. Option E is correct because SageMaker Model Monitor detects drift (data drift, model quality drift) and can emit CloudWatch alarms that trigger a retraining pipeline, closing the detect-to-retrain loop automatically.

Option A is not recommended as a primary strategy because rolling back to an old model does not address the underlying concept drift and the previous version will also degrade on the new distribution. Option C is not recommended because fixed-schedule retraining ignores actual performance signals, wasting compute when drift is absent and lagging when drift is rapid.

Exam trap

The trap here is that candidates may confuse 'detecting drift' with 'responding to drift' and incorrectly choose automatic rollback (Option A) as a best practice, when in reality rollback is a risky operation that should be evaluated carefully, not automated blindly.

206
MCQhard

A company's ML pipeline runs in multiple AWS accounts (dev, test, prod). They want to enforce that only approved models from a central Model Registry can be deployed to the production account. Which combination of services is MOST appropriate to implement this governance?

A.AWS Config, Amazon GuardDuty, and AWS Security Hub.
B.Amazon API Gateway, AWS Step Functions, and Amazon DynamoDB.
C.AWS Service Catalog, AWS KMS, and AWS CloudTrail.
D.AWS Organizations with SCPs, AWS CodePipeline with cross-account actions, and SageMaker Model Registry with approval status.
E.AWS CloudFormation StackSets, Amazon EventBridge, and AWS Lambda.
AnswerD

SCPs in AWS Organizations enforce that only approved registry models deploy to production, while CodePipeline cross-account actions move artefacts between accounts. The Model Registry approval status gates deployment, satisfying the central governance constraint across dev, test and prod accounts.

Why this answer

It combines AWS Organizations with SCPs to enforce cross-account deployment policies, AWS CodePipeline with cross-account actions to orchestrate the pipeline across dev/test/prod accounts, and SageMaker Model Registry with approval status to gate deployments to only approved models. This ensures that only models with an 'Approved' status in the central registry can be deployed to the production account, meeting the governance requirement.

Exam trap

The trap here is that candidates may choose monitoring-focused options like A or C, mistakenly thinking that detecting non-approved deployments is sufficient, when the question explicitly requires enforcement (prevention), which demands a combination of policy-based controls (SCPs) and approval-gated pipelines (CodePipeline + Model Registry).

How to eliminate wrong answers

Option A is wrong because AWS Config, GuardDuty, and Security Hub are monitoring and security services that detect misconfigurations and threats but cannot enforce approval-based gating on model deployments. Option B is wrong because API Gateway, Step Functions, and DynamoDB are used for building serverless workflows and APIs, not for cross-account deployment governance or model approval enforcement. Option C is wrong because AWS Service Catalog manages approved IT service portfolios, KMS handles encryption keys, and CloudTrail logs API activity—none of these services directly enforce that only approved SageMaker models are deployed to production.

Option E is wrong because CloudFormation StackSets deploy infrastructure across accounts, EventBridge routes events, and Lambda runs code—they lack native integration with SageMaker Model Registry approval status to gate deployments.

207
MCQmedium

A data scientist needs to ingest streaming customer clickstream data from a website into an S3 data lake for ML training. The data must be delivered within 1 minute of ingestion, and JSON records must be converted to Parquet. Which AWS service combination should be used?

A.Amazon S3 Transfer Acceleration with direct PUT requests by clients
B.Amazon Kinesis Data Streams (KDS) with a Lambda consumer that writes JSON to S3
C.AWS Glue ETL job triggered every minute to pull from Kinesis Data Streams and write Parquet to S3
D.Amazon Kinesis Data Firehose with a 60-second buffer and Parquet conversion enabled
AnswerD

Firehose buffers incoming records and delivers them within the configured 60-second window, satisfying the one-minute latency constraint, while its built-in record format conversion transforms JSON into Parquet using a Glue table schema before writing to the S3 data lake.

Why this answer

Amazon Kinesis Data Firehose can buffer incoming data and deliver to S3 with a 60-second buffer window, and it supports converting JSON to Parquet. KDS alone does not convert to Parquet. Glue ETL can do the conversion but adds latency.

Lambda with S3 trigger is not streaming-oriented.

208
MCQeasy

A machine learning engineer wants to reduce training costs by using excess EC2 capacity. Which instance purchasing option should they choose for SageMaker training jobs?

A.Reserved Instances
B.On-Demand Instances
C.Dedicated Instances
D.Spot Instances
AnswerD

Spot Instances use spare EC2 capacity at steep discounts, directly satisfying the cost-reduction constraint. SageMaker training jobs support managed spot training with checkpointing, so interrupted instances resume rather than restart. This suits fault-tolerant training workloads, unlike On-Demand or Reserved capacity, which bill for continuity regardless of interruption tolerance.

Why this answer

Spot Instances use AWS's spare EC2 capacity and are discounted up to 90% versus On-Demand, making them the correct choice when the workload is interruptible. SageMaker training jobs support Managed Spot Training, which automatically checkpoints and resumes jobs when capacity is reclaimed, so the engineer gets the cost savings without losing training progress. Reserved Instances and Dedicated Instances do not offer the same spot-level discount for interruptible training.

Exam trap

MLA-C01 often tests the confusion between Reserved Instances (commitment discount for steady workloads) and Spot Instances (discount for interruptible workloads) — candidates pick Reserved because it sounds 'cheaper' without considering the interruptibility requirement.

How to eliminate wrong answers

Option A is wrong because Reserved Instances require a 1- or 3-year commitment and are designed for steady-state, always-on workloads — they provide a billing discount but no mechanism for using excess capacity, and they are not the cost-optimal choice for interruptible training. Option B is wrong because On-Demand Instances are the most expensive option and offer no discount for tolerating interruptions, defeating the stated goal of reducing training costs. Option C is wrong because Dedicated Instances are physical-host-isolated instances priced at a premium for compliance/licensing reasons, not a cost-reduction mechanism for training.

209
Multi-Selecthard

A data scientist is training a large transformer model using SageMaker's model parallelism library. The training job is failing with an out-of-memory (OOM) error. Which two actions can help resolve the OOM error? (Choose two.)

Select 2 answers
A.Reduce the sequence length
B.Enable activation checkpointing
C.Increase the batch size per GPU
D.Switch to a smaller instance type
E.Decrease the pipeline parallelism degree
AnswersA, B

Reducing sequence length shrinks activation memory, which scales linearly with token count, directly relieving the per-device memory pressure causing the OOM. Since model parallelism shards parameters but not activations, this addresses the constraint the stem identifies without altering the sharding configuration or requiring additional instances.

Why this answer

Option A is correct because reducing the sequence length directly lowers the activation memory footprint of a transformer, since attention and intermediate activations scale with sequence length, making it a standard remedy for OOM during SageMaker model-parallel training. Option B is correct because activation checkpointing (gradient checkpointing) recomputes activations during the backward pass instead of storing all of them, substantially reducing memory usage at the cost of some extra compute. Option C is incorrect because increasing the batch size per GPU raises memory consumption and would worsen the OOM.

Option D is incorrect because switching to a smaller instance type reduces available GPU memory, making OOM more likely. Option E is incorrect because decreasing the pipeline parallelism degree spreads the model across fewer stages, increasing per-GPU memory pressure rather than relieving it.

Exam trap

The trap here is that candidates may confuse pipeline parallelism with tensor parallelism, assuming decreasing pipeline degree reduces memory, when in fact it increases per-GPU memory load due to fewer stages.

210
MCQmedium

A data engineer is designing a feature engineering pipeline using Amazon SageMaker Feature Store. The team needs to support both real-time inference (millisecond latency) and batch training jobs that require access to historical feature values at specific points in time. Which configuration should the engineer choose?

A.Create a feature group with only an online store
B.Create separate feature groups — one for online and one for offline — and manage data synchronization manually
C.Create a feature group with both online and offline stores enabled
D.Store features only in the offline store and use a separate low-latency cache like ElastiCache
AnswerC

Enabling both stores gives the feature group a low-latency online store for millisecond real-time inference and an offline store retaining historical values with event times for point-in-time batch training retrieval. This single configuration satisfies both the latency and historical-access constraints stated in the stem.

Why this answer

Feature Store supports dual storage: an online store (low-latency, key-value) for real-time inference and an offline store (S3-backed, queryable) for batch processing and point-in-time queries.

211
MCQmedium

A data scientist wants to train a model on SageMaker using a custom PyTorch script, then register the best model in the SageMaker Model Registry. The training job is part of a SageMaker Pipeline. Which pipeline step should be used to register the model?

A.RegisterModelStep
B.CreateModelStep
C.TrainingStep
D.TransformStep
AnswerA

RegisterModelStep is the dedicated SageMaker Pipelines step that packages the trained model artefact and its metadata into a model package group in the Model Registry, chaining directly from the training step within the pipeline definition.

Why this answer

The `RegisterModelStep` is specifically designed to create a model resource and register it in the SageMaker Model Registry as part of a pipeline. It takes the training output (e.g., model artifacts from a `TrainingStep`) and packages it with the specified inference image and metadata, then creates a model package group version. This is the correct step for registering a model after training, as it directly integrates with the Model Registry for versioning and approval workflows.

Exam trap

The trap here is that candidates confuse `CreateModelStep` (which creates a deployable model resource) with `RegisterModelStep` (which creates a model package version in the registry), assuming both serve the same purpose of model registration.

How to eliminate wrong answers

Option B is wrong because `CreateModelStep` only creates a SageMaker model resource (for deployment or batch inference) but does not register it in the Model Registry; it lacks the versioning and metadata capabilities needed for registry management. Option C is wrong because `TrainingStep` is used to run a training job and produce model artifacts, but it has no built-in functionality to register the model into the Model Registry; registration requires a separate step. Option D is wrong because `TransformStep` is used for batch inference (transform jobs) on existing models, not for registering models into the registry.

212
Multi-Selecthard

A company wants to enable cross-account access to a SageMaker model endpoint. The model is in Account A, and Account B needs to invoke it. Which TWO steps are required? (Select TWO)

Select 2 answers
A.Attach a resource-based policy to the SageMaker model in Account A allowing access from Account B's IAM role
B.Export the model from Account A and re-deploy in Account B
C.Create an IAM role in Account B with permissions to invoke SageMaker endpoints
D.Configure VPC peering between the two accounts
E.Use a SageMaker notebook instance cross-account sharing
AnswersA, C

Resource policies grant cross-account permissions directly on the model.

Why this answer

SageMaker endpoints support resource-based policies that allow cross-account access. By attaching a resource-based policy to the model endpoint in Account A, you can grant the IAM role from Account B explicit permission to invoke the endpoint. This is the standard AWS mechanism for cross-account SageMaker endpoint invocation without needing to duplicate the model.

Exam trap

The trap here is that candidates often confuse network-level connectivity (VPC peering) with IAM-level authorization, or assume that cross-account access requires duplicating resources, when in fact SageMaker's resource-based policies provide a direct and secure solution.

213
MCQmedium

A team is building a recommendation system and wants to store and serve features for online and offline models. The features include user statistics (updated daily) and movie metadata (static). The team needs low-latency inference for real-time recommendations and wants to reuse features across multiple models. Which AWS service should the team use to store, manage, and serve these features?

A.Amazon DynamoDB with TTL.
B.AWS Glue Data Catalog.
C.SageMaker Feature Store.
D.Amazon S3 with AWS Lambda for serving.
AnswerC

SageMaker Feature Store provides an online store for low-latency real-time inference and an offline store for training, with feature groups reusable across models. It directly satisfies the daily-updated user statistics, static movie metadata and cross-model reuse requirements.

Why this answer

Amazon SageMaker Feature Store is purpose-built for storing, managing, and serving ML features with low-latency retrieval for online inference and batch serving for offline training. It supports feature reuse across multiple models by providing a centralized feature registry, consistent feature definitions, and both online (low-latency) and offline (S3-based) stores, which directly matches the team's requirements for real-time recommendations and cross-model reuse.

Exam trap

The trap here is that candidates often confuse a general-purpose database (DynamoDB) or a data catalog (Glue) with a purpose-built ML feature store, overlooking the need for feature-specific capabilities like online/offline consistency, feature versioning, and reuse across models.

How to eliminate wrong answers

Option A is wrong because Amazon DynamoDB with TTL is a key-value and document database that can store features but lacks built-in feature management capabilities such as feature versioning, point-in-time consistency across online/offline stores, and a feature registry; TTL only handles data expiration, not the orchestration needed for ML feature reuse. Option B is wrong because AWS Glue Data Catalog is a metadata repository for data assets (tables, schemas) and does not provide a low-latency online serving endpoint or feature-specific storage; it is used for data discovery and ETL, not for serving features in real-time inference. Option D is wrong because Amazon S3 with AWS Lambda for serving introduces high latency due to Lambda cold starts and S3 GET request overhead, making it unsuitable for low-latency real-time recommendations; additionally, it lacks feature store capabilities like consistent feature definitions, offline/online synchronization, and feature reuse across models.

214
MCQmedium

A machine learning engineer needs to split a time-series dataset for a forecasting model. The data spans 3 years of daily sales. Which splitting strategy should they use to avoid look-ahead bias?

A.k-fold cross-validation with shuffling
B.Random train-test split with 80/20 ratio
C.Stratified sampling based on sales volume
D.Walk-forward validation (time-series split)
AnswerD

Walk-forward validation trains on earlier data and tests on later, chronologically subsequent data, preserving temporal order. This prevents look-ahead bias because the model never sees future observations during training, unlike random k-fold splitting, which would leak future values into past windows.

Why this answer

Walk-forward validation (time-series split) is the correct strategy because it preserves the temporal order of the data, training on past observations and testing on future observations sequentially. This avoids look-ahead bias, where future information leaks into the training set, which would invalidate the forecasting model's performance metrics.

Exam trap

The trap here is that candidates often default to k-fold cross-validation or random splits because they are standard for non-temporal data, failing to recognize that time-series data requires strict temporal ordering to avoid look-ahead bias.

How to eliminate wrong answers

Option A is wrong because k-fold cross-validation with shuffling randomly reorders the data, breaking the temporal sequence and allowing future data to leak into training folds, introducing look-ahead bias. Option B is wrong because a random train-test split with an 80/20 ratio also shuffles the data, disregarding the time order and causing future sales data to appear in the training set. Option C is wrong because stratified sampling based on sales volume does not account for time dependency; it groups data by sales categories, which can mix past and future observations, leading to look-ahead bias.

215
MCQhard

A team is building a time-series forecasting model for daily sales data. They want to evaluate model performance using cross-validation while respecting the temporal order of the data. Which data splitting strategy should they use?

A.Walk-forward validation (time-series split)
B.Holdout with a random 80/20 split
C.Random k-fold cross-validation
D.Stratified sampling
AnswerA

Walk-forward validation, implemented as time-series split, trains on earlier observations and validates on later ones, preserving chronological order. Standard k-fold shuffles data and leaks future information into training, which would produce misleadingly optimistic forecasts for daily sales.

Why this answer

Walk-forward validation (time-series split) respects the temporal order by training on past data and validating on future data, expanding the training window forward in time. This mimics real forecasting conditions and prevents data leakage from future observations into the training set. It is the correct strategy for time-series cross-validation.

Exam trap

MLA-C01 often tests the misconception that standard k-fold or random holdout is acceptable for time-series — candidates overlook temporal leakage and pick random splitting strategies that violate the time order.

How to eliminate wrong answers

Option B is wrong because a random 80/20 holdout ignores temporal order and can train on future data to predict the past, causing leakage and overly optimistic performance estimates. Option C is wrong because random k-fold cross-validation shuffles data, breaking temporal dependencies and allowing future information to leak into training folds. Option D is wrong because stratified sampling preserves class proportions but is designed for classification, not time-series forecasting, and still ignores temporal order.

216
MCQhard

A team is using Amazon SageMaker Data Wrangler to prepare a large dataset. They need to detect potential bias in the data before training. Which capability of Data Wrangler should they use?

A.Integration with Amazon SageMaker Clarify for bias reports
B.Built-in transform for SMOTE oversampling
C.Use of Amazon Athena to query data for bias patterns
D.Export to Amazon SageMaker Feature Store
AnswerA

SageMaker Clarify runs bias detection against the prepared dataset, computing metrics such as class imbalance and disparate impact across facets. Data Wrangler surfaces these Clarify bias reports directly in its analysis view, letting the team identify imbalance before training rather than after model deployment.

Why this answer

Amazon SageMaker Data Wrangler integrates directly with Amazon SageMaker Clarify to detect bias in datasets. This integration allows you to run bias analysis on your data before training, generating reports that highlight potential imbalances or unfairness in features and target variables. It is the correct capability for the team's stated need.

Exam trap

The trap here is that candidates may confuse data preprocessing techniques (like SMOTE for oversampling) with bias detection, or assume that any AWS query service (like Athena) can perform bias analysis, when only SageMaker Clarify provides the dedicated bias detection and reporting capability integrated with Data Wrangler.

How to eliminate wrong answers

Option B is wrong because SMOTE (Synthetic Minority Over-sampling Technique) is a built-in transform for oversampling imbalanced data, not for detecting or reporting bias. Option C is wrong because Amazon Athena is a query service for analyzing data in S3 using SQL; it does not have built-in bias detection capabilities or integration with SageMaker Clarify for bias reports. Option D is wrong because exporting to Amazon SageMaker Feature Store is for storing and sharing features for reuse in training and inference, not for detecting bias in the data.

217
MCQmedium

Refer to the exhibit. A data engineer investigates why a SageMaker endpoint is returning errors. The endpoint configuration has been updated to point to a new model version. What is the MOST likely cause of the error?

A.The endpoint instance type is insufficient.
B.The container image for the new model is not compatible.
C.The IAM role does not have permission to invoke the endpoint.
D.The endpoint is still using the previous configuration.
E.The new model artifact is not properly uploaded to S3.
AnswerD

Updating an endpoint configuration does not automatically redeploy the live endpoint; the existing endpoint keeps serving the previous model version until the update is applied. The mismatch between the new configuration and the running endpoint therefore causes inference errors, making a stale configuration the most likely cause.

Why this answer

When a SageMaker endpoint configuration is updated to point to a new model version, the endpoint itself does not automatically switch to the new configuration unless a subsequent UpdateEndpoint or Deploy call is made. The endpoint continues serving the previous model until the update is explicitly applied, so the errors are most likely caused by the endpoint still using the old configuration.

Exam trap

AWS often tests the distinction between updating an endpoint configuration and actually deploying that configuration to the endpoint, leading candidates to assume that changing the configuration automatically updates the endpoint.

How to eliminate wrong answers

Option A is wrong because an insufficient instance type would typically cause resource exhaustion errors (e.g., OutOfMemory or CPU throttling), not a configuration mismatch error after a model version update. Option B is wrong because if the container image were incompatible, the model would fail to load or produce a different error (e.g., ModelError or ContainerLaunchFailure), not a generic endpoint error. Option C is wrong because the IAM role for invoking the endpoint is separate from the execution role; the invocation role is used by the client, and if it lacked permission, the error would be an AccessDeniedException, not a configuration-related error.

Option E is wrong because if the new model artifact were not properly uploaded to S3, the model creation or endpoint update would fail during deployment, not after the endpoint is already running and returning errors.

218
MCQmedium

A company is using SageMaker Pipelines to orchestrate their ML workflow. They have a Condition step that checks if a model's accuracy exceeds 0.9. If true, they want to register the model in the model registry; otherwise, they want to run a retraining step. Which step type should they use for the decision?

A.Condition step
B.Transform step
C.Processing step
D.Tuning step
AnswerA

A Condition step evaluates a boolean expression against the evaluation step's accuracy metric and branches accordingly: the true branch registers the model, the false branch triggers retraining. It is the only step type providing conditional branching logic in SageMaker Pipelines.

Why this answer

A Condition step in SageMaker Pipelines evaluates a boolean expression against a property (such as a model accuracy metric from a preceding Evaluation step) and branches the pipeline into a true or false path. This is exactly the construct needed to register the model when accuracy > 0.9 and otherwise trigger retraining.

Exam trap

The trap is assuming any step that 'makes a decision' must be a Processing step because it runs code — candidates forget that SageMaker Pipelines has a dedicated Condition step type specifically for branching logic.

How to eliminate wrong answers

Option B is wrong because a Transform step runs batch transform jobs for inference against a dataset, not conditional branching. Option C is wrong because a Processing step runs a containerized data-processing or evaluation job; it produces outputs but does not branch the DAG. Option D is wrong because a Tuning step runs a hyperparameter tuning job that launches multiple training jobs — it has no conditional logic capability.

219
MCQmedium

A company wants to deploy 50 small models (each ~100 MB) for real-time inference. They need to minimize hosting costs while maintaining low latency. Which SageMaker hosting option is most cost-effective?

A.SageMaker Serverless Inference
B.SageMaker Asynchronous Inference
C.SageMaker Multi-Model Endpoint (MME)
D.SageMaker real-time endpoint with one instance per model
AnswerC

Multi-model endpoint loads models on demand into shared memory and caches them, so 50 models share one instance rather than each needing its own endpoint. This satisfies the cost-minimisation constraint while preserving real-time, low-latency inference.

Why this answer

SageMaker Multi-Model Endpoint (MME) hosts many models on a single endpoint and dynamically loads them into memory/disk on invocation, sharing the underlying instance. For 50 small models (~100 MB each), this dramatically reduces hosting cost versus one endpoint per model while still providing real-time, low-latency inference. Serverless Inference is pay-per-invoke but has cold starts and memory limits that can hurt latency for frequent calls.

Exam trap

MLA-C01 often tests the misconception that Serverless Inference is always cheapest, ignoring cold starts and the fact that MME shares one instance across many models for steady real-time traffic.

How to eliminate wrong answers

Option A is wrong because Serverless Inference incurs cold-start latency and is billed per invocation with memory/time caps, which is not ideal for consistently low-latency, high-volume real-time inference across 50 models. Option B is wrong because Asynchronous Inference is designed for large payloads and long processing times with queued requests, not low-latency real-time responses. Option D is wrong because one instance per model for 50 models multiplies hosting costs by 50, directly violating the cost-minimization requirement.

220
MCQmedium

An ML team uses Amazon SageMaker Data Wrangler to prepare a dataset for a binary classification model. They suspect the dataset might contain bias against a certain demographic group. They want to detect and visualize potential bias before training the model. Which feature of SageMaker should they use?

A.SageMaker Debugger
B.SageMaker Experiments
C.SageMaker Model Monitor
D.SageMaker Clarify
AnswerD

SageMaker Clarify runs bias metrics such as class imbalance and disparate impact on the dataset, then renders visual reports before training. This satisfies the requirement to detect and visualise potential bias against a demographic group prior to model training.

Why this answer

SageMaker Clarify is the purpose-built feature for detecting bias in datasets and models, providing pre-training bias metrics (e.g., class imbalance, disparate impact) and post-training metrics, plus visualizations in SageMaker Studio. It analyzes the dataset before training to surface potential bias against demographic groups, exactly matching the team's requirement.

Exam trap

MLA-C01 often tests the confusion between Clarify (bias/explainability) and Model Monitor (drift/quality) — candidates pick Model Monitor thinking 'monitoring' covers bias, but Clarify is the correct pre-training bias tool.

How to eliminate wrong answers

Option A is wrong because SageMaker Debugger focuses on training-job debugging (tensor inspection, vanishing gradients, overfitting detection), not bias analysis. Option B is wrong because SageMaker Experiments tracks and compares training runs, hyperparameters, and metrics — it does not detect bias. Option C is wrong because SageMaker Model Monitor detects data drift and quality issues on deployed endpoints, not pre-training bias.

221
MCQeasy

A retail company stores training datasets, model artifacts, and feature data in Amazon S3. An auditor requires that all objects be encrypted at rest with keys the company controls and that key usage be independently auditable. The team wants minimal operational overhead. Which approach should the ML engineer recommend?

A.Enable default bucket encryption with SSE-S3 and rely on S3 Versioning to protect against unauthorized changes.
B.Use S3 client-side encryption with a key stored in the application's configuration file and rotate it manually each quarter.
C.Use S3 server-side encryption with Amazon S3 managed keys (SSE-S3) and enable S3 access logging.
D.Use S3 server-side encryption with AWS KMS customer-managed keys (SSE-KMS) and enable AWS CloudTrail data events for the bucket.
AnswerD

SSE-KMS with customer-managed keys lets the company control key policies, rotation, and grants, and every use of the key is recorded. CloudTrail data events capture S3 object-level API calls, tying each access to the KMS key usage. This delivers encryption at rest under company-controlled keys with auditable usage and low operational overhead.

Why this answer

SSE-KMS with customer-managed keys gives the company control over key policies, rotation, and access grants while offloading cryptographic operations to AWS. CloudTrail data events record object-level S3 activity, and KMS key usage is logged separately, providing the independent audit trail the auditor requires. SSE-S3, client-side encryption with local keys, and versioning do not meet the customer-controlled key requirement.

Exam trap

The trap here is treating SSE-S3 as equivalent to customer-controlled encryption, when SSE-S3 keys are managed entirely by AWS.

222
Multi-Selectmedium

A team deploys a machine learning model using an Amazon SageMaker endpoint. They need to monitor for data drift and model quality issues. Which AWS services or features should they use? (Choose THREE.)

Select 3 answers
A.AWS Glue DataBrew
B.Amazon SageMaker Clarify
C.Amazon SageMaker Ground Truth
D.Amazon CloudWatch Logs and Metrics
E.Amazon SageMaker Model Monitor
AnswersB, D, E

SageMaker Clarify detects bias and explains feature attributions, and its drift monitoring integrates with Model Monitor to compare live endpoint traffic against a baseline. This directly satisfies the stem's data drift and model quality monitoring requirement for the deployed SageMaker endpoint.

Why this answer

Amazon SageMaker Model Monitor (E) is the purpose-built feature for continuously monitoring deployed SageMaker endpoints, detecting data drift, model quality degradation, bias drift, and feature attribution drift by comparing live traffic against a baseline. Amazon SageMaker Clarify (B) provides bias detection and explainability (SHAP-based feature attributions), and its bias/explainability metrics are integrated with Model Monitor to surface those drift and fairness issues. Amazon CloudWatch Logs and Metrics (D) capture the endpoint's invocation logs and metrics that Model Monitor analyzes and that trigger CloudWatch alarms when violations are detected.

AWS Glue DataBrew (A) is a visual data preparation tool for cleaning and transforming datasets, not for monitoring live endpoints, and SageMaker Ground Truth (C) is a data labeling service for building training datasets, so neither addresses drift or model quality monitoring.

Exam trap

The trap is confusing SageMaker Clarify (bias/explainability) with SageMaker Model Monitor (data drift/quality), as both involve analyzing data distributions, but Clarify focuses on bias and attributions, while Model Monitor handles drift detection.

223
MCQhard

A company uses Amazon SageMaker Data Wrangler to prepare data for ML. The dataset contains a timestamp column and sensor readings from IoT devices. The data scientist needs to create features such as moving averages and rolling statistics over time windows. Which Data Wrangler transformation type should be selected?

A.Join
B.Custom Python script
C.Group by and aggregate
D.Window function
AnswerD

Window functions compute rolling aggregates across ordered partitions, so moving averages and rolling statistics over the timestamp column are produced directly within Data Wrangler. This satisfies the requirement for time-windowed feature engineering, unlike row-level transforms such as numeric or categorical encodings, which cannot aggregate across neighbouring rows.

Why this answer

Window functions in Amazon SageMaker Data Wrangler allow you to compute moving averages, rolling statistics, and other time-window-based aggregations over ordered partitions of data. This is the correct transformation type because it directly supports operations like `SUM() OVER (ORDER BY timestamp ROWS BETWEEN 2 PRECEDING AND CURRENT ROW)` without requiring custom code or losing row-level granularity.

Exam trap

The trap here is that candidates confuse 'Group by and aggregate' with 'Window function' because both involve aggregation, but Group by reduces rows while Window functions preserve row-level detail, which is essential for rolling statistics.

How to eliminate wrong answers

Option A is wrong because Join is used to combine datasets based on a common key, not to compute rolling statistics over a time window. Option B is wrong because while a Custom Python script could technically implement moving averages, Data Wrangler provides a native Window function transformation that is more efficient, easier to maintain, and avoids the overhead of writing and debugging custom code. Option C is wrong because Group by and aggregate collapses rows into summary statistics per group, which loses the individual row-level detail needed for rolling window calculations.

224
MCQmedium

A machine learning engineer is preparing a training dataset in Amazon SageMaker for a binary classification model. The dataset is stored as a single CSV file in Amazon S3 and contains 12 categorical features with high cardinality (thousands of unique values each). The engineer wants to avoid the curse of dimensionality and reduce training time while preserving predictive power. Which preprocessing approach should be used with the SageMaker built-in XGBoost algorithm?

A.Use target encoding (mean encoding) for the categorical features, computed within cross-validation folds to prevent leakage.
B.Convert the categorical features to integer codes using LabelEncoder and pass them directly to XGBoost.
C.Hash the categorical features into a fixed number of buckets using a hashing trick, then one-hot encode the hashed values.
D.Apply one-hot encoding to all categorical features before training.
AnswerA

Target encoding replaces each category with a statistic (e.g., mean of the target) and is well-suited for high-cardinality features. Computing it within cross-validation folds prevents target leakage, which would otherwise inflate validation performance. This reduces dimensionality and training time while retaining predictive signal, directly addressing the engineer's goals without creating thousands of sparse columns.

Why this answer

High-cardinality categorical features require an encoding that avoids exponential feature expansion. Target encoding summarizes each category by its relationship to the target, reducing dimensionality and training time. Performing it within cross-validation folds prevents data leakage, ensuring that validation metrics remain honest.

This approach is compatible with the SageMaker built-in XGBoost algorithm, which accepts numerical features and can benefit from the reduced feature space.

Exam trap

The trap here is assuming that one-hot encoding is always the default for categorical features, overlooking the dimensionality explosion with high-cardinality data.

225
Multi-Selecthard

A team is building a SageMaker Pipeline that trains a model and then registers it in the SageMaker Model Registry. They want the pipeline to automatically deploy the model to a real-time endpoint only after a human approves the model package. Which two actions should the team take to implement this approval-gated deployment? (Choose two.)

Select 2 answers
A.Set the model package status to Approved directly in the RegisterModel step configuration so the endpoint deploys without human intervention.
B.Add a RegisterModel step to the pipeline that creates a model package with a PendingManualApproval status.
C.Use a SageMaker Clarify processing step to approve the model package when bias metrics fall within thresholds.
D.Add a ConditionStep to the pipeline that checks the model package status immediately after RegisterModel and branches to a deployment step.
E.Configure an Amazon EventBridge rule that matches the model package state change to Approved and invokes a target that starts the deployment.
AnswersB, E

RegisterModel creates a model package group entry and sets the package status to PendingManualApproval when manual approval is configured. This produces the artifact that a reviewer evaluates and approves, which is the trigger point the deployment automation watches for.

Why this answer

Manual approval gating relies on the model package lifecycle: RegisterModel creates a package in PendingManualApproval, and a reviewer later approves it. Because the pipeline execution does not pause for review, an EventBridge rule watching for the Approved state change is what triggers the downstream deployment automation, keeping humans in the loop.

Exam trap

The trap here is assuming a pipeline step can wait for human approval, when pipelines are finite executions and approval events must be handled by an external event-driven mechanism.

Page 2

Page 3 of 9

Page 4

All pages