Courseiva

AWS Certified Machine Learning Engineer Associate MLA-C01 (MLA-C01) — Questions 301–375

665 questions total · 9pages · All types, answers revealed

Page 4

Page 5 of 9

Page 6
301
MCQhard

A financial services company operates a real-time inference endpoint for a fraud detection model on Amazon SageMaker. The model was trained on historical transaction data from 2023. Over the past month, the model's precision has dropped from 92% to 78%, while recall remains high at 95%. The data science team suspects data drift and has already enabled SageMaker Model Monitor with data capture and a baseline from the training data. The latest monitoring report indicates no statistically significant drift in any of the input features. The team also verified that the inference code and model artifact have not changed. Despite the stable feature distributions, the model is misclassifying an increasing number of legitimate transactions as fraudulent (false positives). The business is concerned about the impact on customer experience. What is the best course of action?

A.Replace the model with a more complex algorithm such as a gradient-boosted tree.
B.Retrain the model using the most recent 30 days of transaction data with automated retraining pipelines.
C.Increase the data capture sampling percentage from 10% to 100% for more detailed analysis.
D.Investigate recent ground truth labels to check for label drift or changes in the fraud definition.
AnswerD

Label drift occurs when the underlying relationship between features and labels changes. Collecting and analyzing recent labels can confirm if the fraud criteria have shifted.

Why this answer

The model's precision dropped while recall remained high, meaning it is flagging more legitimate transactions as fraudulent (false positives). Since input feature distributions show no drift and the model artifact is unchanged, the issue likely lies in the target variable—label drift or a change in the definition of fraud. Investigating recent ground truth labels will reveal if the labels used for evaluation or retraining have shifted, which would explain the degradation.

This is the best course of action before retraining or changing the model.

Exam trap

MLA-C01 often tests the distinction between feature drift and label drift, so candidates must remember that stable feature distributions do not rule out label drift, which requires investigating ground truth labels.

How to eliminate wrong answers

Option A is wrong because replacing the algorithm does not address the root cause, which is likely label drift, and could introduce new issues. Option B is wrong because retraining with recent data without understanding the label shift may perpetuate the problem if the labels themselves are inconsistent. Option C is wrong because increasing data capture sampling does not address the cause of false positives; the monitoring already shows no feature drift, so more data capture is unnecessary.

302
MCQeasy

Refer to the exhibit. The data scientist wants to update the endpoint to use a new model version without downtime. Which approach should they use?

A.Delete the existing endpoint and create a new one
B.Update the endpoint's model name directly
C.Create a new endpoint configuration with a second variant and update the endpoint
D.Use SageMaker Model Monitor to automatically switch
AnswerC

Adding a second variant to a new endpoint configuration and updating the endpoint enables a blue/green style shift, where the new model version receives traffic alongside the old. This satisfies the no-downtime constraint by avoiding endpoint deletion and recreation.

Why this answer

It uses SageMaker's blue/green deployment pattern: creating a new endpoint configuration with a second variant (the new model version) and updating the endpoint shifts traffic gradually or instantly without downtime. This approach leverages the endpoint's ability to host multiple variants and route traffic between them, ensuring zero interruption during the update.

Exam trap

The trap here is that candidates confuse 'updating the endpoint' with 'modifying the existing configuration directly' (Option B), not realizing that SageMaker enforces immutability of endpoint configurations and requires a new configuration to change the model or variant.

How to eliminate wrong answers

Option A is wrong because deleting the existing endpoint and creating a new one causes downtime—the endpoint is unavailable during the deletion and recreation process. Option B is wrong because updating the endpoint's model name directly is not a supported operation; SageMaker endpoints are immutable once created, and you must use a new endpoint configuration to change the model. Option D is wrong because SageMaker Model Monitor is designed for monitoring data quality and drift, not for automatically switching model versions; it has no mechanism to update endpoint variants.

303
Multi-Selectmedium

A team wants to secure SageMaker endpoints for a healthcare application. They must ensure data is encrypted at rest and in transit, and that the endpoint can only be accessed from within a VPC. Which THREE steps should they take? (Select THREE)

Select 3 answers
A.Use AWS KMS to encrypt the model artifacts and endpoint data
B.Store encryption keys in a public S3 bucket
C.Set the endpoint to use network isolation mode
D.Configure the endpoint to use a VPC and disable public access
E.Enable inter-container traffic encryption using TLS
AnswersA, D, E

KMS provides encryption at rest for data and models.

304
MCQmedium

A data scientist is using Amazon SageMaker Data Wrangler to create a data preparation flow. After completing the flow, the scientist wants to export the processed data to a feature group in Amazon SageMaker Feature Store for reuse in multiple training jobs. Which export option should the scientist choose in Data Wrangler?

A.Export to S3
B.Export as a Python script
C.Save as a Jupyter notebook
D.Export to Feature Store
AnswerD

Export to Feature Store writes the processed flow output directly into a SageMaker Feature Store feature group, registering the features for reuse across multiple training jobs. This satisfies the requirement to make the prepared data reusable rather than exporting a one-off dataset.

Why this answer

Data Wrangler provides a direct 'Export to Feature Store' option that writes the transformed dataset into a SageMaker Feature Store feature group, making it immediately available for reuse across multiple training jobs and real-time inference. This is the purpose-built integration for feature reuse.

Exam trap

MLA-C01 often tests whether candidates confuse generic S3 export with the purpose-built Feature Store export — the key differentiator is that only the Feature Store export registers features for cross-job reuse with online/offline serving.

How to eliminate wrong answers

Option A is wrong because exporting to S3 only stores the processed data as objects — it does not register features in Feature Store, so downstream training jobs would need custom loading logic. Option B is wrong because exporting as a Python script generates reproducible transformation code but does not itself populate a feature group. Option C is wrong because saving as a Jupyter notebook preserves the analysis workflow but does not create or write to a Feature Store feature group.

305
MCQeasy

A data scientist needs to store training data in Amazon S3 and wants to optimize read performance for iterative training jobs. Which S3 feature should they use?

A.S3 Transfer Acceleration
B.S3 Glacier
C.S3 Byte-Range Fetches
D.S3 Select
AnswerC

S3 Byte-Range Fetches let training jobs retrieve only the specific byte ranges needed from large objects, enabling parallelised, partial reads rather than downloading entire files. This satisfies the requirement to optimise read performance for iterative training jobs.

Why this answer

Byte-range fetches allow the data scientist to parallelize reads by requesting specific byte ranges of an object, which significantly improves read performance for iterative training jobs that need to access large datasets stored in S3. This feature enables multiple concurrent requests to different parts of the same object, reducing latency and increasing throughput compared to single-range reads.

Exam trap

AWS often tests the distinction between features that optimize data transfer (like S3 Transfer Acceleration) versus those that optimize data access patterns (like Byte-Range Fetches), leading candidates to confuse upload acceleration with read performance optimization.

How to eliminate wrong answers

Option A is wrong because S3 Transfer Acceleration is designed to speed up uploads over long distances using edge locations and optimized network paths, not to improve read performance for iterative training jobs. Option B is wrong because S3 Glacier is a cold storage class for archival data with retrieval times ranging from minutes to hours, making it unsuitable for frequent iterative training access. Option D is wrong because S3 Select is used to retrieve a subset of data from an object using SQL-like queries, which reduces data transfer but does not optimize parallel read performance for the entire training dataset.

306
MCQmedium

A fraud-detection team trains a model in SageMaker and wants to shift 10 percent of live prediction traffic to a newly retrained model to compare accuracy before a full cutover. The endpoint already serves the current model on one production variant. They need the endpoint to route a controlled fraction of requests to the new model without changing the client application. What should they do?

A.Update the existing endpoint to add a second production variant for the retrained model and set its initial variant weight to 10 while the original variant keeps 90.
B.Enable an inference pipeline on the endpoint and place the retrained model as the second container in the pipeline.
C.Use a SageMaker batch transform job against the live traffic stream to score 10 percent of requests with the retrained model.
D.Create a new endpoint for the retrained model and use Route 53 weighted routing to send 10 percent of DNS queries to the new endpoint.
AnswerA

SageMaker production variants let one endpoint host multiple models with assigned weights, and InvokeEndpoint distributes traffic according to those weights. Setting the new variant to 10 and the existing one to 90 shifts a controlled fraction of requests without any client change, since the client still calls the same endpoint name. This is the built-in mechanism for canary-style traffic shifting.

Why this answer

Production variants with weights are the native SageMaker mechanism for sending a percentage of live traffic to a new model on the same endpoint. Adding a variant for the retrained model and assigning it a weight of 10 while the incumbent holds 90 yields a canary deployment that clients see as a single endpoint. Weights can be adjusted as confidence grows and the old variant removed at full cutover.

Exam trap

The trap here is reaching for DNS-level or batch mechanisms for traffic splitting, when SageMaker production variant weights operate per invocation on a single endpoint.

307
MCQeasy

A team deploys a SageMaker real-time endpoint and configures it with an auto scaling policy targeting a variant. During a flash sale, traffic spikes and the team notices that the number of instances increases, but the average model latency still climbs above the target. The team wants the scaling behavior to react faster to sudden bursts without over-provisioning during steady periods. Which change should they make to the scaling policy?

A.Switch the policy from target tracking to a step scaling policy with a larger scale-out cooldown and a smaller scale-in cooldown.
B.Lower the target value of the predefined InvocationsPerInstance metric and shorten the scale-out cooldown so capacity is added sooner.
C.Replace target tracking with a scheduled scaling policy that adds instances at fixed times each day.
D.Increase the target value of the predefined InvocationsPerInstance metric and enable a longer scale-in cooldown to stabilize the fleet.
AnswerB

Target tracking adjusts capacity to keep the metric near the target. Lowering the InvocationsPerInstance target means the policy scales out at a lower request rate per instance, adding capacity earlier, and shortening the scale-out cooldown allows consecutive scale-out actions to occur more quickly. Together they make the endpoint respond faster to bursts while still scaling in during steady, low-traffic periods.

Why this answer

Target tracking keeps a chosen metric near its target. Lowering the InvocationsPerInstance target causes scale-out to trigger at a lower request rate per instance, so additional capacity arrives earlier during a burst. Shortening the scale-out cooldown lets successive scale-out actions fire sooner.

Scaling in still occurs during quiet periods, so steady-state cost is not inflated.

Exam trap

The trap here is assuming a longer cooldown or a higher metric target improves burst handling, when both actually delay scale-out and let latency rise during sudden spikes.

308
MCQmedium

A company ingests streaming transaction data from multiple sources using Amazon Kinesis Data Streams. The data must be transformed (e.g., JSON parsing, data type conversions) and then stored in Amazon S3 for ML training. The transformation logic may change over time. Which approach provides the greatest flexibility and ease of maintenance?

A.Use a custom application running on Amazon EC2 with the Kinesis Client Library to transform and write to S3.
B.Use Kinesis Data Firehose with a Lambda function for transformation and delivery to S3.
C.Use Kinesis Data Analytics with SQL queries to transform data and output to S3.
D.Use Kinesis Data Streams directly with a Lambda consumer to transform and store in S3.
AnswerB

Kinesis Data Firehose with a Lambda function decouples transformation logic from delivery, so engineers can update the Lambda code without rebuilding the pipeline. This satisfies the requirement that transformation logic may change over time while still landing data in Amazon S3.

Why this answer

Kinesis Data Firehose with a Lambda function for transformation gives you a fully managed delivery pipeline to S3 while letting you change transformation logic by updating the Lambda code, with no servers to manage. This satisfies both the flexibility (logic can evolve) and ease-of-maintenance requirements.

Exam trap

MLA-C01 often tests whether candidates pick the most 'powerful' streaming service — the trap is choosing Kinesis Data Analytics or a custom EC2 consumer when the requirement is flexible, low-maintenance transformation and S3 delivery, which Firehose with Lambda handles natively.

How to eliminate wrong answers

Option A is wrong because a custom EC2 application with the Kinesis Client Library requires you to manage servers, scaling, checkpointing, and deployment, which is high maintenance and inflexible. Option C is wrong because Kinesis Data Analytics with SQL is optimized for streaming SQL analytics and windowed aggregations, not for arbitrary JSON parsing and type conversion logic that changes over time. Option D is wrong because consuming Kinesis Data Streams directly with a Lambda consumer still requires you to build and maintain the S3 delivery, batching, retry, and error-handling logic that Firehose provides out of the box.

309
Multi-Selecthard

A company wants to use SageMaker to fine-tune a foundation model for a text generation task using RLHF (Reinforcement Learning from Human Feedback). Which THREE components are required in the RLHF pipeline?

Select 3 answers
A.A LoRA adapter for parameter-efficient fine-tuning
B.A pre-trained base model
C.A classifier to distinguish generated text from real text
D.A reward model trained on human preferences
E.A reinforcement learning algorithm such as PPO
AnswersB, D, E

RLHF starts from a pre-trained base model, which supplies the generative prior that the policy is later refined from. Without it, there is no initial language model to sample responses from or to update during reinforcement learning, so the pipeline cannot begin.

Why this answer

Option B is correct because RLHF always starts from a pre-trained foundation model (the policy) that already has broad language capabilities; this base model is then fine-tuned with human preference signals rather than trained from scratch. Option D is correct because RLHF requires a reward model trained on human preference comparisons (e.g., chosen vs. rejected responses) to serve as a differentiable proxy for human judgment during optimization. Option E is correct because the policy is optimized against that reward model using a reinforcement learning algorithm such as PPO (Proximal Policy Optimization), typically with a KL penalty to the reference model to prevent reward hacking.

Option A is not required: LoRA is a parameter-efficient fine-tuning technique that can optionally be used, but RLHF does not mandate adapters. Option C is not required: distinguishing generated from real text is a GAN discriminator concept, not part of the RLHF pipeline, which relies on preference-based reward modeling instead.

Exam trap

MLA-C01 often tests whether candidates confuse RLHF components with general fine-tuning techniques like LoRA or with GAN-style discriminators, causing them to pick optional or unrelated options.

310
MCQeasy

A data scientist is using Amazon SageMaker Data Wrangler to prepare a dataset. They need to identify potential bias in the data before training. Which SageMaker feature should they use?

A.Amazon SageMaker Model Monitor
B.Amazon SageMaker Debugger
C.Amazon SageMaker Clarify
D.Amazon SageMaker Pipelines
AnswerC

SageMaker Clarify detects bias through its bias metrics, analysing feature distributions and label correlations before training begins. Data Wrangler's built-in bias report integrates Clarify directly, satisfying the requirement to identify potential bias during data preparation rather than post-training. Other SageMaker features address model training or deployment, not pre-training bias detection.

Why this answer

Amazon SageMaker Clarify is the correct feature because it is specifically designed to detect bias in datasets and machine learning models. It provides built-in bias metrics (e.g., pre-training bias) and can generate bias reports during data preparation, directly addressing the need to identify potential bias before training.

Exam trap

The trap here is that candidates confuse SageMaker Clarify with SageMaker Model Monitor or Debugger, assuming any monitoring or debugging tool can detect bias, but only Clarify provides dedicated bias analysis for both data and models.

How to eliminate wrong answers

Option A is wrong because Amazon SageMaker Model Monitor is used to monitor deployed models for data drift and quality issues in production, not for detecting bias in training data. Option B is wrong because Amazon SageMaker Debugger is designed to debug training jobs by capturing tensors and metrics, not for bias detection in datasets. Option D is wrong because Amazon SageMaker Pipelines is a workflow orchestration service for building and managing ML pipelines, not a tool for bias analysis.

311
MCQeasy

A machine learning engineer has trained a model in SageMaker and wants to deploy it to a real-time endpoint for low-latency inference. The model artifacts are stored in Amazon S3, and the engineer needs to create the endpoint with the least operational effort. Which sequence of actions should the engineer take?

A.Create a model, create an endpoint configuration, and create an endpoint using the SageMaker API or the AWS SDK for Python (Boto3).
B.Package the model into a Docker image, push it to Amazon ECR, and run it on an Amazon EC2 instance behind an Application Load Balancer.
C.Create a batch transform job that reads from Amazon S3 and writes predictions to Amazon S3, then expose the output as a REST API.
D.Register the model in the SageMaker model registry and deploy it through a SageMaker pipeline with a manual approval step.
AnswerA

Deploying to a real-time endpoint in SageMaker requires three resources: a model that points to the artifacts and image, an endpoint configuration that defines the instance type and count, and an endpoint that provisions the compute. Using the SageMaker API or Boto3 performs these steps directly and is the standard low-effort path for a single model deployment.

Why this answer

A SageMaker real-time endpoint is composed of a model, an endpoint configuration, and an endpoint. Creating these three resources through the SageMaker API or Boto3 is the direct, low-effort way to deploy artifacts from Amazon S3 for low-latency inference. Batch transform, model registry pipelines, and self-managed EC2 hosting either do not provide real-time serving or add unnecessary operational burden.

Exam trap

The trap here is conflating model registration or batch transform with real-time hosting, when a real-time endpoint requires the model, configuration, and endpoint resources.

312
MCQmedium

A data scientist needs to run a hyperparameter tuning job for a PyTorch model using SageMaker. They want to use Hyperband for efficient resource allocation. Which tuning strategy should they select in the HyperparameterTuner?

A.Bayesian optimization
B.Hyperband
C.Random search
D.Grid search
AnswerB

Hyperband is the strategy that terminates poorly performing training jobs early and reallocates resources to promising trials. Selecting it in the HyperparameterTuner delivers the efficient resource allocation the data scientist requires for the PyTorch tuning job.

Why this answer

SageMaker's HyperparameterTuner supports Hyperband as a first-class strategy, which implements the multi-armed bandit early-stopping algorithm to allocate more resources to promising configurations and prune poor ones early. Selecting 'Hyperband' as the strategy directly enables this efficient resource allocation. It is distinct from Bayesian, Random, and Grid strategies, which do not perform the successive-halving resource allocation that Hyperband does.

Exam trap

MLA-C01 often tests whether candidates know that Hyperband is a distinct strategy option in HyperparameterTuner, not a synonym for Bayesian optimization or early stopping — picking Bayesian is the most common wrong answer.

How to eliminate wrong answers

Option A is wrong because Bayesian optimization builds a probabilistic surrogate model to choose the next hyperparameters but does not implement the successive-halving resource allocation that defines Hyperband. Option C is wrong because Random search samples hyperparameters uniformly without any early-stopping or resource-allocation mechanism. Option D is wrong because Grid search exhaustively tries every combination and is the least efficient, with no pruning or resource reallocation.

313
MCQeasy

A data science team deploys a real-time inference endpoint on Amazon SageMaker. They want to monitor for data drift in the input features over time. Which AWS service should they use to capture and analyze the input data distribution?

A.Amazon Athena
B.AWS CloudTrail
C.Amazon SageMaker Model Monitor
D.Amazon CloudWatch Logs
AnswerC

Amazon SageMaker Model Monitor captures endpoint request data and compares its distribution against a baseline, detecting drift in input features. It satisfies the requirement for continuous, automated data drift monitoring on a live real-time inference endpoint, publishing violations to Amazon CloudWatch with configurable thresholds.

Why this answer

Amazon SageMaker Model Monitor is the correct service because it is specifically designed to continuously monitor machine learning models in production for data drift and quality issues. It automatically captures input data distributions from real-time inference endpoints and compares them against a baseline to detect statistical changes, alerting the team when drift occurs.

Exam trap

The trap here is that candidates confuse general logging and monitoring services (CloudWatch Logs, CloudTrail) with the specialized model monitoring service, overlooking that SageMaker Model Monitor provides built-in statistical drift detection rather than just raw log storage.

How to eliminate wrong answers

Option A is wrong because Amazon Athena is an interactive query service for analyzing data in Amazon S3 using standard SQL, not a tool for capturing or monitoring data drift from SageMaker endpoints. Option B is wrong because AWS CloudTrail records API activity for auditing and governance, not input data distributions or model performance metrics. Option D is wrong because Amazon CloudWatch Logs stores and monitors log files and metrics, but it lacks built-in capabilities for statistical drift detection or baseline comparison of feature distributions.

314
MCQeasy

A data science team needs to deploy a PyTorch model for real-time inference with low latency. The model requires GPU acceleration. Which SageMaker endpoint configuration should they use?

A.Create a multi-model endpoint using ml.m5.large instances
B.Create a serverless endpoint with memory set to 6144 MB
C.Create a batch transform job using an ml.c5.xlarge instance
D.Create a real-time endpoint using an ml.p3.2xlarge instance
AnswerD

An ml.p3.2xlarge instance provides NVIDIA V100 GPU acceleration, satisfying the stem's GPU requirement, and real-time endpoints deliver the low-latency synchronous inference the team needs. SageMaker real-time endpoints keep the model loaded and respond in milliseconds, unlike serverless or asynchronous options that cap resources or add queuing delay.

Why this answer

Real-time SageMaker endpoints with GPU instances like ml.p3.2xlarge are specifically designed for low-latency, synchronous inference with GPU acceleration. PyTorch models requiring GPU must use instance types that support NVIDIA CUDA, and the ml.p3 family provides the necessary GPU compute for real-time predictions.

Exam trap

The trap here is that candidates may confuse batch transform jobs or serverless endpoints with real-time inference, overlooking the explicit GPU requirement and the need for persistent, low-latency compute resources.

How to eliminate wrong answers

Option A is wrong because multi-model endpoints using ml.m5.large instances are CPU-based and lack GPU acceleration, making them unsuitable for PyTorch models that require GPU for low-latency inference. Option B is wrong because serverless endpoints do not support GPU acceleration; they are limited to CPU compute and cannot meet the GPU requirement. Option C is wrong because batch transform jobs are designed for asynchronous, offline inference on large datasets, not for real-time, low-latency predictions.

315
MCQeasy

A data scientist discovers that a dataset for binary classification contains 95% negative samples and 5% positive samples. Which technique is MOST appropriate to address the class imbalance?

A.Use SMOTE to generate synthetic samples for the minority class
B.Remove all samples from the majority class
C.Increase the learning rate of the model
D.Downsample the majority class to 5% of the original size
AnswerA

SMOTE generates synthetic minority-class samples by interpolating between existing positive instances, rebalancing the 95:5 distribution. This gives the classifier enough positive examples to learn the minority pattern, unlike plain resampling that discards data or duplicates points.

Why this answer

SMOTE (Synthetic Minority Over-sampling Technique) generates new, synthetic minority-class samples by interpolating between existing minority neighbors, which increases the minority representation without simply duplicating rows. This helps the model learn the minority decision boundary and is the standard, most appropriate technique for a 95/5 imbalance. It preserves all majority information while enriching the minority class.

Exam trap

MLA-C01 often tests the misconception that any resampling works — candidates must distinguish SMOTE's synthetic interpolation from naive duplication (RandomOverSampler) and from destructive undersampling that throws away majority data.

How to eliminate wrong answers

Option B is wrong because removing all majority-class samples destroys the dataset and leaves the model with no negative examples to learn from — it is not a viable imbalance remedy. Option C is wrong because learning rate controls optimization step size, not class distribution; raising it can cause divergence and does nothing to fix imbalance. Option D is wrong because downsampling the majority to 5% of its original size discards ~95% of legitimate majority data, drastically reducing the training set and losing information that could improve generalization.

316
MCQmedium

A company wants to serve 200 different PyTorch models. Each model is small (under 1 GB) and only a fraction are used at any time. To minimize cost and management overhead, which SageMaker inference option should be used?

A.Use batch transform for all models
B.Create a separate real-time endpoint for each model
C.Use a multi-container endpoint
D.Use a multi-model endpoint
AnswerD

Multi-model endpoints load models on demand from S3 into shared container memory and unload idle ones, so hosting 200 small PyTorch models costs far less than 200 endpoints. This satisfies the minimise-cost and management-overhead constraint given only a fraction are used concurrently.

Why this answer

A multi-model endpoint (MME) is the correct choice because it allows you to host multiple small PyTorch models (under 1 GB each) on a single endpoint, sharing the underlying compute instance. This minimizes cost by only paying for the active instances, and reduces management overhead since you don't need to create or manage separate endpoints for each model. SageMaker dynamically loads and unloads models from the container's memory based on invocation patterns, which is ideal for a scenario where only a fraction of the 200 models are used at any time.

Exam trap

In AWS SageMaker, the common trap is confusing multi-container endpoints (used for multi-step inference pipelines with different containers) with multi-model endpoints (used for hosting multiple independent models). Candidates often select multi-container endpoints thinking they can serve multiple models, but each container still hosts one model; multi-model endpoints are designed to host many models on a single instance, dynamically loading them as needed.

How to eliminate wrong answers

Option A is wrong because batch transform is designed for offline, asynchronous inference on a complete dataset, not for serving real-time predictions on-demand, and it would incur cost for processing all models even when not needed. Option B is wrong because creating a separate real-time endpoint for each of the 200 models would be prohibitively expensive and introduce significant management overhead, as each endpoint requires its own compute instance and incurs hourly charges regardless of usage. Option C is wrong because a multi-container endpoint is used to run multiple containers (e.g., for pre-processing, inference, post-processing) within a single endpoint, not to host multiple independent models; it does not support dynamic loading/unloading of different model artifacts.

317
Multi-Selecthard

A machine learning engineer is evaluating a binary classification model for detecting fraudulent transactions. The dataset is highly imbalanced, and the cost of false negatives (missing a fraud) is very high. Which two evaluation metrics should the engineer consider? (Choose two.)

Select 2 answers
A.F1-score
B.Accuracy
C.Recall
D.Precision
E.Mean absolute error
AnswersA, C

F1-score is the harmonic mean of precision and recall, so it penalises models that sacrifice either measure. On a highly imbalanced fraud dataset with costly false negatives, it captures the precision-recall trade-off that accuracy would hide behind the dominant negative class.

Why this answer

Recall (Option C) is correct because it measures the proportion of actual fraud cases that the model successfully identifies (TP / (TP + FN)), directly addressing the scenario's high cost of false negatives—maximizing recall minimizes missed frauds. F1-score (Option A) is correct because it is the harmonic mean of precision and recall (2 × (Precision × Recall) / (Precision + Recall)), providing a balanced single metric that remains informative on highly imbalanced datasets where a model could otherwise achieve high recall by flagging everything. Accuracy (Option B) is not appropriate because with severe class imbalance a trivial model predicting 'no fraud' for all transactions can score very high accuracy while catching zero frauds.

Precision (Option D) is relevant but not among the two marked correct; it focuses on false positives rather than the costly false negatives and is already incorporated into the F1-score. Mean absolute error (Option E) is a regression metric and does not apply to binary classification evaluation.

Exam trap

AWS often tests the misconception that accuracy is always a good metric, but in imbalanced classification problems, accuracy is a trap because a naive model can achieve high accuracy by simply predicting the majority class.

318
MCQmedium

An ML team is using Amazon SageMaker Feature Store to serve features for both real-time inference and batch training. They need to ensure that training data uses feature values as they were at the time of each event. Which type of query should they use?

A.Full table scan of the offline store
B.Join query across both stores
C.Latest record query from the online store
D.Point-in-time query from the offline store
AnswerD

Point-in-time queries from the offline store retrieve feature values as they existed at each event's timestamp, preventing label leakage from later updates. This directly satisfies the requirement that training data reflect historical feature states, since the offline store retains the full time-ordered history that the online store does not.

Why this answer

SageMaker Feature Store's offline store supports point-in-time queries, which return feature values as they existed at a specified event timestamp for each record. This prevents 'feature leakage' — using future feature values that were not available when the event occurred — which would otherwise inflate training accuracy and break production parity.

Exam trap

MLA-C01 often tests the confusion between online-store 'latest value' retrieval and offline-store point-in-time retrieval — candidates must recognize that training requires historical, time-correct values, not the freshest ones.

How to eliminate wrong answers

Option A is wrong because a full table scan returns the latest or all versions of features without respecting event time, causing temporal leakage. Option B is wrong because joining online and offline stores is not a supported pattern for training data retrieval and does not solve time-consistency. Option C is wrong because the online store only holds the latest feature values for low-latency inference; it has no historical versioning and cannot reconstruct past states.

319
MCQhard

A company deploys a real-time inference endpoint with auto-scaling using a target tracking policy based on average Invocations per instance. They notice that during a traffic spike, the endpoint scales out too late, causing increased latency. They want to scale proactively before the spike. Which strategy should they implement?

A.Enable provisioned concurrency on the endpoint
B.Pre-warm the endpoint by sending dummy requests
C.Use a scheduled scaling action to add capacity before the expected spike
D.Switch to a step scaling policy with a higher cooldown period
AnswerC

Scheduled scaling adds capacity at predetermined times, so instances are already running before the expected traffic spike. This proactive approach avoids the lag inherent in target tracking, which reacts only after Invocations per instance rises.

Why this answer

A scheduled scaling action adds capacity at a predetermined time before the expected traffic spike, allowing the endpoint to be ready proactively. Target tracking reacts after metrics breach thresholds, which is inherently reactive and too slow for sharp spikes. Scheduled scaling is the correct strategy when spikes are predictable (e.g., known business hours or events).

Exam trap

MLA-C01 often tests reactive versus proactive scaling, so candidates pick target tracking variants or provisioned concurrency, missing that scheduled scaling is the only truly proactive option for predictable spikes.

How to eliminate wrong answers

Option A is wrong because provisioned concurrency is a Lambda feature that pre-initialises execution environments; it does not apply to SageMaker endpoints and does not address proactive capacity for predictable spikes. Option B is wrong because pre-warming with dummy requests is a hack that does not guarantee capacity at spike time and wastes resources. Option D is wrong because step scaling with a higher cooldown period is still reactive and would delay further scaling, worsening the late-scale problem.

320
Multi-Selectmedium

A company wants to deploy a PyTorch model on SageMaker for real-time inference. Which two steps are required? (Select TWO.)

Select 2 answers
A.Upload the training data to an S3 bucket.
B.Register the model in the SageMaker Model Registry.
C.Package the model artifacts into a tar.gz file.
D.Create a SageMaker endpoint configuration with the desired instance type.
E.Set up a SageMaker Notebook instance.
AnswersC, D

SageMaker requires model artefacts in a compressed tar.gz archive in Amazon S3 before deployment; the inference container extracts and loads them. Packaging the PyTorch model this way satisfies the deployment prerequisite for hosting on an endpoint.

Why this answer

Option C is correct because SageMaker requires model artifacts to be packaged as a tar.gz archive (typically containing the model files and an inference script in a directory structure like model.tar.gz) before they can be deployed to a hosting container. Option D is correct because deploying a real-time inference endpoint requires creating an endpoint configuration that specifies the production variant, including the instance type (e.g., ml.m5.large) and the model to serve, which is then used to create the endpoint. Option A is incorrect because uploading training data to S3 is part of the training workflow, not a required step for deploying an already-trained PyTorch model for inference.

Option B is incorrect because the SageMaker Model Registry is optional for governance and versioning; you can deploy a model directly without registering it. Option E is incorrect because a SageMaker Notebook instance is only a development tool and is not required to host a real-time inference endpoint.

Exam trap

The trap here is that candidates often confuse the optional Model Registry step (B) as mandatory for deployment, or mistakenly think uploading training data (A) is needed for inference, when in fact only the model artifact packaging (C) and endpoint configuration (D) are the two required steps for real-time inference on SageMaker.

321
MCQmedium

A company is building a sentiment analysis model for customer reviews. The text data contains many typos, abbreviations, and informal language. Which text preprocessing step would be most beneficial to support the ML model?

A.Apply stemming and lowercasing only
B.Remove all stop words
C.Tokenize using whitespace tokenizer
D.Apply spell-check and normalize abbreviations
AnswerD

Spell-checking and abbreviation normalisation collapse noisy variants such as "gr8" and "thx" onto canonical tokens, shrinking vocabulary sparsity so the sentiment model learns consistent signal from informal reviews rather than treating each typo as a distinct feature.

Why this answer

Typos, abbreviations, and informal language introduce noise that can cause the model to treat the same word as different tokens (e.g., 'gr8' vs 'great'), degrading sentiment classification. Applying spell-check and normalizing abbreviations standardizes the text so the model learns consistent patterns. This directly addresses the specific data quality issue described.

Exam trap

MLA-C01 often tests the difference between generic preprocessing steps (stemming, stop-word removal) and the step that addresses the specific data quality problem — candidates may pick stemming because it is common, but it does not fix typos or abbreviations.

How to eliminate wrong answers

Option A is wrong because stemming and lowercasing do not fix typos or abbreviations; 'luv' would still be a different token from 'love' after stemming. Option B is wrong because removing stop words does not address misspellings or informal shorthand, and stop words can carry sentiment (e.g., 'not' in 'not good'). Option C is wrong because whitespace tokenization is a basic tokenization method that does not correct typos or expand abbreviations; it would actually produce poor tokens for informal text with punctuation.

322
MCQmedium

A company is using SageMaker Debugger to monitor a training job for a deep learning model. They want to detect when gradients become extremely large, which may cause training instability. Which built-in rule should they use?

A.DeadRelu
B.ExplodingGradients
C.VanishingGradients
D.Overfit
AnswerB

The ExplodingGradients built-in rule monitors gradient values across training iterations and raises an alert when they exceed a threshold, indicating instability. This matches the stem's requirement to detect gradients becoming extremely large during the SageMaker Debugger training job.

Why this answer

SageMaker Debugger provides a built-in rule named ExplodingGradients that monitors gradient tensors during training and triggers when gradient magnitudes exceed a threshold, indicating training instability. It is purpose-built to detect the exact condition described — gradients becoming extremely large. The rule can emit a warning or stop the job based on configuration.

Exam trap

MLA-C01 often tests the distinction between ExplodingGradients and VanishingGradients — candidates must read the symptom (extremely large vs. extremely small) carefully, as both are gradient-related Debugger rules.

How to eliminate wrong answers

Option A is wrong because DeadRelu detects ReLU neurons that are permanently inactive (outputting zero), which is a different failure mode related to activation, not gradient magnitude. Option C is wrong because VanishingGradients detects gradients that become too small, the opposite problem of exploding gradients. Option D is wrong because Overfit detects a divergence between training and validation loss, which is about generalization, not gradient magnitude.

323
MCQeasy

A machine learning engineer is using Amazon SageMaker Experiments to track multiple training runs. They want to compare the performance of different hyperparameter configurations visually. Which SageMaker tool provides an interactive interface to compare experiments?

A.SageMaker Studio
B.SageMaker Model Monitor
C.SageMaker Experiments SDK
D.SageMaker Debugger Insights
AnswerA

SageMaker Studio provides the interactive visual interface for comparing experiment runs, satisfying the requirement to compare hyperparameter configurations graphically. Within Studio, the Experiments pane renders trial component metrics as charts and tables, enabling side-by-side analysis of training runs without custom code.

Why this answer

SageMaker Studio provides an interactive, web-based interface that allows you to visually compare experiment runs, including hyperparameter configurations and performance metrics, through built-in experiment management and visualization tools. This is the correct answer because the question specifically asks for an interactive interface, which Studio offers natively, unlike the other options which are programmatic or monitoring-focused.

Exam trap

The trap here is that candidates confuse the SageMaker Experiments SDK (a programmatic tool) with the interactive visual interface provided by SageMaker Studio, leading them to select option C because they think 'Experiments' implies a visual tool, but the SDK is code-only.

How to eliminate wrong answers

Option B is wrong because SageMaker Model Monitor is designed for detecting data drift and model quality degradation over time, not for comparing hyperparameter configurations across training runs. Option C is wrong because the SageMaker Experiments SDK is a programmatic API for logging and querying experiment data, not an interactive visual interface. Option D is wrong because SageMaker Debugger Insights provides debugging and profiling information during training, such as gradients and tensors, but lacks the interactive experiment comparison capabilities needed for hyperparameter analysis.

324
MCQhard

A team is fine-tuning a foundation model using reinforcement learning from human feedback (RLHF) on SageMaker. They have a dataset of human preferences. Which SageMaker capability is most suitable for the reward model training step?

A.SageMaker JumpStart
B.SageMaker Ground Truth
C.SageMaker Autopilot
D.SageMaker Training with a custom PyTorch container
AnswerD

SageMaker Training with a custom PyTorch container suits reward-model training because RLHF reward models are bespoke regression networks, not standard supervised tasks. A custom container lets you define the pairwise preference loss, initialise from the fine-tuned foundation model's weights, and control hyperparameters — flexibility that built-in algorithms and Autopilot cannot provide for this scenario.

Why this answer

Training a reward model for RLHF requires a custom training loop over human preference pairs, typically using a PyTorch model with a regression or ranking loss. SageMaker Training with a custom PyTorch container gives full control over the training script, loss function, and data pipeline needed for this step. The other SageMaker capabilities are higher-level or data-labeling services that do not provide this flexibility.

Exam trap

MLA-C01 often tests whether candidates confuse data-labeling (Ground Truth) or pre-built model hubs (JumpStart) with the custom training capability needed for RLHF reward models — the key is recognizing the need for a custom training script.

How to eliminate wrong answers

Option A is wrong because JumpStart provides pre-trained models and fine-tuning templates for common tasks, not a custom reward-model training workflow for RLHF. Option B is wrong because Ground Truth is a data labeling service for generating human annotations, not for training a reward model. Option C is wrong because Autopilot automates model selection and hyperparameter tuning for tabular data, and does not support custom RLHF reward-model training.

325
MCQeasy

A data scientist needs to annotate a large dataset of images for an object detection model. The team wants to minimize manual labeling effort and cost. Which Amazon SageMaker feature should they use?

A.SageMaker Ground Truth
B.SageMaker Feature Store
C.SageMaker Studio Classic
D.SageMaker Data Wrangler
AnswerA

SageMaker Ground Truth provides managed labelling workflows with automatic data labelling and human review, reducing manual annotation effort and cost. This directly satisfies the requirement to annotate a large image dataset efficiently for object detection.

Why this answer

SageMaker Ground Truth provides labeling workflows with built-in active learning, which automatically selects the most informative images for human review and uses machine learning to label the rest, reducing manual effort and cost.

326
MCQeasy

A company is setting up a data pipeline to ingest streaming clickstream data from their website for real-time analytics and machine learning. The data must be reliably ingested, transformed, and stored in Amazon S3 for batch processing. Which combination of AWS services should be used?

A.Amazon Kinesis Data Analytics to Amazon S3
B.AWS Glue ETL job to Amazon S3
C.Amazon Kinesis Data Firehose to Amazon S3
D.Amazon Kinesis Data Streams to Amazon SageMaker
AnswerC

Kinesis Data Firehose handles the ingestion, optional transformation, and reliable delivery of streaming clickstream data directly into Amazon S3, satisfying the requirement for near-real-time capture with durable batch storage. It needs no custom consumer code, unlike Kinesis Data Streams, which requires separate processing to land data in S3.

Why this answer

Amazon Kinesis Data Firehose is the correct choice because it is a fully managed service designed to reliably ingest streaming data, transform it (e.g., convert to Parquet/ORC, compress, or invoke AWS Lambda for custom transformations), and automatically deliver it to Amazon S3 without requiring custom code or manual scaling. This directly meets the requirement for real-time ingestion, transformation, and storage in S3 for batch processing.

Exam trap

The trap here is that candidates confuse Kinesis Data Streams (a raw streaming layer requiring custom consumers) with Kinesis Data Firehose (a managed delivery service to destinations like S3), often picking Data Streams because it is more commonly discussed, but it does not directly write to S3 without additional components.

How to eliminate wrong answers

Option A is wrong because Amazon Kinesis Data Analytics is used for real-time SQL or Apache Flink-based analytics on streaming data, not for reliably ingesting and storing raw data to S3; it lacks built-in delivery to S3. Option B is wrong because AWS Glue ETL jobs are batch-oriented processing tools that run on a schedule or trigger, not designed for real-time streaming ingestion from a website. Option D is wrong because Amazon Kinesis Data Streams is a real-time data streaming service that requires a separate consumer (e.g., Lambda, Kinesis Data Firehose, or custom application) to write to S3, and Amazon SageMaker is a machine learning platform, not a storage destination for raw clickstream data.

327
MCQmedium

A team is fine-tuning a Hugging Face transformer model on SageMaker. They need to use a custom training script with the Hugging Face Estimator. Which SageMaker feature does this represent?

A.Built-in algorithm
B.SageMaker Autopilot
C.SageMaker Debugger
D.Script mode
AnswerD

Script mode lets the Hugging Face Estimator run a user-supplied Python training script inside the managed container, satisfying the custom training script requirement. The estimator passes hyperparameters and data channels to that entry point, so no prebuilt algorithm image is needed.

Why this answer

Using a custom training script with the Hugging Face Estimator in SageMaker represents Script Mode. Script Mode allows you to bring your own training script and run it within a pre-built framework container, such as Hugging Face, PyTorch, or TensorFlow. This provides flexibility to customize training logic while leveraging SageMaker's managed infrastructure.

Exam trap

MLA-C01 often tests the distinction between built-in algorithms and Script Mode, causing candidates to confuse custom script execution with automated ML features like Autopilot or debugging tools like Debugger.

How to eliminate wrong answers

Option A is wrong because a built-in algorithm refers to SageMaker's pre-built algorithms (e.g., XGBoost, Linear Learner) that do not require custom code. Option B is wrong because SageMaker Autopilot is an automated machine learning feature that automatically builds and tunes models, not a method for running custom scripts. Option C is wrong because SageMaker Debugger is a tool for monitoring and debugging training jobs, not a feature for executing custom training scripts.

328
Multi-Selectmedium

A machine learning team notices an increase in 5XXError count for a SageMaker endpoint. They want to set up automated remediation. Which THREE actions should they take? (Select THREE)

Select 3 answers
A.Add an SNS topic as the alarm action
B.Increase the endpoint instance count manually
C.Create a CloudWatch Alarm on the 5XXError metric
D.Enable detailed monitoring on the endpoint
E.Configure a Lambda function to restart the endpoint or scale out
AnswersA, C, E

An SNS topic as the alarm action provides the notification and integration channel that downstream automation, such as a Lambda function, subscribes to, enabling the automated remediation workflow the team requires when the 5XXError threshold is breached.

Why this answer

Option C is correct because a CloudWatch Alarm on the SageMaker endpoint's 5XXError metric is the foundational detection mechanism that turns the rising error count into an actionable state change. Option A is correct because attaching an SNS topic as the alarm action provides the notification/event distribution channel that can fan out the alert to subscribers or trigger downstream automation. Option E is correct because a Lambda function invoked by the alarm (directly or via SNS) performs the automated remediation, such as restarting the endpoint or scaling out its instance count to restore availability.

Option B is wrong because manually increasing the instance count is a human intervention, not automated remediation, which is what the team requires. Option D is wrong because enabling detailed monitoring only increases metric granularity (e.g., 1-minute CloudWatch metrics) and does not by itself detect-and-remediate the 5XXError increase.

Exam trap

The trap here is that candidates often confuse enabling detailed monitoring (which only increases metric frequency) with automated remediation, or they mistakenly think manual scaling counts as automated remediation.

329
MCQeasy

A data scientist needs to version and manage multiple models for a team of five. The team frequently experiments with different algorithms and hyperparameters. They need a centralized registry to store, deploy, and compare model versions. Which AWS service should the data scientist use?

A.Store each model artifact in Amazon S3 with manual versioning in the key name.
B.Use AWS Config to track model version changes.
C.Use AWS CodeArtifact to store model packages.
D.Use Amazon SageMaker Model Registry.
AnswerD

Amazon SageMaker Model Registry provides a centralised catalogue for versioning, comparing and approving models, and integrates with deployment endpoints. It satisfies the stem's need for one shared registry across five team members experimenting with algorithms and hyperparameters, unlike per-notebook or S3-only storage.

Why this answer

Amazon SageMaker Model Registry is the correct choice because it provides a centralized repository specifically designed for cataloging, versioning, approving, and deploying machine learning models. It integrates natively with SageMaker pipelines and endpoints, enabling the team to compare model versions, manage metadata (e.g., hyperparameters, metrics), and promote models through stages (e.g., from staging to production) with approval workflows.

Exam trap

The trap here is that candidates confuse AWS CodeArtifact (a package manager for code libraries) with a model registry, overlooking that SageMaker Model Registry is purpose-built for ML model versioning, metadata tracking, and deployment orchestration.

How to eliminate wrong answers

Option A is wrong because manual versioning in S3 key names lacks built-in model metadata tracking, approval workflows, and deployment integration, making it error-prone and unscalable for a team of five. Option B is wrong because AWS Config is a service for auditing and evaluating resource configurations (e.g., compliance rules), not for versioning or managing ML model artifacts. Option C is wrong because AWS CodeArtifact is a package management service for software libraries (e.g., Python packages, Maven artifacts), not for storing and versioning trained ML model artifacts or their metadata.

330
MCQmedium

A data scientist suspects that a deep learning model is overfitting. They enable SageMaker Debugger and want to detect overfitting automatically. Which built-in rule should they use?

A.ExplodingGradients
B.PoorWeightInitialization
C.Overfit
D.DeadRelu
AnswerC

The Overfit rule directly satisfies the requirement to detect overfitting automatically. It monitors training and validation loss across steps, raising an issue when validation loss stops decreasing while training loss continues falling. This divergence is the defining signature of overfitting, so it flags the problem without manual threshold tuning.

Why this answer

SageMaker Debugger includes a built-in rule called Overfit that monitors the gap between training and validation loss (or accuracy) and triggers when the model begins to overfit. It is the direct, purpose-built rule for automatic overfitting detection. Enabling it requires the training script to emit both training and validation metrics via the debugger hook.

Exam trap

MLA-C01 often tests the exact built-in rule names for Debugger — candidates may pick a plausible-sounding rule like 'PoorWeightInitialization' or confuse Overfit with a general loss-monitoring rule, but the exam expects the precise 'Overfit' rule.

How to eliminate wrong answers

Option A is wrong because ExplodingGradients detects large gradient magnitudes, which is a training instability issue, not overfitting. Option B is wrong because PoorWeightInitialization detects bad initial weight distributions that hinder convergence, not overfitting. Option D is wrong because DeadRelu detects inactive ReLU neurons, which is an activation pathology, not a generalization gap.

331
MCQmedium

A machine learning engineer is monitoring a deployed model for data drift. The input features are a mix of categorical and numerical columns. The baseline is from the training data. Which SageMaker Model Monitor feature should they enable to detect changes in the distribution of each feature over time?

A.Bias drift monitoring
B.Data quality monitoring
C.Model quality monitoring
D.Feature attribution drift monitoring
AnswerB

Data quality monitoring computes distribution metrics per feature against the training baseline, handling numerical and categorical columns separately, and emits violations when distributions shift. This satisfies the requirement to detect per-feature distribution changes over time.

Why this answer

Data quality monitoring in SageMaker Model Monitor compares the statistical distribution of each input feature (both numerical and categorical) against a baseline computed from the training data, detecting drift in feature distributions over time. It supports categorical and numerical columns and is the correct feature for detecting per-feature distribution changes.

Exam trap

MLA-C01 often tests the confusion between data quality monitoring and model quality monitoring — candidates pick C because 'model' sounds right, but model quality needs ground-truth labels and measures accuracy, while data quality measures input feature distributions.

How to eliminate wrong answers

Option A is wrong because bias drift monitoring (SageMaker Clarify integration) detects changes in bias metrics like disparate impact across groups, not general feature distribution drift. Option C is wrong because model quality monitoring compares model predictions against ground-truth labels to detect accuracy/regression degradation, not input feature distribution changes. Option D is wrong because feature attribution drift monitoring (also Clarify-based) tracks changes in feature importance (SHAP values) relative to baseline, not the raw distribution of feature values.

332
MCQmedium

A machine learning team deploys a model for loan approval. They want to monitor data drift on the real-time endpoint using SageMaker Model Monitor. Which set of actions should they take to set up data quality monitoring?

A.Use SageMaker Clarify to detect data drift on the endpoint
B.Enable data capture on the endpoint, generate a baseline from training data, create a data quality monitoring schedule, and set up a CloudWatch Alarm on violations
C.Create a model quality monitoring schedule directly on the endpoint without any baseline
D.Enable data capture and rely on SageMaker Model Monitor to automatically infer drift without a baseline
AnswerB

Data quality monitoring requires capturing endpoint requests and responses, generating a baseline statistics file from the training dataset, scheduling the monitor against that baseline, and alarming on violations. Together these detect drift in real-time inference data and notify the team via CloudWatch.

Why this answer

SageMaker Model Monitor requires a baseline from training data, then schedules monitoring jobs that compare live endpoint captures against that baseline. Alerts are sent via CloudWatch Alarms.

333
MCQmedium

A company is using SageMaker endpoints for inference. To reduce costs, they want to use Automatic Scaling. However, they observe that scaling up takes several minutes, causing latency spikes during traffic bursts. What should they do to mitigate this?

A.Optimize the model to reduce inference time.
B.Use larger instance types to handle more requests per instance.
C.Configure the endpoint with a target tracking scaling policy and pre-warm additional instances during expected traffic surges.
D.Set the endpoint to scale down slowly to maintain capacity.
AnswerC

Target tracking maintains utilisation near a setpoint, but new instances still need minutes to provision and load the model. Pre-warming instances ahead of predicted surges removes that cold-start delay, so capacity exists before the burst arrives rather than scaling reactively.

Why this answer

It combines a target tracking scaling policy with pre-warming additional instances during expected traffic surges. Pre-warming (or provisioning) instances ahead of time ensures that capacity is available when the burst occurs, mitigating the latency spike caused by the several-minute scaling-up delay. This approach directly addresses the cold-start problem in SageMaker auto-scaling.

Exam trap

AWS often tests the misconception that optimizing model performance or using larger instances can eliminate scaling delays, but the real issue is the provisioning time, which requires proactive capacity management like pre-warming.

How to eliminate wrong answers

Option A is wrong because optimizing the model to reduce inference time does not solve the scaling delay; it only reduces per-request latency, not the time to provision new instances. Option B is wrong because using larger instance types increases per-instance throughput but does not eliminate the provisioning delay when scaling up; the same cold-start latency applies. Option D is wrong because scaling down slowly maintains capacity but does not help during traffic bursts; it may even increase costs without addressing the scaling-up delay.

334
MCQeasy

Which SageMaker built-in algorithm is best suited for detecting anomalous login attempts based on IP addresses and user behavior?

A.XGBoost
B.IP Insights
C.PCA
D.K-Means
AnswerB

IP Insights learns associations between IP addresses and user identities, flagging logins from unusual IP-user pairings. This directly satisfies the scenario's need to detect anomalous login attempts from IP and behaviour patterns, unlike supervised classification algorithms requiring labelled fraud data.

Why this answer

IP Insights is a SageMaker built-in algorithm purpose-built for learning the patterns of IP address usage by users and resources, making it ideal for detecting anomalous login attempts. It uses a neural network to embed IP addresses and entities (like user IDs) into a vector space, flagging unusual combinations as anomalies. XGBoost, PCA, and K-Means are general-purpose algorithms not designed for IP-entity behavioral analysis.

Exam trap

MLA-C01 often tests whether candidates can distinguish purpose-built SageMaker algorithms (like IP Insights for anomaly detection) from general-purpose algorithms (XGBoost, PCA, K-Means) that require custom feature engineering for the same task.

How to eliminate wrong answers

Option A is wrong because XGBoost is a supervised gradient-boosting algorithm for classification/regression on tabular data, not designed to model IP-user behavioral patterns without labeled anomaly data. Option C is wrong because PCA is an unsupervised dimensionality-reduction technique, not an anomaly-detection algorithm for login behavior. Option D is wrong because K-Means is a clustering algorithm that groups similar data points but does not model IP-entity relationships or produce anomaly scores for login attempts.

335
MCQeasy

A company needs to deploy a model that processes large payloads (up to 1 GB) asynchronously. The results should be written to S3, and the team needs SNS notifications upon completion. Which SageMaker inference option is MOST suitable?

A.Asynchronous Inference
B.Batch Transform
C.Real-time endpoint
D.Serverless Inference
AnswerA

Asynchronous Inference queues requests and supports payloads up to 1 GB, writing results to Amazon S3 and publishing completion notifications via Amazon SNS. This directly satisfies the stem's large-payload, asynchronous and notification constraints, unlike real-time endpoints, which cap payloads far smaller.

Why this answer

SageMaker Asynchronous Inference is purpose-built for large payloads (up to 1 GB) and long processing times (up to 15 minutes), queuing requests and writing results to S3. It natively supports SNS notifications on completion or failure, matching the requirement exactly. This makes it the correct fit for asynchronous, large-payload processing with S3 output and SNS alerts.

Exam trap

MLA-C01 often tests the payload/timeout limits of each inference type — candidates who don't memorize the 1 GB/15 min (async), 6 MB/60 s (real-time), and 4 MB/60 s (serverless) limits pick the wrong option.

How to eliminate wrong answers

Option B is wrong because Batch Transform is for offline batch scoring of an entire dataset, not for individual asynchronous requests with per-request SNS notifications. Option C is wrong because Real-time endpoints have a 6 MB payload limit and a 60-second timeout, so they cannot handle 1 GB payloads. Option D is wrong because Serverless Inference has a 4 MB payload limit and a 60-second timeout, and does not natively write results to S3 or send SNS notifications.

336
MCQeasy

A data engineer needs to ingest streaming data from IoT devices into Amazon S3 for machine learning. The data arrives continuously and must be available for querying within minutes. Which service should be used to collect and deliver the streaming data to S3?

A.AWS Database Migration Service (DMS)
B.AWS Glue ETL job triggered by an event
C.Amazon Kinesis Data Firehose
D.Amazon S3 Transfer Acceleration
AnswerC

Amazon Kinesis Data Firehose is a fully managed delivery service that ingests continuous streaming data and loads it into Amazon S3, with configurable buffering that keeps data queryable within minutes. It requires no consumer application to manage.

Why this answer

Amazon Kinesis Data Firehose is the correct choice because it is a fully managed service designed to ingest streaming data in real time and automatically deliver it to destinations like Amazon S3 with near-real-time latency (typically 60 seconds minimum buffering). It handles data transformation, compression, and partitioning without requiring custom code, making it ideal for continuously arriving IoT data that must be queryable within minutes.

Exam trap

The trap here is that candidates confuse AWS Glue (a batch ETL service) with a streaming ingestion tool, or mistakenly think S3 Transfer Acceleration can handle continuous streaming data, when in fact only Kinesis Data Firehose provides the necessary buffering and automatic delivery for near-real-time streaming to S3.

How to eliminate wrong answers

Option A is wrong because AWS Database Migration Service (DMS) is designed for migrating databases to AWS, not for ingesting streaming data from IoT devices. Option B is wrong because AWS Glue ETL jobs are batch-oriented and triggered by events (e.g., S3 PUT), not designed to continuously collect and deliver streaming data in near real time. Option D is wrong because Amazon S3 Transfer Acceleration speeds up uploads over long distances using edge locations but does not provide streaming ingestion or buffering capabilities.

337
MCQmedium

A company uses Amazon SageMaker Feature Store to manage features for real-time inference. They need to store customer transaction data that is updated frequently and must be available for low-latency lookups. Which type of Feature Store should be used?

A.Neither; use S3 directly
B.Both online and offline store
C.Online store only
D.Offline store only
AnswerC

An online store provides low-latency reads for real-time inference, backed by a low-latency store such as Amazon ElastiCache or DynamoDB. Frequent transaction updates are served immediately, satisfying the requirement without the offline store's analytical latency.

Why this answer

The online store is designed for low-latency access, making it suitable for real-time inference lookups. The offline store is used for batch processing and analytics, not for low-latency requirements. Therefore, only the online store is needed.

Option C is correct. Options A, B, and D are incorrect because A bypasses Feature Store, B includes the unnecessary offline store, and D uses only the offline store which lacks low-latency capability.

338
MCQmedium

A company runs an online retail business and wants to build a product recommendation system. They have a dataset of customer purchases stored in Amazon S3 as CSV files. The dataset includes columns: 'customer_id', 'product_id', 'purchase_date', 'quantity', 'price', and 'category'. The data science team plans to use Amazon SageMaker to train a factorization machines model. During data exploration, they discover that the 'category' column has 1,200 unique values, and many categories appear only a few times. The 'product_id' column has 50,000 unique values. They want to include both features in the model. The team is concerned about the high cardinality of these features. Which approach should they take to prepare these features for the factorization machines model?

A.Apply one-hot encoding to both 'product_id' and 'category' columns.
B.Drop the 'category' column and only use 'product_id' since it has more granularity.
C.Encode both columns as integer indices and feed them directly to the factorization machines algorithm as categorical features.
D.Apply principal component analysis (PCA) to reduce the dimensionality of the categorical features.
AnswerC

Factorization machines handle high-cardinality categoricals by learning latent factor vectors per index, so integer-encoding category and product_id directly satisfies the stem's concern. One-hot encoding would explode dimensionality; index encoding lets the algorithm capture interactions between sparse features efficiently.

Why this answer

Amazon SageMaker's factorization machines algorithm natively supports categorical features encoded as integer indices (0-based). This avoids the explosion of features from one-hot encoding (which would create 51,200 columns) and leverages the algorithm's ability to learn interactions between high-cardinality features via factorized parameters, making it both memory-efficient and effective for sparse data.

Exam trap

The trap here is that candidates default to one-hot encoding (Option A) as the standard categorical encoding technique, not realizing that factorization machines are specifically designed to avoid that explosion by accepting raw integer indices as categorical features.

How to eliminate wrong answers

Option A is wrong because one-hot encoding 1,200 categories and 50,000 products would create 51,200 binary columns, causing extreme sparsity and memory blowup, which undermines the factorization machine's efficiency and can lead to poor generalization. Option B is wrong because dropping the 'category' column discards valuable hierarchical information (e.g., product type) that could improve recommendation quality; factorization machines are designed to handle high-cardinality features, so there is no need to drop it. Option D is wrong because PCA is a linear dimensionality reduction technique for continuous features, not suitable for categorical data; applying PCA to one-hot encoded categories would destroy the interpretability of interactions and is not a standard preprocessing step for factorization machines.

339
MCQhard

A company deploys a large NLP model on a SageMaker real-time endpoint using an ml.p3.2xlarge instance. To reduce inference cost without sacrificing throughput, they want to compile the model for their target hardware. Which service should they use?

A.SageMaker Neo
B.Triton Inference Server on SageMaker
C.Amazon Elastic Inference
D.SageMaker Inference Recommender
AnswerA

SageMaker Neo compiles the model for the specific target instance's instruction set, producing an optimised artefact that runs faster on ml.p3.2xlarge. This satisfies the reduce-cost-without-losing-throughput constraint by improving hardware utilisation rather than shrinking the instance.

Why this answer

SageMaker Neo is the correct service because it compiles trained machine learning models into an optimized binary for a specific target hardware (e.g., ml.p3.2xlarge with NVIDIA GPUs). This reduces inference latency and cost by applying hardware-specific optimizations such as kernel fusion and memory layout tuning, while preserving the original model's throughput. The compilation process uses Apache TVM under the hood to generate efficient code for the target instance type.

Exam trap

AWS often tests the distinction between model compilation (Neo) and runtime serving optimizations (Triton) or hardware acceleration (Elastic Inference), leading candidates to confuse a compile-time optimization service with a runtime serving framework or a hardware add-on.

How to eliminate wrong answers

Option B is wrong because Triton Inference Server on SageMaker is a model serving framework that supports multiple backends (e.g., TensorRT, ONNX Runtime) and dynamic batching, but it does not perform ahead-of-time model compilation for a specific hardware target; it optimizes runtime execution, not compile-time optimization. Option C is wrong because Amazon Elastic Inference attaches a separate accelerator to a CPU instance for cost savings, but it is not a model compilation service; it is a hardware attachment that does not compile the model for the target instance. Option D is wrong because SageMaker Inference Recommender is a tool for benchmarking and recommending instance types and endpoint configurations based on load tests, not for compiling or optimizing the model binary for a specific hardware target.

340
Multi-Selecteasy

A company ingests daily log data into an S3 bucket. They need to update the existing ML training dataset with new data without reprocessing the entire history. Which two strategies should they adopt? (Choose two.)

Select 2 answers
A.Store all data in a single large file and use append operations
B.Use AWS Glue to incrementally process new partitions
C.Use a partition key such as date to add new partitions
D.Manually copy new files to the same S3 bucket
E.Overwrite the entire existing dataset with the new data
AnswersB, C

AWS Glue incrementally processes only newly arrived S3 partitions, satisfying the requirement to update the training dataset without reprocessing the entire history. By tracking partition metadata in the Data Catalog, Glue reads just the fresh daily logs, cutting compute cost and runtime compared with full-dataset reprocessing.

Why this answer

Option B is correct because AWS Glue can perform incremental processing by reading only newly added partitions (using job bookmarks and partition predicates) rather than reprocessing the full history, which directly satisfies the requirement to update the dataset without re-running the entire pipeline. Option C is correct because partitioning the S3 data by a key such as date lets new daily log data land in a new partition, so downstream training jobs or Glue ETL can scan only the new partition and append it to the existing dataset. Option A is incorrect because S3 objects are immutable and cannot be appended to in place; a single large file would require rewriting the whole object.

Option D is incorrect because manually copying files into the same bucket does not provide incremental processing logic or partition awareness, and is error-prone. Option E is incorrect because overwriting the entire dataset forces full reprocessing, which is exactly what the scenario wants to avoid.

Exam trap

AWS often tests the misconception that S3 supports append operations or that simply copying new files to the same bucket constitutes an incremental update strategy, when in reality S3 objects are immutable and a proper processing framework like AWS Glue with job bookmarks is required.

341
MCQhard

A healthcare company uses Amazon SageMaker to deploy a real-time inference endpoint for a diagnostic model. The endpoint is configured with a single ml.p3.2xlarge instance. The model processes patient data and returns a risk score. Recently, the endpoint has been experiencing intermittent 504 errors along with increased latency. The team uses Amazon CloudWatch to monitor the endpoint's InvocationsPerInstance and ModelLatency metrics. They observe that InvocationsPerInstance is well below the throttling threshold, but ModelLatency shows periodic spikes lasting 5-10 seconds. The endpoint's CPU utilization remains below 60%, but memory utilization occasionally spikes to 90% during those spikes. The team has checked the inference code and found no obvious memory leaks or performance bottlenecks in the custom logic. The model itself is a deep neural network hosted using Apache MXNet. The team suspects that the issue might be related to resource contention or an external dependency. What should the team do FIRST to diagnose and resolve the issue?

A.Implement request batching to increase throughput and reduce the number of inference requests.
B.Increase the instance type to a more memory-intensive instance like ml.p3.8xlarge to handle memory spikes.
C.Set up SageMaker Model Monitor to track data drift and model quality metrics.
D.Enable SageMaker Debugger rules and profiling to monitor memory and CPU utilization at a fine-grained level during inference.
AnswerB

Increasing to a memory-intensive instance directly addresses the observed memory spikes and associated endpoint timeouts.

Why this answer

The endpoint memory spikes to 90% and correlates with 504 errors and latency spikes, so the team should first resolve the resource pressure by moving to a larger memory instance such as ml.p3.8xlarge. The current correct option D is technically inaccurate: Amazon SageMaker Debugger rules and profiling are for training jobs, not real-time inference endpoints. Option A does not address memory pressure.

Option C is for data drift and model quality, not performance diagnostics.

342
MCQhard

A machine learning engineer is preparing a dataset for training a SageMaker built-in Linear Learner model for binary classification. The dataset contains a highly imbalanced target with only 2% positive examples. They want to improve the model's ability to detect positives without collecting more data. Which SageMaker Linear Learner hyperparameter should they adjust to assign more weight to the positive class?

A.loss
B.mini_batch_size
C.positive_example_weight_mult
D.balance_multiplier
AnswerC

SageMaker Linear Learner includes positive_example_weight_mult, which multiplies the weight of positive examples during training. For imbalanced binary classification, setting this above 1 increases the loss contribution of positives, encouraging the model to detect them better. This directly addresses the scenario without resampling or collecting more data, and it is a documented hyperparameter of the built-in Linear Learner algorithm.

Why this answer

SageMaker Linear Learner provides positive_example_weight_mult to scale the contribution of positive examples in the loss. For a binary target with only 2% positives, increasing this multiplier pushes the model to prioritize positive class recall. Other hyperparameters like loss, mini_batch_size, or nonexistent balance_multiplier do not implement class weighting in this algorithm.

Exam trap

The trap here is assuming any class-weighting parameter works, when SageMaker Linear Learner specifically names it positive_example_weight_mult.

343
MCQmedium

A machine learning model is deployed on SageMaker and its predictions are used in a production application. The model's accuracy has degraded over time. What is the most likely cause?

A.The training data was not shuffled properly.
B.The model was not compiled for inference.
C.The model experienced concept drift.
D.The endpoint instance type is too small.
AnswerC

Concept drift occurs when the statistical relationship between input features and target labels changes in production, so the model's learned mapping no longer matches reality. This degrades accuracy over time even though the model and code are unchanged.

Why this answer

Concept drift occurs when the statistical properties of the target variable change over time, causing the model's predictions to become less accurate. In production ML systems on SageMaker, this is a common issue as real-world data distributions evolve, and the model does not automatically adapt without retraining.

Exam trap

The trap here is that candidates confuse performance degradation due to resource constraints (e.g., instance size) with accuracy degradation caused by data distribution shifts, which is a core concept in ML monitoring.

How to eliminate wrong answers

Option A is wrong because not shuffling training data affects model training convergence and generalization, but it does not cause accuracy to degrade over time after deployment; it is a one-time training issue. Option B is wrong because compiling a model for inference (e.g., using SageMaker Neo) optimizes latency and throughput, not accuracy; it has no impact on prediction quality degradation. Option D is wrong because an endpoint instance type that is too small would cause performance issues like high latency or throttling, not a gradual decline in model accuracy.

344
MCQmedium

A data scientist is training a model on text data from customer reviews. The dataset contains a mix of English and Spanish reviews. The scientist wants to convert the text into numerical features for a classification model. Which approach is MOST appropriate for this multilingual dataset?

A.Translate all Spanish reviews to English using Amazon Translate, then apply TF-IDF on the English text
B.Use pre-trained multilingual word embeddings (e.g., multilingual BERT or FastText) to generate feature vectors
C.Apply TF-IDF separately for English and Spanish reviews, then concatenate the feature matrices
D.Use one-hot encoding on character n-grams for both languages
AnswerB

Multilingual embeddings map English and Spanish tokens into one shared vector space, so semantically equivalent reviews land close together regardless of language. Monolingual embeddings would place the two languages in disjoint regions, forcing the classifier to learn each separately and wasting the mixed-language signal.

Why this answer

Pre-trained multilingual embeddings like multilingual BERT or FastText are trained on corpora spanning many languages and map semantically similar words from different languages into a shared vector space. This allows the classifier to learn from both English and Spanish reviews without translation, preserving language-specific nuance and avoiding translation artifacts. It is the most appropriate approach for a genuinely multilingual dataset.

Exam trap

MLA-C01 often tests the misconception that translation is a required preprocessing step for multilingual NLP, when shared multilingual embedding spaces are the modern, more accurate solution.

How to eliminate wrong answers

Option A is wrong because translating Spanish to English introduces translation errors and latency, and collapses two languages into one, losing native signal and doubling preprocessing cost. Option C is wrong because applying TF-IDF separately per language produces disjoint vocabularies whose concatenated matrices have no shared semantic alignment, so the model cannot generalize across languages. Option D is wrong because one-hot character n-grams create extremely high-dimensional sparse vectors with no semantic similarity between related words, making classification on multilingual text ineffective.

345
Multi-Selecthard

A fraud-detection team runs a SageMaker real-time endpoint in a production account. Their security team requires that the endpoint be reachable only from within a specific Amazon VPC and that access to invoke the endpoint be governed by identity-based policies with least privilege. Which TWO configurations should the ML engineer implement to meet these requirements? (Choose two.)

Select 2 answers
A.Set the endpoint's DataCaptureConfig to capture 100 percent of requests and responses for monitoring.
B.Enable network isolation on the endpoint configuration to block all outbound traffic from the model container.
C.Configure the endpoint to use a customer-managed KMS key for volume encryption and rotate the key annually.
D.Attach an IAM policy to the calling role that allows sagemaker:InvokeEndpoint only for the specific endpoint ARN and denies other SageMaker actions.
E.Create an interface VPC endpoint (AWS PrivateLink) for the SageMaker Runtime service in the VPC and invoke the endpoint through it.
AnswersD, E

An identity-based policy that grants sagemaker:InvokeEndpoint scoped to the specific endpoint ARN enforces least privilege for callers. This restricts which principals can invoke which endpoint and prevents broader SageMaker permissions. It complements the network control by governing authorization independently of the network path.

Why this answer

Private connectivity through an interface VPC endpoint for SageMaker Runtime keeps invocation traffic inside the VPC and off the public internet, while an identity-based IAM policy scoped to the specific endpoint ARN enforces least-privilege authorization. Together they satisfy the network and identity requirements. Network isolation, volume encryption, and data capture address container egress, data-at-rest protection, and observability respectively.

Exam trap

The trap here is assuming that network isolation on the container restricts who can invoke the endpoint, when it only limits the container's outbound traffic.

346
MCQeasy

An ML engineer needs to monitor the operational health of a SageMaker endpoint, specifically the time taken for the container to process an inference request and the overhead added by SageMaker. Which two CloudWatch metrics should they examine?

A.ModelLatency and 4XXError
B.Latency and 5XXError
C.ModelLatency and OverheadLatency
D.Invocations and Latency
AnswerC

ModelLatency captures the time the container spends processing the inference request, while OverheadLatency measures the additional time SageMaker adds for request routing and response handling. Together they decompose total endpoint latency into model and platform components.

Why this answer

ModelLatency is the time taken by the model to respond, and OverheadLatency is the additional time added by SageMaker infrastructure. Invocations is count, not duration; Latency is total latency (ModelLatency + OverheadLatency).

347
MCQeasy

Refer to the exhibit. A data scientist is trying to use AWS Glue to read data from the S3 bucket `ml-data-bucket`. The Glue job fails with an access denied error. What is the most likely cause?

A.The policy allows s3:PutObject but the job only reads
B.The policy does not specify the bucket ARN without /*
C.The Glue job role does not have the required permissions
D.The policy does not include s3:ListBucket permission on the bucket
AnswerD

Without s3:ListBucket, Glue cannot enumerate the bucket's objects, so its read attempt fails with AccessDenied even when s3:GetObject is granted. The stem's constraint is a Glue job reading from `ml-data-bucket`; Glue's crawler and job bootstrap require ListBucket to resolve prefixes before fetching any object.

Why this answer

The error occurs because the IAM policy attached to the Glue job role grants s3:GetObject on the bucket objects (via the `arn:aws:s3:::ml-data-bucket/*` resource) but does not include the s3:ListBucket permission on the bucket itself (`arn:aws:s3:::ml-data-bucket`). When AWS Glue reads data from S3, it first performs a ListBucket operation to enumerate objects in the bucket or prefix, and without that permission, the request is denied even if GetObject is allowed.

Exam trap

AWS often tests the subtle distinction between bucket-level permissions (like s3:ListBucket) and object-level permissions (like s3:GetObject), where candidates assume that granting GetObject on objects is sufficient for reading data, forgetting that listing the bucket is a prerequisite for discovering those objects.

How to eliminate wrong answers

Option A is wrong because the error is an access denied on a read operation, not a write operation; s3:PutObject is irrelevant to reading data. Option B is wrong because the policy does specify the bucket ARN without `/*` for the s3:ListBucket permission (as required), but the issue is that the s3:ListBucket permission itself is missing entirely. Option C is wrong because the Glue job role does have some permissions (as shown in the exhibit), but the specific missing permission is s3:ListBucket, not a general lack of permissions.

348
MCQmedium

A team is evaluating classification models for a medical diagnosis application. The cost of a false negative is much higher than the cost of a false positive. Which metric should be optimized during model selection?

A.Recall
B.Accuracy
C.F1 score
D.Precision
AnswerA

Recall measures the proportion of actual positives correctly identified, so optimising it directly reduces false negatives. In medical diagnosis, where missing a condition is costlier than a false alarm, recall is the metric that satisfies the stem's asymmetric cost constraint.

Why this answer

Recall (sensitivity) measures the proportion of actual positives correctly identified, which directly minimizes false negatives. In medical diagnosis, missing a disease (false negative) is far more costly than a false alarm, so optimizing recall ensures the model captures as many true positive cases as possible.

Exam trap

The trap here is that candidates often default to F1 score as a 'balanced' metric, forgetting that when costs are asymmetric, the metric must reflect the specific business or clinical cost structure, not a generic harmonic mean.

How to eliminate wrong answers

Option B (Accuracy) is wrong because accuracy treats false positives and false negatives equally, which is inappropriate when the cost of false negatives is much higher; a model with high accuracy could still miss many positive cases. Option C (F1 score) is wrong because it balances precision and recall, but when false negatives are far more costly, the optimal trade-off should heavily favor recall over precision, not balance them equally. Option D (Precision) is wrong because precision focuses on minimizing false positives, which is the opposite of the requirement; optimizing precision would reduce false alarms but could increase false negatives.

349
MCQmedium

A machine learning team trains a model in SageMaker and wants to track every step — from dataset version to hyperparameters to final model artifact — for reproducibility and audit compliance. Which SageMaker feature should they use?

A.SageMaker Feature Store
B.SageMaker ML Lineage Tracking
C.SageMaker Experiments
D.SageMaker Model Registry
AnswerB

ML Lineage Tracking automatically captures entities and artefacts across the ML workflow, recording dataset versions, hyperparameters and model artefacts as a queryable graph. This satisfies the reproducibility and audit compliance constraint by preserving end-to-end traceability.

Why this answer

SageMaker ML Lineage Tracking is the correct choice because it is specifically designed to create a directed acyclic graph (DAG) of every step in the ML workflow, including dataset versions, hyperparameters, training jobs, and model artifacts. This enables full reproducibility and audit compliance by capturing the provenance of each entity and their relationships, which is exactly what the question requires.

Exam trap

The trap here is that candidates confuse SageMaker Experiments (which tracks trial metrics and parameters) with ML Lineage Tracking (which captures the full end-to-end provenance graph), leading them to pick Experiments when the question explicitly asks for tracking every step from dataset to final artifact for audit compliance.

How to eliminate wrong answers

Option A is wrong because SageMaker Feature Store is a centralized repository for storing, managing, and sharing features (input data) for ML models, but it does not track the lineage of training steps, hyperparameters, or model artifacts. Option C is wrong because SageMaker Experiments focuses on organizing and comparing multiple training runs (trials) with their parameters and metrics, but it does not automatically capture the full lineage graph connecting datasets, models, and endpoints for audit trails. Option D is wrong because SageMaker Model Registry is a catalog for managing model versions, approvals, and deployments, but it does not track the upstream lineage of how a model was trained (e.g., which dataset version and hyperparameters were used).

350
MCQeasy

A data scientist trained a model using SageMaker and wants to automate the retraining process when new data becomes available. Which AWS service is best suited to trigger a SageMaker training job based on an S3 event?

A.AWS Step Functions with a scheduled trigger.
B.Amazon Simple Workflow Service (SWF) decider.
C.Amazon EventBridge with a rule matching S3 object creation.
D.Amazon Simple Queue Service (SQS) with a polling script.
AnswerC

EventBridge rules can match S3 object-created events and invoke a target that starts the training job, giving event-driven retraining without polling. This satisfies the requirement to trigger SageMaker training automatically whenever new data lands in the bucket.

Why this answer

Amazon EventBridge is the correct choice because it can directly capture S3 events (such as ObjectCreated) via a rule and invoke a SageMaker training job as a target. This event-driven architecture eliminates the need for polling or scheduled checks, enabling immediate retraining when new data arrives in S3.

Exam trap

The trap here is that candidates often confuse scheduled triggers (Step Functions) with event-driven triggers, overlooking that EventBridge is the native AWS service for reacting to S3 events in real time without polling or custom scripts.

How to eliminate wrong answers

Option A is wrong because AWS Step Functions with a scheduled trigger relies on a fixed time interval (e.g., cron expression), not on real-time S3 events, so it cannot react immediately to new data. Option B is wrong because Amazon Simple Workflow Service (SWF) is a legacy workflow orchestration service designed for human-in-the-loop tasks and does not natively integrate with S3 events to trigger SageMaker jobs. Option D is wrong because Amazon Simple Queue Service (SQS) with a polling script requires a separate compute resource (e.g., EC2 or Lambda) to poll the queue and trigger the job, adding latency and complexity compared to EventBridge's direct push-based integration.

351
MCQmedium

A data scientist is using SageMaker Experiments to track multiple training runs. They want to compare the F1 scores across runs. Which component should they use to log the F1 score?

A.Parameter
B.Hyperparameter
C.Artifact
D.Metric
AnswerD

Metrics are the SageMaker Experiments component that logs numeric values such as F1 score against a run, satisfying the requirement to compare F1 scores across runs. They are recorded via the log_metric call within a run.

Why this answer

In SageMaker Experiments, metrics are logged using the SageMaker SDK's log_metric method or by reporting through the training job's metric definitions. Hyperparameters are logged separately. Artifacts are for model files or datasets.

352
Multi-Selecthard

A company is fine-tuning a large language model using reinforcement learning from human feedback (RLHF). Which THREE components are typically required?

Select 3 answers
A.A discriminative classifier
B.A reference model
C.A reward model
D.A policy model (the LLM)
E.A value function
AnswersB, C, D

RLHF needs a frozen reference model to compute the KL-divergence penalty against the fine-tuned policy, preventing reward hacking and keeping outputs close to the original pretrained distribution. This satisfies the requirement for a stable baseline during PPO optimisation.

Why this answer

In the standard RLHF pipeline, the policy model (D) is the large language model being fine-tuned; it generates responses and is updated via reinforcement learning (typically PPO) to maximize reward. A reward model (C) is trained on human preference comparisons to output a scalar score predicting human preference, and it supplies the reward signal that guides the policy's optimization. A reference model (B) is a frozen copy of the pre-RLHF model used to compute a KL-divergence penalty, keeping the policy from drifting too far from the original model and collapsing into degenerate, high-reward outputs.

The other options are not required components: a discriminative classifier (A) is not part of RLHF (the reward model is a regression-style preference predictor, not a classifier), and a value function (E) is only an internal component of the PPO algorithm used to estimate advantages, not a standalone required model in the RLHF architecture.

Exam trap

The trap is including the value function (critic) as a required component because it appears in PPO; the question asks for the three architectural components of RLHF—policy, reward, and reference models.

353
MCQhard

Refer to the exhibit. A SageMaker training job logs show training AUC increasing but validation AUC plateauing at 0.880. What is the most likely issue?

A.Overfitting
B.Learning rate too high
C.Underfitting
D.Insufficient training data
AnswerA

Training AUC keeps rising while validation AUC plateaus at 0.880, meaning the model fits training noise rather than generalisable patterns. This divergence between training and validation curves is the classic signature of overfitting, satisfying the stem's observed symptom.

Why this answer

The training AUC increasing while validation AUC plateaus at 0.880 is a classic sign of overfitting. The model is learning noise and patterns specific to the training data that do not generalize to unseen validation data, causing the validation metric to stall despite continued improvement on the training set.

Exam trap

The AWS ML Engineer Associate exam often tests the distinction between overfitting and underfitting by showing a divergence in training vs. validation metrics, where candidates mistakenly attribute the plateau to a learning rate issue or insufficient data rather than recognizing the hallmark of overfitting.

How to eliminate wrong answers

Option B is wrong because a learning rate that is too high typically causes the loss to diverge or oscillate, not a plateau in validation AUC while training AUC continues to rise. Option C is wrong because underfitting would show both training and validation AUC being low and similar, not a divergence between the two. Option D is wrong because insufficient training data usually leads to high variance and poor generalization, but the specific pattern of training AUC increasing while validation AUC plateaus is more characteristic of overfitting than simply having too little data.

354
MCQeasy

A company deploys a deep learning model to a real-time SageMaker endpoint. After deployment, users report high inference latency. Which action is the MOST effective first step to reduce latency?

A.Switch to a larger instance type with more GPU memory.
B.Compile the model using SageMaker Neo to optimize for the target instance.
C.Enable SageMaker Model Monitor to capture inference data.
D.Increase the number of instances in the endpoint to handle more requests.
AnswerB

SageMaker Neo compiles the model into optimised machine code for the target instance's specific processor architecture, cutting inference latency on the existing endpoint. This directly addresses the reported latency without redeploying to different hardware, making it the most effective first step for the deployed deep learning model.

Why this answer

SageMaker Neo compiles the trained model to optimize it for the target instance hardware, reducing inference latency without requiring additional resources. This is the most effective first step because it directly addresses model execution efficiency, often yielding significant speedups for deep learning models.

Exam trap

The trap here is that candidates often confuse latency reduction with throughput improvement, incorrectly choosing horizontal scaling (Option D) or vertical scaling (Option A) as the first step, when model optimization via compilation is the most direct and cost-effective approach.

How to eliminate wrong answers

Option A is wrong because switching to a larger instance type with more GPU memory may reduce latency if the model is memory-bound, but it is not the most effective first step—it increases cost and does not address software-level inefficiencies. Option C is wrong because SageMaker Model Monitor is used for capturing inference data to detect data drift and model quality issues, not for reducing latency. Option D is wrong because increasing the number of instances (horizontal scaling) improves throughput and handles more concurrent requests, but it does not reduce the latency of individual inference requests; it may even add network overhead.

355
MCQeasy

A company is training a binary classifier in SageMaker and observes that the training loss decreases but validation loss increases after a few epochs. What is the most likely issue?

A.Learning rate too high
B.Overfitting
C.Underfitting
D.Data imbalance
AnswerB

Diverging loss curves — training loss falling while validation loss rises — is the defining signature of overfitting, where the model memorises training noise rather than generalising. The gap widens after several epochs, matching the stem's observation precisely.

Why this answer

The training loss decreasing while validation loss increasing after a few epochs is the classic signature of overfitting. The model is memorizing the training data (including noise) rather than learning generalizable patterns, which causes it to perform poorly on unseen validation data. In SageMaker, this often occurs when the model has too many parameters relative to the dataset size, or when regularization techniques like dropout or L2 weight decay are insufficient.

Exam trap

Candidates often mistake the pattern of decreasing training loss with increasing validation loss for a high learning rate, but a high learning rate would typically cause both losses to oscillate or diverge, not show this specific pattern.

How to eliminate wrong answers

Option A is wrong because a learning rate that is too high typically causes both training and validation loss to diverge or oscillate, not a decrease in training loss with an increase in validation loss. Option C is wrong because underfitting is characterized by both training and validation loss remaining high and not decreasing sufficiently, which is the opposite of the observed behavior. Option D is wrong because data imbalance primarily affects the model's ability to learn the minority class and often manifests as poor recall or precision on that class, not as a divergence between training and validation loss curves.

356
MCQmedium

Which feature scaling method is most robust to outliers in the data?

A.Normalization (L2)
B.Standardization (Z-score)
C.Robust scaling
D.Min-max scaling
AnswerC

Robust scaling centres data using the median and divides by the interquartile range, so extreme values do not distort the transformation. Standardisation and min-max scaling both rely on the mean and range, which outliers skew heavily, making robust scaling the appropriate choice here.

Why this answer

Robust scaling is the most robust to outliers because it uses the median and interquartile range (IQR) instead of the mean and standard deviation. The median and IQR are not influenced by extreme values, so the scaled features remain stable even when the data contains significant outliers.

Exam trap

AWS often tests the misconception that Standardization (Z-score) is robust because it uses standard deviation, but candidates forget that both mean and standard deviation are outlier-sensitive, whereas robust scaling explicitly uses median and IQR.

How to eliminate wrong answers

Option A is wrong because Normalization (L2) scales each sample to have unit Euclidean norm, which does not account for outliers and can be heavily skewed by extreme values in the feature space. Option B is wrong because Standardization (Z-score) uses the mean and standard deviation, both of which are highly sensitive to outliers, causing the scaled values to be distorted. Option D is wrong because Min-max scaling uses the minimum and maximum values, which are directly affected by outliers, leading to compression of the majority of the data into a narrow range.

357
MCQeasy

A company has a dataset of 2 billion records stored as text files in Amazon S3. The data is partitioned by year and month. The data science team wants to read only the last 6 months of data for model training using SageMaker. To minimize data scanned and reduce costs, which approach should the team use?

A.Use S3 Select to retrieve only the last 6 months of data by applying an SQL expression on each object.
B.Use AWS Glue to create a catalog table with partitions, then query with Athena to create a filtered dataset in S3.
C.Use SageMaker Processing with a script that lists all objects in the bucket and reads only those with the desired prefixes.
D.Use SageMaker Processing with Input Mode 'File' and specify the S3 prefix for the last 6 months.
AnswerB

Partition pruning is the mechanism: the Glue catalog exposes year and month partitions, so Athena reads only the six relevant partitions rather than scanning the full 2 billion records, directly minimising bytes scanned and cost.

Why this answer

AWS Glue can crawl the S3 data to create a catalog table with partitions by year and month. Athena can then query only the partitions corresponding to the last 6 months, scanning minimal data and writing the filtered results back to S3 for SageMaker training. This approach leverages partition pruning to reduce costs and avoids loading or processing the full 2 billion records.

Exam trap

AWS often tests the misconception that SageMaker's Input Mode 'File' or S3 Select can efficiently filter partitioned data, but the key trap is that partition pruning requires a catalog service (like Glue) and a query engine (like Athena) to avoid scanning all objects or listing the entire bucket.

How to eliminate wrong answers

Option A is wrong because S3 Select operates on a single object at a time and cannot filter across multiple objects or partitions; applying it to 2 billion records would require iterating over all objects, negating cost savings. Option C is wrong because listing all objects in the bucket and reading only those with desired prefixes still requires enumerating the entire bucket, which incurs significant API costs and does not minimize data scanned (the script must still list all objects). Option D is wrong because SageMaker Processing with Input Mode 'File' downloads the entire dataset to the training instance; specifying a prefix for the last 6 months would still download all files under that prefix, but the data is partitioned by year and month, so using the prefix alone does not guarantee partition pruning—the team would need to explicitly list only the relevant prefixes, which is inefficient compared to Glue+Athena.

358
MCQmedium

A team monitors a production endpoint and notices a sudden increase in 5XXError count. Which of the following is the most likely cause?

A.The endpoint is under-provisioned and requests are throttled
B.The input data format has changed
C.The model container is out of memory or crashing
D.The model is returning predictions with high latency
AnswerC

5XX errors are server-side failures returned by the endpoint itself, so the fault lies in the container rather than the client request. Memory exhaustion or a container crash prevents the model server from completing inference, producing exactly this error class, whereas throttling or bad input would surface as 4XX responses.

Why this answer

A sudden increase in 5XX errors, particularly HTTP 503 or 502, typically indicates that the model container is failing to process requests due to resource exhaustion (e.g., OOM kills) or a crash in the inference process. In a production ML endpoint, such errors often stem from the container running out of memory, leading to the container being terminated by the orchestrator (e.g., Kubernetes OOMKill) or the application crashing internally, which directly causes 5XX responses.

Exam trap

In AWS, 5XX errors on a SageMaker endpoint indicate server-side failures (e.g., container crash, out-of-memory). A common trap is to confuse 5XX errors with client-side errors like throttling (HTTP 429) or input format issues (HTTP 400), but only 5XX errors point to a problem within the model container or inference code.

How to eliminate wrong answers

Option A is wrong because under-provisioning leading to throttling typically results in 429 (Too Many Requests) errors, not 5XX errors; 5XX indicates server-side failures, not rate limiting. Option B is wrong because a change in input data format would likely cause 400 (Bad Request) errors or prediction failures, not a sudden spike in 5XX errors, as the server would reject malformed inputs at the request validation layer. Option D is wrong because high latency does not inherently generate 5XX errors; it may cause timeouts (e.g., 504 Gateway Timeout) if the load balancer or API gateway has a timeout setting, but a general increase in latency alone does not produce a broad 5XX error count unless the container crashes under load.

359
MCQeasy

A data scientist is performing feature selection for a linear regression model and wants to remove features that are highly correlated with each other to reduce multicollinearity. Which technique is BEST suited for this purpose?

A.Correlation analysis
B.Lasso regularization
C.Mutual information
D.Recursive Feature Elimination (RFE)
AnswerA

Correlation analysis directly quantifies pairwise linear dependence between features, so highly correlated pairs can be identified and pruned before fitting. This satisfies the stem's requirement to reduce multicollinearity in linear regression, where correlated predictors inflate coefficient variance. Other techniques address feature importance or dimensionality differently, not pairwise correlation detection.

Why this answer

Correlation analysis directly measures the linear relationship between pairs of features, producing a correlation matrix that identifies highly correlated pairs (e.g., |r| > 0.8) for removal. Since multicollinearity in linear regression is specifically about linear dependence among predictors, correlation analysis is the most direct and interpretable technique for detecting and removing redundant features.

Exam trap

MLA-C01 often tests the confusion between feature-target relevance techniques (mutual information, RFE) and feature-feature redundancy techniques (correlation, VIF) — candidates pick a selection method that optimizes prediction rather than one that detects multicollinearity.

How to eliminate wrong answers

Option B is wrong because Lasso regularization performs feature selection by shrinking coefficients to zero, but it does not explicitly target pairwise correlation and may retain correlated features depending on the penalty strength. Option C is wrong because mutual information measures any (including non-linear) dependence between a feature and the target, not correlation between features, so it does not address multicollinearity. Option D is wrong because RFE removes features based on model performance/importance iteratively, which can handle multicollinearity indirectly but is not specifically designed to detect correlated feature pairs.

360
MCQhard

A team is using SageMaker Pipelines to train a model. The pipeline has multiple steps: data processing, training, evaluation, and registration. They use a Condition step to evaluate the model's accuracy and if it exceeds a threshold, register the model. They run the pipeline and the training step succeeds, but the pipeline fails at the Condition step with an error: 'Unable to evaluate condition: the property 'Accuracy' does not exist.' The evaluation step output is a JSON file with key 'accuracy'. What is the most likely cause?

A.The evaluation step did not produce the output correctly.
B.The training step output is being used instead of the evaluation step output.
C.The pipeline definition has a syntax error.
D.The Condition step is referencing the wrong property name.
AnswerD

The evaluation step writes JSON keyed 'accuracy' (lowercase), but the Condition step queries 'Accuracy'. SageMaker property references are case-sensitive, so the mismatch produces the 'property does not exist' error; aligning the property name resolves it.

Why this answer

The Condition step in SageMaker Pipelines evaluates a property from the output of a previous step. The error 'Unable to evaluate condition: the property 'Accuracy' does not exist' indicates that the Condition step is looking for a property named 'Accuracy' (capital A), but the evaluation step outputs a JSON file with the key 'accuracy' (lowercase a). This mismatch in property name casing causes the condition to fail, even though the evaluation step produced the correct output.

Exam trap

AWS often tests the subtlety of case sensitivity in property names when referencing step outputs in SageMaker Pipelines, leading candidates to incorrectly assume the evaluation step failed or that the pipeline definition has a syntax error.

How to eliminate wrong answers

Option A is wrong because the evaluation step did produce the output correctly (a JSON file with key 'accuracy'), and the pipeline fails only at the Condition step, not at the evaluation step itself. Option B is wrong because the training step output is typically a model artifact, not a JSON property like 'accuracy', and the error specifically mentions the property 'Accuracy' not existing, which points to a naming mismatch rather than using the wrong step's output. Option C is wrong because a syntax error in the pipeline definition would likely cause a different error (e.g., parsing or validation error) at pipeline creation or execution start, not a runtime error at the Condition step that specifically references a missing property.

361
Multi-Selecthard

A company is building a real-time inference pipeline for an ML model. The raw data arrives in JSON format via Amazon Kinesis Data Streams. Before invoking the SageMaker endpoint, the data must be preprocessed to match the training data format. Which THREE steps should be included in the preprocessing function? (Select THREE)

Select 3 answers
A.Ensure that missing values are handled consistently with the training phase
B.Convert the data to a CSV string for model input
C.Apply the same feature engineering transformations (e.g., scaling, encoding) that were used during training
D.Re-train the model periodically using new data
E.Parse the JSON payload
AnswersA, C, E

Handling missing values identically to training preserves the feature distribution the model learned, preventing inference-time skew. Because Kinesis delivers raw JSON that may contain nulls, this step satisfies the requirement that preprocessed data match the training format before the SageMaker endpoint is invoked.

Why this answer

The preprocessing function must first parse the JSON payload (E), because the raw records arrive from Kinesis Data Streams in JSON format and the individual feature fields cannot be accessed until the JSON is decoded into a usable structure. It must also apply the same feature engineering transformations used during training (C), such as scaling and encoding, since the SageMaker endpoint expects inputs in the exact distribution and representation the model learned; using different transformations would cause training-serving skew and degrade predictions. Missing values must be handled consistently with the training phase (A), because imputation or drop logic applied at inference must mirror training-time behavior to keep the feature semantics identical.

Option B is not required because the model's input format is not stated to be CSV, and forcing a CSV string could conflict with the actual serialization the endpoint expects. Option D is incorrect because retraining the model is a separate MLOps activity, not part of the per-record preprocessing function invoked before calling the endpoint.

Exam trap

The trap here is that candidates confuse the preprocessing function's scope with broader MLOps tasks like model retraining, or assume a specific serialization format like CSV is required when JSON is natively supported by SageMaker endpoints.

362
MCQmedium

A team uses SageMaker Clarify to monitor bias drift in production. They schedule weekly analysis. After a month, Clarify reports a significant increase in a bias metric. What should the team do first?

A.Disable the bias monitor because the metric may be noisy.
B.Immediately retrain the model with a balanced dataset.
C.Increase the frequency of analysis to daily.
D.Review the analysis report to understand which feature and segment contributed to the drift.
AnswerD

The report identifies the specific feature and segment driving the metric shift, which is needed before choosing any remediation. Acting on the metric value alone risks misdirected fixes. Reviewing the contributing attributes first establishes the cause the team must address.

Why this answer

When SageMaker Clarify reports a significant increase in a bias metric, the first step is to review the analysis report to identify which feature and segment drove the drift. This diagnostic step informs whether the drift is due to data distribution shift, a specific subgroup, or a false positive. Only after understanding the cause should the team consider retraining or adjusting monitoring.

Exam trap

MLA-C01 often tests the impulse to immediately retrain or disable monitoring — the correct first step is always to investigate the report to understand the root cause before acting.

How to eliminate wrong answers

Option A is wrong because disabling the monitor ignores a potentially real bias issue and violates responsible AI practices. Option B is wrong because immediately retraining with a balanced dataset is premature without knowing the root cause — it may not address the actual drift. Option C is wrong because increasing analysis frequency does not solve the problem; it just provides more data points without diagnosis.

363
MCQhard

A financial services firm is training a fraud detection model using SageMaker. The dataset is highly imbalanced (0.1% fraudulent transactions). The model currently achieves 99.9% accuracy but only catches 5% of fraud cases. Which metric should the team prioritize to evaluate model performance?

A.Accuracy
B.Precision
C.Recall
D.F1-score
AnswerC

Recall measures the proportion of actual fraudulent transactions the model identifies, so it exposes the 5% detection rate that 99.9% accuracy conceals. With 0.1% positives, accuracy is dominated by the majority class, making recall the metric aligned to catching fraud.

Why this answer

Recall is the proportion of actual fraud cases that the model correctly identifies, so it directly measures the model's ability to catch fraud. With only 5% of fraud cases caught, recall is extremely low, and improving it is the priority. Accuracy is misleading here because a model that predicts 'not fraud' for every transaction would still achieve 99.9% accuracy on a 0.1% fraud dataset.

Exam trap

MLA-C01 often tests whether candidates default to accuracy or F1 without considering that in highly imbalanced datasets, recall is the metric that directly reflects the model's ability to catch the minority class.

How to eliminate wrong answers

Option A is wrong because accuracy is dominated by the majority class in an imbalanced dataset and can be 99.9% while missing nearly all fraud, making it a poor metric here. Option B is wrong because precision measures how many predicted frauds are actually fraud, which is important but secondary when the immediate problem is missing 95% of fraud cases. Option D is wrong because F1-score is the harmonic mean of precision and recall and is useful as a balanced metric, but the question specifically asks which metric to prioritize given the low catch rate, and recall is the direct measure of that failure.

364
MCQhard

A data engineer is ingesting streaming clickstream data from a website into Amazon S3 for ML training. The data arrives at a rate of 10,000 events per second, and the team needs near-real-time availability with minimal transformation. Which AWS service should the engineer use to ingest the data into S3 with the LEAST operational overhead?

A.Amazon Kinesis Data Streams with a custom consumer application writing to S3
B.Amazon Kinesis Data Firehose with S3 as the destination
C.AWS Glue ETL job running on a schedule to pull data from a message queue
D.AWS Database Migration Service (DMS) with S3 as target
AnswerB

Amazon Kinesis Data Firehose buffers and delivers streaming records directly into Amazon S3 without managing consumers, shards or custom code, satisfying the 10,000 events per second throughput and near-real-time requirement. Its fully managed, serverless delivery model removes the operational overhead of provisioning and scaling infrastructure that other ingestion services demand.

Why this answer

Amazon Kinesis Data Firehose is a fully managed service that ingests streaming data and delivers it directly to destinations like Amazon S3 with built-in buffering, compression, and optional transformation. It requires no consumer application to write or manage, so operational overhead is minimal. It easily handles 10,000 events per second and provides near-real-time delivery to S3.

Exam trap

MLA-C01 often tests the distinction between Kinesis Data Streams (requires custom consumer) and Kinesis Data Firehose (fully managed delivery to S3), so candidates must recognize that 'least operational overhead' points to Firehose.

How to eliminate wrong answers

Option A is wrong because Kinesis Data Streams requires a custom consumer application (e.g., using KCL) to read from the stream and write to S3, adding significant operational overhead. Option C is wrong because AWS Glue ETL jobs run on a schedule (batch), not near-real-time, and pulling from a message queue adds complexity. Option D is wrong because AWS DMS is designed for database migration and replication, not for ingesting high-volume streaming clickstream data into S3.

365
MCQeasy

A company has deployed a machine learning model on Amazon SageMaker and wants to automatically detect when the distribution of input features deviates significantly from the training data distribution. Which SageMaker feature should they use?

A.SageMaker Clarify
B.SageMaker Edge Manager
C.SageMaker Model Monitor – Model Quality Monitoring
D.SageMaker Model Monitor – Data Quality Monitoring
AnswerD

SageMaker Model Monitor's data quality monitoring compares live inference traffic against a baseline computed from the training dataset, raising CloudWatch alerts when feature distributions drift beyond configured thresholds. This directly satisfies the requirement to detect input feature deviation from the training distribution automatically, without custom code.

Why this answer

SageMaker Model Monitor – Data Quality Monitoring is the correct choice because it is specifically designed to detect deviations in the distribution of input features compared to the training data distribution. It continuously monitors incoming inference requests and compares statistical properties (e.g., mean, variance, or histogram) against a baseline computed from the training dataset, alerting when drift is detected.

Exam trap

The trap here is that candidates often confuse 'Data Quality Monitoring' with 'Model Quality Monitoring', mistakenly thinking that monitoring prediction accuracy covers input distribution drift, whereas Data Quality Monitoring is explicitly for input features and Model Quality Monitoring is for output predictions.

How to eliminate wrong answers

Option A is wrong because SageMaker Clarify is used for bias detection and explainability of model predictions, not for monitoring input feature distribution drift. Option B is wrong because SageMaker Edge Manager manages and optimizes models on edge devices, focusing on deployment and inference at the edge, not on monitoring input data quality in a cloud-based SageMaker endpoint. Option C is wrong because SageMaker Model Monitor – Model Quality Monitoring tracks prediction quality metrics (e.g., accuracy, precision) against a ground truth, not the distribution of input features.

366
Multi-Selecteasy

A company wants to use SageMaker Clarify to analyze bias in their training data and model predictions. Which TWO types of bias can Clarify detect? (Choose TWO.)

Select 2 answers
A.Algorithmic bias
B.Pre-training bias
C.Inference bias
D.Deployment bias
E.Post-training bias
AnswersB, E

Pre-training bias measures imbalance or skew in the training dataset itself, such as label or feature distribution disparities, before any model is fitted. Clarify reports this alongside post-training bias, satisfying the requirement to analyse bias in training data.

Why this answer

SageMaker Clarify organizes its bias metrics into two categories that map directly to the two marked options: pre-training bias and post-training bias. Option B (Pre-training bias) is correct because Clarify analyzes the training dataset itself before model training, computing metrics such as Class Imbalance (CI), Difference in Proportions of Labels (DPL), and Kolmogorov-Smirnov (KS) to detect imbalances or label skew in the input data. Option E (Post-training bias) is correct because Clarify evaluates model predictions after training, using metrics like Disparate Impact (DI), Difference in Conditional Acceptance (DCA), and Accuracy Difference (AD) to detect bias in predicted outcomes across groups.

The remaining options are not Clarify bias categories: algorithmic bias (A) is a general concept rather than a Clarify metric group, while inference bias (C) and deployment bias (D) are not terms Clarify uses to classify the bias it detects.

367
MCQmedium

A data science team is using Amazon SageMaker to train and deploy a binary classification model. They want to continuously monitor the model for data drift in production. Which combination of AWS services and SageMaker features should they use to implement automated drift detection with minimal operational overhead?

A.SageMaker Debugger and Amazon SNS
B.SageMaker Pipelines and AWS Lambda
C.SageMaker Clarify and AWS Config
D.SageMaker Model Monitor and Amazon CloudWatch
AnswerD

SageMaker Model Monitor continuously evaluates production data against baselines to detect drift, publishing metrics and alerts to Amazon CloudWatch. This managed integration requires no custom infrastructure, satisfying the requirement for automated drift detection with minimal operational overhead.

Why this answer

SageMaker Model Monitor is the native SageMaker feature designed specifically for continuously monitoring deployed models for data drift, bias drift, and feature attribution drift. It automatically captures inference requests and responses, computes statistics, and publishes metrics to Amazon CloudWatch, which can trigger alarms for drift detection. This combination provides automated drift detection with minimal operational overhead because it requires no custom infrastructure or manual scheduling.

Exam trap

The trap here is that candidates confuse SageMaker Debugger (training debugging) with SageMaker Model Monitor (production drift detection), or they overcomplicate the solution by adding unnecessary services like Lambda or Config when the native integration with CloudWatch already provides automated alerting.

How to eliminate wrong answers

Option A is wrong because SageMaker Debugger is used for debugging training jobs (e.g., monitoring gradients, weights, and loss during training), not for monitoring data drift in production inference. Option B is wrong because SageMaker Pipelines is a CI/CD orchestration tool for building and managing ML workflows, not a continuous monitoring service; while AWS Lambda could be used to process drift alerts, the core drift detection capability is missing. Option C is wrong because SageMaker Clarify is designed for bias detection and explainability (SHAP values) on datasets or during training, not for real-time drift monitoring of production endpoints; AWS Config tracks resource configuration changes, not model performance or data drift.

368
MCQmedium

A data scientist is using Amazon SageMaker Data Wrangler to prepare a dataset. The dataset contains a column with date strings in the format 'YYYY-MM-DD'. The data scientist wants to extract the year, month, and day as separate features. Which Data Wrangler transform should be used?

A.Encode categorical transform.
B.Scale values transform.
C.Parse date transform.
D.Handle missing transform.
AnswerC

The Parse date transform interprets the 'YYYY-MM-DD' string as a datetime type, from which Data Wrangler can derive year, month and day components. This satisfies the requirement to extract those three separate features without custom code.

Why this answer

The 'Parse date' transform in Amazon SageMaker Data Wrangler is specifically designed to convert date strings into structured datetime components. By applying this transform to the 'YYYY-MM-DD' column, the data scientist can automatically extract year, month, and day as separate features, enabling downstream feature engineering without manual string parsing.

Exam trap

The trap here is that candidates may confuse 'Parse date' with 'Encode categorical' because dates can be treated as categorical features, but the question specifically asks for extracting year, month, and day as separate features, which requires parsing the date string into its components, not encoding the entire date as a category.

How to eliminate wrong answers

Option A is wrong because 'Encode categorical' transform is used to convert categorical variables into numerical representations (e.g., one-hot encoding), not to parse date strings. Option B is wrong because 'Scale values' transform normalizes or standardizes numerical features (e.g., min-max scaling, z-score), which is irrelevant for extracting date components. Option D is wrong because 'Handle missing' transform addresses null or missing values through imputation or deletion, not date parsing.

369
MCQhard

A machine learning engineer is using Amazon SageMaker Data Wrangler to prepare a dataset for a regression model. The dataset contains a categorical feature with high cardinality (over 10,000 unique values). The engineer wants to encode this feature efficiently without creating thousands of binary columns, which would explode the feature space. Which encoding technique should the engineer use in Data Wrangler?

A.Ordinal encoding
B.Binary encoding
C.Target encoding
D.One-hot encoding
AnswerC

Target encoding replaces each category with the mean of the target variable for that category, producing a single numeric feature. This handles high cardinality without expanding the feature space. Data Wrangler supports target encoding, and it is effective for regression tasks when properly regularized to avoid overfitting.

Why this answer

Target encoding is designed for high-cardinality categorical features, replacing each category with a statistic of the target variable. This produces a single feature, avoiding the curse of dimensionality. In SageMaker Data Wrangler, target encoding is available as a transformation and can be configured with smoothing to reduce overfitting, making it the appropriate choice for this regression scenario.

Exam trap

The trap here is assuming that binary encoding is sufficient for high cardinality, but it still creates multiple columns and may not capture the target relationship as effectively as target encoding.

370
MCQhard

A company deploys a model for fraud detection. They need to monitor for bias after deployment, specifically whether the model's false positive rate changes across demographic groups over time. Which SageMaker feature should they use?

A.SageMaker Model Monitor – Model Quality
B.SageMaker Model Monitor – Feature Attribution Drift
C.SageMaker Clarify (post-deployment bias monitoring)
D.SageMaker Model Monitor – Data Quality
AnswerC

SageMaker Clarify's post-deployment bias monitoring continuously evaluates live endpoint traffic, computing metrics such as false positive rate disparity across demographic groups over time. This directly satisfies the requirement to detect bias drift after deployment, which static pre-training analysis cannot address.

Why this answer

SageMaker Clarify provides post-deployment bias monitoring by analyzing predictions against ground truth labels for defined facets. It can track metrics like false positive rate differences over time.

371
MCQmedium

A data scientist is using SageMaker Automatic Model Tuning to optimize hyperparameters for an XGBoost model. They want to maximize AUC. Which search strategy is MOST appropriate for efficient exploration?

A.Random search
B.Grid search
C.Bayesian optimization
D.Hyperband
AnswerC

Bayesian optimization builds a probabilistic surrogate model of the objective and selects hyperparameter configurations that maximise expected improvement, converging in far fewer training jobs than grid or random search. This suits maximising AUC efficiently given Automatic Model Tuning's cost per trial.

Why this answer

Bayesian optimization is the most appropriate search strategy for efficiently exploring hyperparameter space because it builds a probabilistic surrogate model of the objective function (AUC) and uses an acquisition function to intelligently select the next hyperparameter combination to evaluate. This makes it far more sample-efficient than grid or random search, which is critical when each training run is expensive. SageMaker Automatic Model Tuning supports Bayesian optimization as its default and recommended strategy.

Exam trap

The trap is confusing Hyperband's early-stopping efficiency with search efficiency — candidates may pick Hyperband because it sounds 'efficient,' but the question asks for the most appropriate search strategy for exploring hyperparameter space, which is Bayesian optimization.

How to eliminate wrong answers

Option A is wrong because random search, while better than grid search, does not learn from previous trials and therefore wastes compute on unpromising regions of the hyperparameter space — it is less sample-efficient than Bayesian optimization. Option B is wrong because grid search exhaustively evaluates every combination in a predefined grid, which scales exponentially with the number of hyperparameters and is computationally prohibitive for expensive model training. Option D is wrong because Hyperband is a bandit-based early-stopping strategy that prunes poor trials — it is efficient for resource allocation but is not the primary search strategy SageMaker recommends for maximizing a metric like AUC; Bayesian optimization is the standard answer for efficient exploration.

372
Multi-Selecthard

A machine learning engineer is preparing data for a SageMaker training job and needs to split a large dataset into training, validation, and test sets while avoiding data leakage from the same entity appearing in multiple splits. The dataset contains multiple rows per customer, and the target is customer churn. Which TWO strategies are appropriate? (Choose two.)

Select 2 answers
A.Split the data by a hash of the customer ID so that all rows for a given customer land in exactly one split.
B.Duplicate rows for customers in the minority class so that each split contains examples of every customer.
C.Sort the dataset by timestamp and take the first 80 percent of rows for training and the last 20 percent for validation.
D.Use GroupShuffleSplit from scikit-learn with the customer ID as the group parameter to partition customers across splits.
E.Use a random row-level split with a fixed seed, because a fixed seed makes the split reproducible and therefore leakage-free.
AnswersA, D

Hashing the customer ID and assigning splits by hash range guarantees that every row for a customer goes to the same split, which prevents the same entity from appearing in both training and validation. This is a standard grouped-split technique for entity-level leakage and is deterministic and reproducible across reruns when the same hash function and boundaries are used.

Why this answer

Entity-level leakage occurs when rows from the same customer appear in more than one split, letting the model memorize customer-specific patterns and inflating validation metrics. Hashing the customer ID into split buckets and using GroupShuffleSplit with the customer ID as the group both ensure that all rows for a customer stay together, which is the correct way to partition this churn dataset.

Exam trap

The trap here is confusing reproducibility with leakage prevention, since a fixed random seed produces repeatable splits but still scatters a customer's rows across training and validation.

373
MCQhard

A company is building a fraud detection model on credit card transactions. The dataset contains a column 'merchant_id' with 50,000 unique values, many with low frequency. The team wants to avoid overfitting while preserving predictive signal. Which feature engineering approach is most appropriate?

A.Drop the 'merchant_id' column to avoid overfitting
B.Apply label encoding and treat it as a numeric feature
C.Apply target encoding with smoothing based on global mean
D.One-hot encode the 'merchant_id' column
AnswerC

Target encoding replaces each of the 50,000 merchant_id values with a statistic, avoiding the high dimensionality of one-hot encoding. Smoothing blends each category's mean toward the global mean, so low-frequency merchants are not overfitted while their predictive signal is retained.

Why this answer

Target encoding with smoothing replaces each category with a blend of the category's mean target value and the global mean, weighted by frequency. This preserves predictive signal from high-frequency merchants while reducing overfitting for low-frequency ones by shrinking their encoded values toward the global mean. It is a standard technique for high-cardinality categorical features.

Exam trap

MLA-C01 often tests the trade-off between preserving signal and avoiding overfitting with high-cardinality features, tricking candidates into choosing one-hot encoding or dropping the feature instead of using target encoding with smoothing.

How to eliminate wrong answers

Option A is wrong because dropping merchant_id discards potentially valuable signal that could improve fraud detection. Option B is wrong because label encoding assigns arbitrary integers, implying an ordinal relationship that does not exist, which can mislead the model. Option D is wrong because one-hot encoding 50,000 unique values would create a very sparse, high-dimensional dataset, leading to overfitting and computational inefficiency.

374
Multi-Selectmedium

A company is building a recommender system using implicit feedback (clicks) and explicit feedback (ratings). They plan to use Amazon SageMaker to train a model. The data includes user ID, item ID, timestamp, and rating (if any). Which TWO data preparation steps should the team perform? (Choose TWO.)

Select 2 answers
A.Convert user ID and item ID to integer indices for matrix factorization
B.Use target encoding on user ID based on average rating
C.Normalize ratings using StandardScaler
D.One-hot encode user ID and item ID
E.Sort the data by timestamp and use a time-based split for training and validation
AnswersA, E

Matrix factorization algorithms (e.g., in SageMaker's built-in Factorization Machines) require user and item IDs as integers.

Why this answer

Matrix factorization algorithms in Amazon SageMaker (e.g., the built-in Factorization Machines algorithm or the Apache Spark-based collaborative filtering) require user and item identifiers to be converted to contiguous integer indices starting from 0. This is necessary for efficient embedding lookup and to avoid memory blowup from sparse categorical features. SageMaker's implementation expects the input data in recordIO-wrapped protobuf format with integer-encoded user and item columns.

Exam trap

The trap here is that candidates confuse one-hot encoding (which is common in linear models) with the integer indexing required for embedding-based models like matrix factorization, leading them to select Option D instead of Option A.

375
MCQeasy

A machine learning engineer is preparing a dataset for training a model on Amazon SageMaker. The dataset contains numerical features with varying scales, and the engineer wants to ensure that all features contribute equally during training. Which data preparation step should the engineer take?

A.Apply one-hot encoding to all numerical features.
B.Apply feature scaling, such as standardization or normalization, to the numerical features.
C.Remove outliers from the dataset.
D.Apply principal component analysis (PCA) to reduce dimensionality.
AnswerB

Feature scaling transforms numerical features to a similar scale, preventing features with larger magnitudes from dominating the learning process. Standardization (z-score) or normalization (min-max) are common methods. This ensures all features contribute equally, which is the goal stated in the scenario.

Why this answer

Feature scaling adjusts the range or distribution of numerical features so that they are on a similar scale. This is crucial for algorithms that rely on distance calculations or gradient descent, as unscaled features can lead to slow convergence or biased results. Standardization or normalization are appropriate methods to achieve equal contribution from all features.

Exam trap

The trap here is confusing feature scaling with other preprocessing steps like encoding or dimensionality reduction, which do not address varying scales.

Page 4

Page 5 of 9

Page 6

All pages