Courseiva

AWS Certified Machine Learning Engineer Associate MLA-C01 (MLA-C01) — Questions 376–450

665 questions total · 9pages · All types, answers revealed

Page 5

Page 6 of 9

Page 7
376
Multi-Selecteasy

Which TWO actions are recommended best practices for securing an Amazon SageMaker notebook instance? (Select TWO.)

Select 2 answers
A.Use network ACLs to restrict API calls to the SageMaker API.
B.Enable Multi-AZ deployment for the notebook instance.
C.Use AWS KMS to encrypt the notebook instance's storage volume.
D.Associate the notebook instance with a public subnet that has an internet gateway.
E.Disable direct internet access for the notebook instance.
AnswersC, E

AWS KMS encryption protects the notebook instance's attached EBS storage volume at rest, satisfying the requirement to secure stored data, including notebooks and datasets. SageMaker integrates natively with customer managed KMS keys, so encryption applies to the volume without altering the instance's runtime behaviour or interrupting interactive development sessions.

Why this answer

Option C is correct because SageMaker notebook instances store data on an attached EBS volume, and AWS best practices recommend encrypting that storage with an AWS KMS customer managed key to protect data at rest, including notebooks, scripts, and artifacts. Option E is correct because disabling direct internet access forces the notebook instance to route traffic through a VPC (with NAT gateway or VPC endpoints), which prevents unauthorized outbound access and helps protect against data exfiltration while still allowing controlled access to AWS services. Option A is not a recommended best practice for securing a notebook instance because network ACLs are stateless subnet-level controls and do not govern SageMaker API calls, which are authorized via IAM policies and VPC endpoint policies.

Option B is incorrect because Multi-AZ deployment is not a feature of SageMaker notebook instances; they run in a single Availability Zone. Option D is incorrect because placing a notebook instance in a public subnet with an internet gateway exposes it to inbound internet traffic and violates the recommendation to keep notebook instances in private subnets without direct internet access.

Exam trap

The trap here is that candidates often confuse network-level controls (network ACLs) with API-level controls (IAM/VPC endpoints), or they mistakenly think Multi-AZ applies to all AWS services, when in fact it is specific to database and high-availability services.

377
Multi-Selectmedium

A company is training a large NLP model on SageMaker and wants to reduce costs by using Spot Instances. Which TWO configurations should they implement to handle Spot interruptions gracefully?

Select 2 answers
A.Use a single large instance to reduce interruption probability
B.Set `use_spot_instances=True` and `max_wait` in the estimator
C.Increase the `max_run` parameter to allow longer training
D.Use `keep_alive_period` to keep the instance alive after training
E.Enable checkpointing to save model state periodically
AnswersB, E

Setting `use_spot_instances=True` with `max_wait` enables managed Spot training, where SageMaker checkpoints to Amazon S3 and resumes automatically after interruption, within the specified waiting window. This satisfies the stem's requirement to handle interruptions gracefully while cutting costs, since training continues rather than restarting from scratch.

Why this answer

Option B is correct because in the SageMaker SDK estimator you must explicitly set use_spot_instances=True to enable managed Spot training, and max_wait defines the maximum wall-clock time SageMaker will wait for the Spot capacity (including interruptions and restarts), which is required to let training resume after an interruption. Option E is correct because enabling checkpointing (e.g., via checkpoint_s3_uri) periodically saves model state to Amazon S3, so when a Spot instance is reclaimed the job can restart from the last checkpoint instead of from scratch, which is the core mechanism for handling interruptions gracefully. Option A is not correct because using a single large instance does not meaningfully reduce interruption probability and actually increases the cost/impact of a single interruption.

Option C is not correct because increasing max_run only sets the maximum training duration; it does not help the job survive or recover from a Spot interruption. Option D is not correct because keep_alive_period is used for managed warm pools to reduce cold-start latency between jobs, not to preserve training state across Spot interruptions.

Exam trap

MLA-C01 often tests Spot Instance handling, and candidates may confuse max_wait with max_run or overlook the need for checkpointing.

378
Multi-Selectmedium

A company wants to deploy a new model using a canary deployment strategy on SageMaker. Which two actions should they take? (Select TWO.)

Select 2 answers
A.Register both models in the Model Registry with 'Approved' status
B.Use SageMaker Model Monitor to compare model performance
C.Create a new endpoint with two production variants
D.Enable data capture on the endpoint
E.Set the initial traffic weights for the variants (e.g., 95% and 5%)
AnswersC, E

A canary deployment on SageMaker requires a single endpoint hosting two production variants, with traffic initially weighted heavily toward the existing model. Creating that endpoint satisfies the stem's requirement, letting a small percentage of inference requests shift to the new model.

Why this answer

Option C is correct because a SageMaker canary deployment is implemented by creating an endpoint configuration with two production variants (the existing/stable model and the new model) and deploying them behind a single endpoint, which is the foundational mechanism for splitting traffic between versions. Option E is correct because the canary pattern requires assigning initial traffic weights to those variants—typically a small percentage (e.g., 5%) to the new model and the remainder (e.g., 95%) to the stable model—so the new model receives limited live traffic before being promoted. Option A is not required for the deployment itself; Model Registry 'Approved' status is a governance/approval step, not a technical prerequisite for creating canary variants.

Option B is not required because Model Monitor is used for detecting data/quality drift and bias, not for comparing model performance during a canary rollout. Option D is not required because data capture is an optional monitoring feature that logs request/response data, not a necessary action to perform a canary deployment.

Exam trap

MLA-C01 often tests whether candidates confuse the mechanics of canary deployment (two variants + traffic weights) with supporting features like Model Monitor or Model Registry — the trap is selecting monitoring/governance options as if they were required to create the canary.

379
MCQmedium

A team has a SageMaker Pipeline that trains a model and registers it in the Model Registry. They want to automate the deployment of the approved model to a staging environment. Which event-driven approach should they use?

A.Use an SQS queue to store approval messages and have a cron job process them
B.Set up a CloudWatch alarm on the Model Registry's ApprovalStatus metric
C.Use Amazon EventBridge to listen for Model Registry approval events and trigger an AWS Lambda function that deploys the model
D.Configure an AWS Step Functions state machine to poll the Model Registry every minute
AnswerC

Amazon EventBridge subscribes to Model Registry approval state-change events, and a rule invokes a Lambda function that performs the deployment. This event-driven pattern reacts automatically to approval, satisfying the requirement to deploy the approved model to staging without polling.

Why this answer

Amazon EventBridge natively integrates with SageMaker Model Registry and emits events such as 'SageMaker Model Package State Change' when a model package's approval status changes to Approved. An EventBridge rule can match that event pattern and invoke a Lambda function (or Step Functions, CodePipeline, etc.) to deploy the model to staging. This is the canonical event-driven pattern for automating post-approval deployment in SageMaker MLOps workflows.

Exam trap

MLA-C01 often tests whether candidates confuse CloudWatch metrics/alarms with EventBridge events — Model Registry approval is an event, not a metric, so any answer referencing a CloudWatch alarm on ApprovalStatus is a distractor.

How to eliminate wrong answers

Option A is wrong because SQS is a pull-based queue with no native awareness of Model Registry state changes, and a cron job introduces polling latency and unnecessary infrastructure. Option B is wrong because CloudWatch does not expose an 'ApprovalStatus' metric for the Model Registry — approval is an event, not a metric, so an alarm cannot be created on it. Option D is wrong because polling every minute is inefficient, adds cost and latency, and Step Functions is not the recommended trigger mechanism when EventBridge delivers the event directly.

380
MCQmedium

A company is using SageMaker Automatic Model Tuning to optimize a regression model. They want to minimize the root mean squared error (RMSE). The tuner has completed 20 jobs, and the RMSE has plateaued. Which action should the data scientist take to potentially improve the results?

A.Increase the maximum number of training jobs
B.Increase the number of parallel training jobs
C.Decrease the range of hyperparameters to focus on promising areas
D.Switch the objective metric to mean absolute error (MAE)
AnswerC

Narrowing each hyperparameter's search range concentrates the tuner's sampling around regions that previously produced low RMSE, increasing the chance of finding better values. This satisfies the stem's plateaued-after-20-jobs constraint, where broad ranges waste trials on unpromising areas.

Why this answer

When RMSE has plateaued, narrowing the hyperparameter ranges to focus on promising areas can help the tuner explore more finely around good values. This is a common technique in Bayesian optimization to refine the search. Increasing jobs or parallelism may not help if the search space is too broad.

Exam trap

The trap is thinking more jobs or parallelism will always improve results; candidates may not realize that refining hyperparameter ranges is a more effective strategy when progress stalls.

How to eliminate wrong answers

Option A is wrong because simply increasing the maximum number of jobs may not help if the tuner is already stuck in a plateau; it may just waste resources. Option B is wrong because increasing parallel jobs can actually reduce tuning effectiveness due to less sequential learning. Option D is wrong because switching the objective metric changes the goal, not necessarily improving RMSE.

381
Multi-Selectmedium

A data scientist wants to fine-tune a Llama 2 7B model using SageMaker for a text summarization task. The dataset is 10 GB. The budget is limited, so cost efficiency is important. Which THREE steps should the data scientist take? (Choose THREE.)

Select 3 answers
A.Use SageMaker Debugger to reduce training time
B.Use the SageMaker built-in BlazingText algorithm
C.Use LoRA to reduce the number of trainable parameters
D.Use managed spot training
E.Use the SageMaker HuggingFace estimator
AnswersC, D, E

LoRA freezes the base Llama 2 weights and injects small trainable low-rank matrices into attention layers, cutting trainable parameters and optimiser memory dramatically. This directly addresses the limited budget by reducing GPU hours and memory needed for fine-tuning the 7B model.

Why this answer

Option C is correct because LoRA (Low-Rank Adaptation) freezes the base Llama 2 7B weights and injects small trainable low-rank matrices into the attention layers, cutting the number of trainable parameters and GPU memory/compute needed for fine-tuning, which directly supports the limited budget. Option D is correct because SageMaker managed spot training uses spare EC2 capacity at up to a 90% discount and supports checkpointing to S3 so training resumes after interruptions, making it a key cost-efficiency measure. Option E is correct because the SageMaker HuggingFace estimator provides a prebuilt PyTorch/TensorFlow container with the transformers, datasets, and peft libraries needed to fine-tune Llama 2 7B for summarization without building a custom container.

Option A is not correct because SageMaker Debugger is a monitoring and profiling tool for detecting training issues such as vanishing gradients or resource bottlenecks; it does not itself reduce training time. Option B is not correct because BlazingText is a built-in algorithm for text classification and word2vec embeddings, not for fine-tuning large generative LLMs like Llama 2.

Exam trap

MLA-C01 often tests whether candidates can distinguish cost-optimization techniques for LLM fine-tuning (LoRA, spot training, HuggingFace estimator) from unrelated tools (Debugger) or algorithms that cannot handle LLMs (BlazingText), so the trap is picking Debugger thinking it speeds up training.

382
MCQmedium

A machine learning engineer is using SageMaker to train a model and wants to automatically stop a training job when the validation loss has not improved for 10 consecutive epochs, while still saving the best model artifacts. The engineer is using the SageMaker training toolkit in a custom container. Which combination of actions should the engineer take?

A.Configure the estimator's stopping condition with the MaxRuntimeInSeconds parameter and set the checkpoint_s3_uri to save intermediate models.
B.Use SageMaker Automatic Model Tuning with the Bayesian strategy and set the max_parallel_jobs parameter to 1 to enable early stopping across trials.
C.Implement early stopping in the training script by tracking validation loss and calling the SageMaker training toolkit's save_model or exiting the loop when the patience threshold is reached.
D.Enable SageMaker Debugger with the vanishing_gradient rule and configure a stop condition on the rule to halt training when validation loss plateaus.
AnswerC

SageMaker does not provide a built-in early stopping callback for custom training scripts. The engineer must implement patience logic in the script, monitoring validation loss and breaking the training loop when there is no improvement for 10 epochs. Saving the best model before exiting ensures the artifacts reflect the best checkpoint rather than the final epoch.

Why this answer

Early stopping based on validation loss patience is logic that belongs in the training script when using a custom container. The script should evaluate validation loss each epoch, track the best value, and stop after 10 epochs without improvement, saving the best model artifacts. Estimator time limits, tuning strategies, and Debugger rules do not implement this per-epoch behavior.

Exam trap

The trap here is expecting SageMaker to provide a built-in early stopping feature for custom training scripts, when it must be coded explicitly.

383
MCQeasy

A company wants to reduce costs for a real-time inference endpoint that experiences predictable traffic spikes during business hours and low traffic at night. Which auto-scaling policy is MOST cost-effective while maintaining performance?

A.Step scaling based on CPU utilization
B.Manual scaling by the operations team
C.Scheduled scaling that increases instances before business hours and decreases after
D.Target tracking with a custom metric for response time
AnswerC

Scheduled scaling provisions capacity ahead of the known business-hours peak and scales down at night, so instances are never idle during predictable low-traffic periods. This matches the predictable spike pattern directly, unlike reactive target-tracking, which lags demand and over-provisions.

Why this answer

Scheduled scaling directly aligns capacity with the predictable traffic pattern (business hours vs. night), allowing you to proactively add instances before demand increases and remove them afterward. This avoids the cost of over-provisioning during low-traffic periods and the latency of reactive scaling, making it the most cost-effective approach for a known, recurring schedule.

Exam trap

The trap here is that candidates often choose reactive scaling options (like step scaling or target tracking) because they seem 'automated,' but they fail to recognize that for predictable, time-based traffic patterns, scheduled scaling is both more cost-effective and more performant than any reactive policy.

How to eliminate wrong answers

Option A is wrong because step scaling based on CPU utilization is reactive—it only adds capacity after a spike begins, which can cause latency or throttling during the initial surge, and it may keep instances running longer than needed due to cooldown periods, increasing cost. Option B is wrong because manual scaling by the operations team is error-prone, requires 24/7 staffing, and cannot react quickly enough to maintain performance during sudden traffic changes, leading to either over-provisioning or under-provisioning. Option D is wrong because target tracking with a custom metric for response time is also reactive and may cause oscillations (hunting) as the system tries to maintain a target, and it does not leverage the known schedule to pre-emptively scale, resulting in higher costs from delayed or excessive scaling actions.

384
MCQmedium

A healthcare analytics team trains models in SageMaker and stores artifacts in an S3 bucket that contains protected health information. An auditor asks how the team can prove which training dataset and container image produced the model currently deployed to production, and wants the evidence retained even if someone deletes the training job. Which SageMaker capability should the team rely on to capture and retain this metadata automatically?

A.SageMaker Model Registry, which stores model versions and their approval status but does not capture dataset or container provenance automatically.
B.SageMaker Experiments, which groups runs and logs metrics and parameters but does not persist a provenance graph after the trial components are deleted.
C.AWS CloudTrail management events, which record API calls such as CreateTrainingJob and CreateModel and can be queried for who performed each action.
D.SageMaker ML Lineage Tracking, which automatically records entities and associations such as datasets, training jobs, and model artifacts as the workflow runs.
AnswerD

ML Lineage Tracking automatically creates entities and associations for data, training jobs, and models, forming a queryable graph that links the deployed model back to its training dataset and container image. Because the lineage graph is retained independently of the training job's lifecycle, it provides durable evidence for the audit even if the job is deleted, satisfying the reproducibility requirement.

Why this answer

ML Lineage Tracking automatically builds a graph of entities and associations across the ML workflow, linking datasets, training jobs, and model artifacts. This graph persists independently of the training job, so the team can trace the deployed model back to its inputs and container image even after the job is deleted, which is exactly the durable provenance an auditor needs.

Exam trap

The trap here is assuming Model Registry or Experiments captures full provenance automatically, when lineage is the service that records dataset-to-model associations and retains them independently of the training job.

385
Multi-Selectmedium

A data scientist is using SageMaker Experiments to track multiple training runs for a PyTorch model. They want to compare metrics across runs and identify the best hyperparameters. Which TWO capabilities should they use? (Choose TWO.)

Select 2 answers
A.SageMaker Experiments list and search API to query runs by metric
B.SageMaker SDK's experiment logging capabilities
C.SageMaker Autopilot
D.SageMaker Clarify
E.SageMaker Model Monitor
AnswersA, B

The Experiments list and search API lets you query runs programmatically and filter or sort them by logged metric values, so you can retrieve the top-performing runs and compare their hyperparameter configurations. This directly supports identifying the best hyperparameters across many training runs.

Why this answer

Option A is correct because the SageMaker Experiments list and search API (e.g., search() with filters on metric values) lets the data scientist query and compare runs across an experiment by metric, which is exactly what is needed to rank runs and identify the best hyperparameters. Option B is correct because the SageMaker SDK's experiment logging capabilities (Run.log_metric, log_parameter, and log_artifact) record the metrics and hyperparameters for each training run so they can later be analyzed and compared. Option C (SageMaker Autopilot) is an automated machine learning service that builds and tunes models automatically, not a tool for comparing metrics across existing runs.

Option D (SageMaker Clarify) provides bias detection and explainability, and Option E (SageMaker Model Monitor) detects drift in deployed models; neither supports run-to-run metric comparison or hyperparameter selection.

Exam trap

MLA-C01 often tests the confusion between experiment tracking (SageMaker Experiments) and automated model building or monitoring services (Autopilot, Clarify, Model Monitor), so candidates must map each service to its exact purpose.

386
Multi-Selecthard

A data scientist is preparing a dataset for a multi-class classification problem. The dataset contains a categorical feature with 50,000 unique values (high cardinality). The scientist wants to reduce dimensionality while preserving predictive information. Which TWO approaches are appropriate? (Choose 2)

Select 2 answers
A.Target encoding
B.One-hot encoding
C.Count encoding (frequency encoding)
D.Ordinal encoding based on alphabetical order
E.Label encoding
AnswersA, C

Target encoding replaces each of the 50,000 categories with a statistic derived from the target, collapsing the feature to a single numeric column. This reduces dimensionality while retaining predictive signal, though regularisation or cross-fitting is needed to limit target leakage.

Why this answer

Target encoding (Option A) is appropriate because it replaces each of the 50,000 categories with a statistic derived from the target variable (e.g., the mean target value per category), collapsing the feature into a single numeric column and thereby reducing dimensionality while retaining predictive signal about the classes. Count encoding (Option C) is also appropriate because it maps each category to its frequency in the dataset, producing one numeric column that captures useful information (rare vs. common categories) without expanding the feature space. One-hot encoding (Option B) is not suitable here since it would create roughly 50,000 sparse binary columns, which is the opposite of dimensionality reduction.

Ordinal encoding based on alphabetical order (Option D) imposes an arbitrary, meaningless ordering on the categories and does not reduce dimensionality in a way that preserves predictive information. Label encoding (Option E) similarly assigns arbitrary integer IDs to categories, introducing a false ordinal relationship and offering no principled preservation of predictive information.

Exam trap

MLA-C01 often tests the misconception that one-hot or label encoding is suitable for high-cardinality features, when in fact they either explode dimensionality or impose false ordinality, making target and count encoding the correct choices.

387
MCQeasy

A team wants to automatically retrain a model when new labeled data arrives. Which SageMaker feature can orchestrate this workflow?

A.SageMaker Pipelines
B.SageMaker Model Monitor
C.SageMaker Debugger
D.SageMaker Autopilot
AnswerA

SageMaker Pipelines orchestrates the retraining workflow by chaining processing, training and evaluation steps into a directed acyclic graph, triggered automatically when new labelled data lands in Amazon S3 via EventBridge. This satisfies the stem's requirement for automatic, event-driven retraining rather than manual intervention.

Why this answer

SageMaker Pipelines is a purpose-built CI/CD service for machine learning that allows you to define, orchestrate, and automate end-to-end ML workflows, including retraining models when new labeled data arrives. You can create a pipeline that triggers on new data events (e.g., via an S3 event notification or a Lambda function) and automatically executes steps such as data processing, training, evaluation, and model registration. This makes it the correct choice for orchestrating an automated retraining workflow.

Exam trap

The trap here is that candidates confuse monitoring services (Model Monitor, Debugger) with orchestration services, or assume Autopilot's automation includes workflow orchestration, when in fact only Pipelines provides the explicit DAG-based orchestration needed to chain retraining steps on new data events.

How to eliminate wrong answers

Option B (SageMaker Model Monitor) is wrong because it is designed for detecting data drift, model drift, and bias in production, not for orchestrating retraining workflows; it can alert you to drift but cannot automatically trigger a retraining pipeline. Option C (SageMaker Debugger) is wrong because it provides real-time monitoring and debugging of training jobs (e.g., capturing tensors, gradients, and metrics) but has no capability to orchestrate multi-step workflows or trigger retraining. Option D (SageMaker Autopilot) is wrong because it automates the process of building, training, and tuning models from a tabular dataset, but it does not provide a programmable orchestration framework for chaining steps or reacting to new data events.

388
MCQmedium

A team uses SageMaker ML Lineage Tracking to capture the metadata of their ML workflow. They want to query the lineage to see which model version was trained from a specific dataset. Which Lineage Tracking entity represents the dataset?

A.Association
B.Action
C.Context
D.Artifact
AnswerD

In SageMaker Lineage Tracking, an artifact represents a specific versioned object or data resource, such as an S3 dataset version, produced or consumed by a trial component. Querying the lineage graph for the dataset therefore means locating the artifact entity.

Why this answer

In SageMaker ML Lineage Tracking, an Artifact represents a tangible object or data — datasets, models, model versions, and endpoints are all artifacts. The dataset is therefore represented as an Artifact, and lineage queries traverse associations between artifacts, actions, and contexts to trace which model version was trained from it.

Exam trap

MLA-C01 often tests whether candidates can distinguish the four lineage entity types — the common mistake is picking Association or Action because they appear in the lineage graph, but only Artifact represents the actual dataset object.

How to eliminate wrong answers

Option A is wrong because an Association is the relationship (edge) between two lineage entities — for example, the link between a training action and the dataset artifact — not the dataset itself. Option B is wrong because an Action represents an activity or step in the workflow, such as a training job, processing job, or model deployment, not the data consumed by it. Option C is wrong because a Context is a logical grouping of entities, such as an experiment, project, or model package group, used to organize lineage, not to represent a dataset.

389
MCQeasy

A data scientist wants to normalize a feature to have a range between 0 and 1 for a neural network. Which scaling technique should be applied?

A.RobustScaler
B.StandardScaler
C.MaxAbsScaler
D.MinMaxScaler
AnswerD

MinMaxScaler applies the transformation x' = (x - min) / (max - min), linearly mapping each feature onto the [0, 1] interval. This satisfies the stem's explicit 0-to-1 range requirement, unlike StandardScaler which centres on zero with unit variance and produces negative values.

Why this answer

MinMaxScaler scales each feature to a given range, typically [0,1], by subtracting the minimum and dividing by the range. This makes it the correct choice for normalizing data to [0,1]. RobustScaler uses median and IQR, StandardScaler uses mean and standard deviation (resulting in zero mean, unit variance), and MaxAbsScaler scales to [-1,1] by dividing by the maximum absolute value.

Therefore, only MinMaxScaler guarantees a [0,1] range.

390
MCQmedium

A machine learning engineer is preparing a dataset for a binary classification model. The dataset has 10,000 samples with a 1:100 class imbalance. The engineer needs to balance the classes before training. Which technique would create a balanced dataset without discarding majority class samples and without generating synthetic data?

A.Cost-sensitive learning with class weights
B.Random oversampling of the minority class
C.Synthetic Minority Over-sampling Technique (SMOTE)
D.Random undersampling of the majority class
AnswerB

Random oversampling duplicates existing minority class records until both classes are equally represented, satisfying the stem's constraints: no majority samples are discarded and no synthetic data is generated. It differs from SMOTE, which interpolates new synthetic points, and from undersampling, which removes majority data.

Why this answer

Random oversampling of the minority class duplicates existing minority samples until the class distribution is balanced, which increases the minority count without discarding any majority samples and without creating synthetic data. This satisfies all three constraints: balanced dataset, no majority samples removed, and no synthetic generation. It is a straightforward resampling technique commonly used as a baseline before considering more advanced methods.

Exam trap

MLA-C01 often tests the distinction between techniques that balance data (oversampling, undersampling, SMOTE) and techniques that adjust learning (class weights), causing candidates to pick cost-sensitive learning when the question explicitly requires a balanced dataset.

How to eliminate wrong answers

Option A is wrong because cost-sensitive learning with class weights does not create a balanced dataset; it adjusts the loss function to penalize minority misclassification more heavily, leaving the data distribution unchanged. Option C is wrong because SMOTE generates synthetic minority samples by interpolating between existing ones, which violates the 'without generating synthetic data' constraint. Option D is wrong because random undersampling discards majority class samples, which violates the 'without discarding majority class samples' constraint.

391
MCQhard

A data scientist is building a time-series forecasting model for daily sales data. The data spans two years. To evaluate the model's performance, the data scientist needs to simulate a realistic rolling forecast scenario. Which data splitting strategy should be used?

A.Walk-forward validation
B.Random 80/20 train-test split
C.Stratified k-fold cross-validation
D.Hold-out split based on time (e.g., train on first 18 months, test on last 6 months)
AnswerA

Walk-forward validation repeatedly trains on all data up to a point and tests on the immediately following window, then advances that split forward through time. This preserves chronological order and mimics how a deployed forecaster is retrained and rolled forward, unlike random or single-holdout splits.

Why this answer

Walk-forward validation (also called rolling-origin or expanding-window validation) repeatedly trains on past data and tests on the next time period, then moves the window forward. This mimics a realistic rolling forecast where the model is retrained as new data arrives, making it the correct choice for time-series evaluation.

Exam trap

MLA-C01 often tests the data leakage risk of random splits on time-series data, baiting candidates toward standard k-fold or hold-out methods that ignore temporal ordering.

How to eliminate wrong answers

Option B is wrong because a random 80/20 split ignores temporal order, causing data leakage where future data is used to predict the past, inflating performance metrics. Option C is wrong because stratified k-fold is designed for classification with imbalanced classes, not for time-series forecasting, and it also shuffles data, breaking temporal order. Option D is wrong because a single hold-out split based on time evaluates only one forecast horizon and does not simulate the rolling/retraining aspect of a realistic forecasting scenario.

392
MCQeasy

A machine learning engineer needs to standardize features to have zero mean and unit variance before training a support vector machine. Which scaling method should they apply?

A.StandardScaler
B.Normalizer
C.RobustScaler
D.MinMaxScaler
AnswerA

StandardScaler subtracts the feature mean and divides by the standard deviation, producing zero mean and unit variance per feature. That is exactly the transformation requested, and it suits SVMs because they rely on distances and are sensitive to differing feature scales; MinMaxScaler would only bound values to a range.

Why this answer

StandardScaler transforms data to have zero mean and unit variance, which is required for SVM and many other algorithms.

393
MCQeasy

A company has a SageMaker endpoint that uses a trained model to classify images. The endpoint is experiencing high latency and the team suspects it is due to the model size. Which action can the team take to reduce latency without significantly impacting accuracy?

A.Switch to a compute-optimized instance type
B.Use SageMaker Neo to compile the model for the target instance
C.Reduce the batch size of inference requests
D.Convert the model to ONNX format
AnswerB

SageMaker Neo compiles the model into optimised machine code for the target instance family, cutting inference latency through graph-level and operator-level optimisations while preserving accuracy. This directly addresses the model-size-driven latency without retraining or changing the endpoint.

Why this answer

SageMaker Neo compiles trained models into an optimized binary for the target hardware, applying techniques like operator fusion, memory layout optimization, and quantization. This reduces model size and inference latency while preserving accuracy, making it the correct choice for addressing high latency caused by model size.

Exam trap

AWS often tests the misconception that converting to an open format like ONNX inherently optimizes performance, when in reality it is just a serialization format and requires a separate compilation step (e.g., Neo) to reduce latency.

How to eliminate wrong answers

Option A is wrong because switching to a compute-optimized instance (e.g., c5) may improve CPU-bound processing but does not reduce model size or memory footprint; the latency issue stems from the model itself, not insufficient compute. Option C is wrong because reducing batch size can lower throughput and increase per-request overhead, potentially worsening latency; it does not address the root cause of model size. Option D is wrong because converting to ONNX format alone does not guarantee latency reduction; ONNX is an interchange format that requires a compatible runtime (e.g., ONNX Runtime) and may still need optimization like Neo to achieve performance gains.

394
MCQhard

A hospital deploys a model to predict patient readmission risk. To comply with regulations, they must ensure that the model's predictions do not show bias against any demographic group over time. Which service should they use for ongoing monitoring?

A.SageMaker Clarify
B.AWS Audit Manager
C.SageMaker Model Monitor
D.Amazon Macie
AnswerA

SageMaker Clarify provides bias detection with configurable metrics such as demographic parity difference, and its monitoring schedules run continuously against live endpoint traffic, satisfying the requirement for ongoing bias monitoring across demographic groups rather than one-off analysis.

Why this answer

SageMaker Clarify is the correct service because it is specifically designed to detect bias in ML model predictions and can be configured for ongoing monitoring. It provides bias metrics (e.g., difference in positive proportion, disparate impact) and can run on a schedule to continuously evaluate predictions against demographic groups, ensuring regulatory compliance over time.

Exam trap

The trap here is confusing SageMaker Model Monitor (which tracks data drift) with SageMaker Clarify (which tracks bias), leading candidates to choose Model Monitor because they think 'monitoring' covers all aspects of model health, but bias detection requires a separate, specialized tool.

How to eliminate wrong answers

Option B (AWS Audit Manager) is wrong because it is designed to audit AWS resource usage and compliance against frameworks (e.g., SOC 2, PCI DSS), not to monitor ML model bias. Option C (SageMaker Model Monitor) is wrong because it focuses on detecting data drift and feature distribution changes, not bias in predictions against demographic groups. Option D (Amazon Macie) is wrong because it is a data security service that discovers and protects sensitive data using machine learning, not a tool for monitoring model bias.

395
MCQmedium

A data scientist is using SageMaker Data Wrangler to prepare a large dataset. The data contains duplicate rows, which could bias the model. Which built-in step in Data Wrangler can automatically detect and remove duplicates?

A.Amazon QuickSight duplicate detection
B.Handle Duplicates transform in Data Wrangler
C.AWS Glue Studio FindDuplicates transform
D.Amazon DataZone catalog
AnswerB

The Handle Duplicates transform operates directly on the imported dataset within Data Wrangler, detecting and removing duplicate rows without external code. It satisfies the stem's requirement for a built-in step that automatically eliminates duplicates, preventing the bias they would introduce during model training.

Why this answer

The Handle Duplicates transform is a built-in step in SageMaker Data Wrangler specifically designed to detect and remove duplicate rows from a dataset. It provides configurable options such as selecting a subset of columns for duplicate detection and choosing whether to keep the first or last occurrence, directly addressing the bias risk from duplicate rows in ML training data.

Exam trap

The trap here is that candidates confuse AWS Glue Studio transforms (like FindDuplicates) with SageMaker Data Wrangler's built-in steps, as both are AWS data preparation services but operate in different environments and have distinct feature sets.

How to eliminate wrong answers

Option A is wrong because Amazon QuickSight is a business intelligence (BI) service for visualization and dashboards, not a data preparation tool with built-in duplicate detection for ML pipelines. Option C is wrong because AWS Glue Studio FindDuplicates is a transform available in AWS Glue Studio (a separate ETL service), not within SageMaker Data Wrangler's interface or step library. Option D is wrong because Amazon DataZone is a data catalog and governance service for managing data assets across an organization, not a data preparation tool that detects or removes duplicates.

396
MCQhard

A financial services company is developing a fraud detection model using Amazon SageMaker. They have a dataset with 10 million transactions, each with 300 features. The dataset is highly imbalanced (0.1% fraud). They have performed feature engineering and now need to split the data for training, validation, and test sets. The data is stored in CSV files in Amazon S3. They plan to use SageMaker's built-in XGBoost algorithm. To ensure proper evaluation and avoid data leakage, which data splitting strategy should they use?

A.Randomly shuffle the entire dataset and then split into 80% training, 10% validation, 10% test.
B.Use k-fold cross-validation on the entire dataset and average the results.
C.Perform a stratified split on the target variable to ensure each set has the same fraud ratio.
D.Apply SMOTE to balance the dataset first, then split randomly into training, validation, and test sets.
AnswerC

A stratified split preserves the 0.1% fraud ratio across training, validation and test sets, preventing the minority class from being absent or severely under-represented in any split. This satisfies the stem's requirement for proper evaluation of the imbalanced target, since random splitting could yield validation sets with too few fraud cases to assess model performance reliably.

Why this answer

A stratified split preserves the original 0.1% fraud ratio across training, validation, and test sets, which is critical for imbalanced datasets. This ensures each subset is representative of the population, allowing SageMaker's XGBoost to be evaluated fairly without data leakage. Random splits (Option A) could accidentally create a validation or test set with zero fraud cases, making evaluation meaningless.

Exam trap

The trap here is that candidates often choose random splitting (Option A) out of habit, forgetting that imbalanced datasets require stratified sampling to avoid evaluation sets with zero positive cases, which would render metrics like precision and recall undefined.

How to eliminate wrong answers

Option A is wrong because random shuffling and splitting an imbalanced dataset (0.1% fraud) risks producing validation or test sets with no fraud examples, leading to misleading accuracy metrics and inability to detect model overfitting. Option B is wrong because k-fold cross-validation on the entire dataset would leak information from future folds into training when used for final model selection, and it does not provide a held-out test set for unbiased final evaluation. Option D is wrong because applying SMOTE before splitting introduces synthetic data that can leak information across the split boundaries, causing data leakage and overly optimistic performance estimates; SMOTE should only be applied to the training set after splitting.

397
MCQhard

A company runs a regression model to predict house prices. They have 50 features including 'zip_code' (high cardinality), 'square_footage', and 'year_built'. They want to select the most important features to reduce overfitting. Which feature selection method is computationally efficient for high-dimensional data and can handle multicollinearity?

A.Lasso regression (L1 regularization)
B.Mutual information
C.Recursive feature elimination (RFE)
D.Principal component analysis (PCA)
AnswerA

Lasso efficiently selects features by shrinking coefficients to zero.

Why this answer

Lasso regression (L1 regularization) is computationally efficient for high-dimensional data because it performs both feature selection and regularization simultaneously by shrinking less important feature coefficients to zero. It can handle multicollinearity by selecting only one feature from a correlated group, effectively reducing overfitting while maintaining model interpretability.

Exam trap

The AWS ML Engineer Associate exam often tests the distinction between feature selection (keeping original features) and dimensionality reduction (creating new features), so candidates mistakenly choose PCA thinking it handles multicollinearity, but PCA transforms features rather than selecting them, which violates the requirement to 'select the most important features'.

How to eliminate wrong answers

Option B is wrong because mutual information is a filter-based method that measures dependency between features and the target, but it does not inherently handle multicollinearity and can be computationally expensive for high-cardinality features like 'zip_code' without proper binning. Option C is wrong because recursive feature elimination (RFE) is computationally intensive for 50 features as it requires repeatedly training the model and eliminating features one by one, making it inefficient for high-dimensional data. Option D is wrong because principal component analysis (PCA) is a dimensionality reduction technique that creates new orthogonal components, not a feature selection method; it transforms the original features, losing interpretability and not directly selecting the most important original features.

398
Multi-Selecteasy

A team uses SageMaker Ground Truth to create labeled datasets. They need to ensure labeling jobs are cost-effective. Which TWO measures should they take? (Select TWO.)

Select 2 answers
A.Use a smaller instance type for the labeling job.
B.Use a smaller workforce type.
C.Set up a labeling workflow with 'Incremental training'.
D.Enable the 'Consolidated billing' for labeling costs.
E.Use the 'Automated data labeling' feature.
AnswersC, E

Incremental training reuses labels from earlier Ground Truth jobs to pre-annotate subsequent data, reducing human labelling effort and cost. It suits iterative datasets where prior work informs new batches, directly serving the stem's cost-effectiveness requirement.

Why this answer

Option C is correct because SageMaker Ground Truth supports incremental training, which reuses labels and model artifacts from previous labeling jobs so that the auto-labeling model starts from an already-trained state, reducing the amount of human labeling and therefore cost on subsequent jobs. Option E is correct because Automated data labeling (active learning) uses a machine learning model to label a portion of the dataset automatically and only sends low-confidence items to human workers, which lowers the number of human annotations billed. Option A is not correct because instance type affects throughput and speed, not the fundamental per-label cost, and Ground Truth manages the labeling instances.

Option B is not correct because workforce type (public, private, vendor) changes who labels and the price per label, but choosing a 'smaller' workforce type is not a defined cost-optimization measure and does not inherently reduce cost. Option D is not correct because Consolidated billing is an AWS Organizations billing feature for aggregating charges across accounts, not a Ground Truth labeling cost-reduction mechanism.

Exam trap

AWS often tests the misconception that reducing compute instance size or workforce type directly lowers labeling costs, when in reality, Ground Truth costs are driven by the number of human annotations and the use of automated labeling features.

399
MCQhard

A machine learning team is building a model to predict customer churn. They have historical data that includes customer activity logs, each with a timestamp. The team wants to ensure that the training data does not contain any data leakage from the future. Which approach should they take when preparing the training and validation datasets?

A.Use stratified sampling based on churn label
B.Randomly split the data 80/20 for training and validation
C.Use k-fold cross-validation with shuffling
D.Split the data by time, using data before a certain date for training and after for validation
AnswerD

Time-based split ensures no future data influences training.

Why this answer

Splitting by time (chronological split) prevents data leakage by ensuring that the validation set contains only future data relative to the training set. In time-series or timestamped data, random splits can allow the model to learn from future patterns, artificially inflating performance. This approach respects the temporal dependency inherent in customer churn prediction.

Exam trap

AWS often tests the concept of data leakage in time-series contexts, where candidates mistakenly choose random splits or cross-validation with shuffling, overlooking that temporal order must be preserved to avoid future data leaking into training.

How to eliminate wrong answers

Option A is wrong because stratified sampling based on churn label preserves class distribution but does not address temporal leakage; it can still mix future and past data. Option B is wrong because random splitting ignores the timestamp order, allowing future data to leak into the training set and causing the model to learn from events that haven't occurred yet. Option C is wrong because k-fold cross-validation with shuffling randomly reorders the data, which breaks the time sequence and introduces future information into training folds.

400
MCQeasy

A company stores its raw IoT sensor data in Amazon S3. The data is in CSV format and contains timestamps, sensor IDs, and readings. A data engineer needs to catalog this data for discoverability and querying by other team members. Which AWS service should they use to create a searchable metadata catalog?

A.Amazon DynamoDB
B.Amazon Athena data catalog
C.Amazon RDS
D.AWS Glue Data Catalog
AnswerD

AWS Glue Data Catalog provides a centralised, searchable metadata repository storing table definitions, schemas and locations for data in Amazon S3. Crawlers infer CSV structure automatically, making the IoT data discoverable and queryable by other team members through Athena and related services.

Why this answer

The AWS Glue Data Catalog is a managed metadata repository that stores table definitions, schema information, and locations. It integrates with other services like Athena, EMR, and Redshift Spectrum for querying.

401
MCQmedium

A company is using SageMaker to train a neural network for image classification. The training job is taking too long. The team wants to reduce training time without sacrificing model accuracy. Which approach should they recommend?

A.Increase the batch size to the maximum possible
B.Use a GPU-based instance such as ml.p3.2xlarge
C.Use a learning rate scheduler that reduces the learning rate over time
D.Add more convolutional layers to the model
AnswerB

GPU instances such as ml.p3.2xlarge provide massively parallel floating-point throughput, which accelerates the matrix multiplications and convolutions dominating neural network training. This directly satisfies the stem's constraint of reducing training time while preserving model accuracy, since the same algorithm and hyperparameters run unchanged — only the compute hardware differs.

Why this answer

GPU-based instances like ml.p3.2xlarge are specifically designed for parallel processing of matrix operations, which are fundamental to neural network training. By offloading compute-intensive tensor operations to GPU cores, training time can be significantly reduced without altering the model architecture or data, thus preserving accuracy.

Exam trap

AWS often tests the misconception that any change to hyperparameters or architecture can reduce training time without side effects, but the trap here is that candidates confuse 'reducing training time' with 'improving convergence speed'—only hardware acceleration (GPU) directly reduces wall-clock time without risking accuracy degradation.

How to eliminate wrong answers

Option A is wrong because increasing batch size to the maximum possible can lead to degraded model accuracy due to reduced gradient noise, causing the model to converge to sharp minima or even fail to converge; it also risks out-of-memory errors. Option C is wrong because a learning rate scheduler that reduces the learning rate over time helps with convergence stability and final accuracy, but it does not directly reduce training time—it may even extend it if the learning rate becomes too small too early. Option D is wrong because adding more convolutional layers increases model complexity and the number of parameters, which typically increases training time and can lead to overfitting without guaranteeing improved accuracy.

402
MCQhard

A model deployed on SageMaker is returning inaccurate predictions for certain customer segments. The team suspects data drift. Which SageMaker feature should they use to continuously monitor input data distribution?

A.SageMaker Clarify
B.SageMaker Debugger
C.SageMaker Model Monitor
D.SageMaker Feature Store
AnswerC

SageMaker Model Monitor continuously captures endpoint input data and compares its distribution against a baseline, detecting drift via statistical tests such as KL divergence or L-infinity distance. This directly satisfies the requirement to monitor input data distribution across customer segments, triggering CloudWatch alerts when violations exceed thresholds.

Why this answer

SageMaker Model Monitor is the correct choice because it is specifically designed to continuously monitor the input data distribution of a deployed model and detect data drift over time. It automatically captures and analyzes the statistical properties of incoming inference requests against a baseline, alerting you when significant deviations occur.

Exam trap

The trap here is that candidates often confuse SageMaker Clarify's bias detection capabilities with data drift monitoring, but Clarify analyzes static datasets for fairness and explainability, not continuous production data distribution shifts.

How to eliminate wrong answers

Option A is wrong because SageMaker Clarify is used for bias detection and explainability of model predictions, not for monitoring input data distributions over time. Option B is wrong because SageMaker Debugger is designed to debug training jobs by capturing tensors and metrics during training, not to monitor inference data drift in production. Option D is wrong because SageMaker Feature Store is a centralized repository for storing, sharing, and managing features for ML training and inference, not a monitoring tool for data drift.

403
MCQeasy

A team wants to apply a custom container for inference on SageMaker. The container needs to implement a web server that responds to API requests. Which protocol and port must the container listen on to be compatible with SageMaker hosting?

A.The container must listen on port 8080 and use HTTPS protocol.
B.The container must listen on port 8080 and use HTTP protocol.
C.The container can listen on any port as long as the port is specified in the endpoint configuration.
D.The container must listen on port 8000 and use HTTP protocol.
AnswerB

SageMaker hosting requires custom inference containers to run a web server listening on port 8080 over HTTP, which the platform's routing layer expects for health checks and invocations. Listening on any other port breaks the container-to-endpoint contract, so 8080/HTTP satisfies the stem's compatibility constraint.

Why this answer

SageMaker requires custom inference containers to listen on port 8080 and communicate over HTTP (not HTTPS). The SageMaker hosting service uses a proxy that terminates HTTPS and forwards plain HTTP requests to the container on port 8080. This ensures compatibility with the built-in model serving infrastructure.

Exam trap

The trap here is that candidates assume SageMaker requires HTTPS for security, but the service actually handles encryption externally, so the container must use plain HTTP on port 8080.

How to eliminate wrong answers

Option A is wrong because SageMaker's proxy handles TLS termination, so the container must use HTTP, not HTTPS; using HTTPS would cause a protocol mismatch and connection failure. Option C is wrong because SageMaker mandates port 8080 for custom containers; the endpoint configuration does not allow overriding this port. Option D is wrong because the required port is 8080, not 8000; port 8000 is not recognized by SageMaker's hosting proxy.

404
MCQmedium

A startup wants to deploy a containerized ML application that includes both a model inference server and a preprocessing component in the same endpoint. Which SageMaker endpoint type supports running multiple containers?

A.Asynchronous Inference
B.Multi-container endpoint
C.Multi-model endpoint
D.Real-time endpoint
AnswerB

Multi-container endpoints run up to fifteen containers behind one endpoint, letting the inference server and preprocessing component be packaged and invoked together. This satisfies the stem's requirement to host both components within a single endpoint.

Why this answer

SageMaker multi-container endpoints allow you to run up to 15 containers on a single endpoint, with containers invoked in a defined sequence (inference pipeline) or directly. This is the correct choice when a preprocessing component and a model inference server must coexist in one endpoint, since the preprocessing container can transform the request before passing it to the inference container. This pattern is ideal for encapsulating feature engineering with the model.

Exam trap

MLA-C01 often tests the confusion between multi-container endpoints (multiple containers, one endpoint) and multi-model endpoints (multiple models, one container) — candidates who conflate the two pick the wrong option.

How to eliminate wrong answers

Option A is wrong because Asynchronous Inference is a deployment mode for long-running or large-payload requests with queuing, not a mechanism for running multiple containers. Option C is wrong because Multi-model endpoints host many models behind a single container using a shared serving stack — they do not run multiple distinct containers. Option D is wrong because Real-time endpoint is a hosting mode (synchronous, low-latency) that can host a single container or a multi-container pipeline, but 'Real-time endpoint' alone does not describe the multi-container capability being asked about.

405
MCQeasy

A company uses Amazon SageMaker to deploy a real-time inference endpoint. They notice increased latency in predictions during peak hours. Which should they investigate first to address the issue?

A.Review the endpoint auto-scaling policy
B.Check the data labeling job status
C.Modify the training instance type
D.Increase the model artifact size
AnswerA

Peak-hour latency typically indicates insufficient capacity, so reviewing the endpoint auto-scaling policy reveals whether instance counts and target-tracking thresholds keep pace with demand. This satisfies the requirement to investigate the most likely cause first, before examining model artefacts or payload sizes.

Why this answer

Increased latency during peak hours is a classic symptom of insufficient compute capacity to handle the request volume. The first step is to review the endpoint's auto-scaling policy to ensure it is configured to scale out instances proactively or reactively based on a relevant metric like 'SageMakerVariantInvocationsPerInstance'. If the policy has a high cooldown period or a low target metric value, it may not add instances quickly enough, causing requests to queue and latency to spike.

Exam trap

The trap here is that candidates confuse training infrastructure (instance type, artifact size) with inference infrastructure, or assume that data labeling quality affects inference speed, when the immediate cause of peak-hour latency is almost always insufficient endpoint capacity due to misconfigured auto-scaling.

How to eliminate wrong answers

Option B is wrong because data labeling job status has no impact on the runtime performance of a deployed inference endpoint; labeling is a separate offline process. Option C is wrong because modifying the training instance type affects model training time and cost, not the inference endpoint's serving capacity or latency during peak hours. Option D is wrong because increasing the model artifact size would likely increase latency further due to longer load times and larger memory footprint, not reduce it.

406
Multi-Selecthard

A company uses SageMaker to train a model. They want to ensure that training data is encrypted at rest and in transit, and that only authorized users can access the training artifacts. Which three steps should they take? (Choose three.)

Select 3 answers
A.Configure IAM policies to restrict access to SageMaker resources
B.Use SageMaker Model Monitor
C.Use a VPC with private subnets and VPC endpoints
D.Enable S3 server-side encryption for training data
E.Use SageMaker Network Isolation
AnswersA, C, D

IAM policies define who may invoke SageMaker training jobs and read model artefacts in S3, directly satisfying the stem's requirement that only authorised users access training artefacts. Combined with KMS encryption at rest and TLS in transit, least-privilege IAM policies are the access-control mechanism; Microsoft Entra ID is irrelevant here since SageMaker uses AWS IAM natively.

Why this answer

Option A is correct because IAM policies are the mechanism that enforces least-privilege authorization over SageMaker resources, including training jobs, endpoints, and the S3 buckets holding training artifacts, ensuring only authorized users can access them. Option C is correct because placing training jobs in a VPC with private subnets and interface VPC endpoints (AWS PrivateLink) for services like S3 and SageMaker keeps traffic on the AWS private network, protecting data in transit and preventing exposure to the public internet. Option D is correct because enabling S3 server-side encryption (SSE-S3, SSE-KMS, or SSE-C) encrypts training data at rest in the bucket where it is stored and read by the training job.

Option B is not correct because SageMaker Model Monitor detects data drift and quality issues in deployed models; it does not provide encryption or access control for training data. Option E is not correct because SageMaker Network Isolation disables outbound network access for training containers, which is a containment control rather than a mechanism that encrypts data at rest or in transit.

Exam trap

The trap here is that candidates often confuse network isolation (Option E) with encryption or access control, but network isolation only restricts network connectivity, not data encryption or authorization.

407
MCQeasy

A retail company is building a machine learning model to predict customer churn. The data engineering team has extracted customer transaction data from Amazon Aurora and stored it as CSV files in Amazon S3. The data includes customer IDs, transaction amounts, timestamps, and product categories. A data scientist discovers that the dataset contains several missing values in the 'transaction_amount' column for about 15% of the records. The data scientist also notices that the 'customer_id' column has some duplicate entries. The team wants to prepare the data for training a churn model using Amazon SageMaker. The data is approximately 50 GB in size. What should the data scientist do to handle the missing values and duplicates efficiently while preparing the data for training?

A.Use a SageMaker notebook instance with Pandas to load the entire dataset into memory, fill missing values with the median, and drop duplicate customer IDs.
B.Use an AWS Glue ETL job to read the data from S3, apply transformations to fill missing values with the mean or median, and drop duplicate customer IDs, then write the cleaned data back to S3.
C.Drop all records with missing values in the transaction_amount column and remove duplicate customer IDs using an Athena SQL query, then store the result in S3.
D.Use an Amazon EMR cluster with Spark to read the CSV files, impute missing transaction amounts with the mean or median, and remove duplicate customers.
AnswerB

AWS Glue handles the 50 GB scale serverlessly, and its ETL transforms can impute missing transaction_amount values and drop duplicate customer IDs in one pass, writing cleaned output back to S3 for SageMaker training. This satisfies the efficiency constraint that single-node pandas processing would struggle with.

Why this answer

AWS Glue ETL jobs are serverless and designed to handle large-scale data transformations (like 50 GB) without requiring manual cluster management. Glue can read CSV files from S3, apply transformations to impute missing values with the mean or median, drop duplicate customer IDs, and write the cleaned data back to S3, all while scaling automatically to handle the data volume efficiently.

Exam trap

The trap here is that candidates often choose Option A (Pandas in a notebook) because it seems simple, but they overlook the memory limitations of a single-instance notebook when processing 50 GB of data, which is a classic 'scale vs. simplicity' trick in the MLA-C01 exam.

How to eliminate wrong answers

Option A is wrong because loading a 50 GB dataset into memory using Pandas in a SageMaker notebook instance is inefficient and likely to cause out-of-memory errors, as Pandas is single-threaded and not designed for distributed processing of large datasets. Option C is wrong because dropping all records with missing values (15% of data) would discard a significant portion of the dataset, potentially biasing the model, and Athena SQL queries do not natively support imputation of missing values with mean or median without complex workarounds. Option D is wrong because while Amazon EMR with Spark could handle the task, it requires provisioning and managing a cluster, which is more complex and less cost-effective than the serverless AWS Glue approach for this specific data preparation task.

408
MCQeasy

A data scientist is preparing a large dataset for training a machine learning model. The dataset contains missing values in several columns. Which approach is the MOST efficient for handling missing values in a large dataset using AWS services?

A.Use AWS Glue ETL to write a custom Python script that imputes missing values with the mean.
B.Use Amazon SageMaker Data Wrangler to impute missing values using built-in transforms.
C.Use pandas in a SageMaker notebook to impute missing values with the median.
D.Remove all rows with missing values from the dataset.
AnswerB

Data Wrangler provides built-in imputation transforms that run as scalable Spark processing, avoiding custom code for a large dataset. This satisfies the efficiency constraint by handling missing values across many columns in one visual flow, with results exportable directly to SageMaker training.

Why this answer

Amazon SageMaker Data Wrangler provides a visual interface and built-in transforms for handling missing values efficiently at scale, without writing custom code. Glue ETL is more code-heavy, and imputation with pandas is not scalable for large datasets. Removing all rows with missing values is not always optimal and may not be efficient.

409
MCQeasy

A data engineer needs to ingest streaming clickstream data from a website into an S3 data lake for ML training. The data arrives continuously and must be written to S3 in near real-time. Which AWS service is best suited for this task?

A.AWS Lambda function writing to S3 on every click event
B.Amazon Athena queries running on the website's source database
C.Amazon Kinesis Data Firehose with S3 as destination
D.AWS Glue ETL job triggered by a cron job every 5 minutes
AnswerC

Kinesis Data Firehose buffers incoming records and delivers them continuously to S3, providing the near real-time ingestion the clickstream pipeline demands without managing consumers. Managed scaling and native S3 delivery satisfy the continuous-write constraint for ML training data.

Why this answer

Amazon Kinesis Data Firehose is the most appropriate service for loading streaming data into S3 with minimal effort and near-real-time latency. It can buffer, transform, and compress data before delivery.

410
MCQmedium

A machine learning team deploys a custom container image for an Amazon SageMaker training job. The container needs to access an S3 bucket that contains sensitive data. The team wants to follow the principle of least privilege. How should the team grant access?

A.Create an IAM role with S3 access and assign it as the SageMaker execution role for the training job.
B.Attach an IAM instance profile to the training instance with permissions to the bucket.
C.Configure an S3 bucket policy that grants access to the training job's ARN.
D.Store AWS access keys in the container image and use them to access the bucket.
AnswerA

The SageMaker execution role is assumed by the training container, so attaching an IAM role scoped to only the required S3 bucket grants least-privilege access. Credentials are supplied automatically without embedding keys in the custom image.

Why this answer

SageMaker training jobs use an IAM execution role to grant permissions to AWS services like S3. By creating a dedicated IAM role with only the necessary S3 actions (e.g., s3:GetObject, s3:PutObject) and assigning it as the SageMaker execution role, the team follows the principle of least privilege. SageMaker automatically assumes this role via AWS Security Token Service (STS) to access the S3 bucket on behalf of the container, without embedding credentials.

Exam trap

The trap here is that candidates confuse SageMaker's execution role mechanism with EC2 instance profiles, assuming you can attach an IAM role directly to the underlying instance, but SageMaker abstracts instance management and only supports execution roles for granting permissions.

How to eliminate wrong answers

Option B is wrong because SageMaker training jobs do not support attaching an IAM instance profile directly to the training instance; SageMaker manages the underlying EC2 instances and uses the execution role instead. Option C is wrong because a training job does not have an ARN that can be used in an S3 bucket policy; bucket policies grant access to IAM principals (users, roles, accounts) or VPC endpoints, not to job ARNs. Option D is wrong because storing AWS access keys in a container image violates security best practices (e.g., AWS IAM recommends never embedding long-term credentials) and makes key rotation difficult, increasing the risk of exposure.

411
MCQeasy

A company wants to update an existing SageMaker real-time endpoint to serve a new model version. They need to route a small percentage of traffic to the new version initially and monitor for errors before switching fully. Which deployment pattern supports this?

A.Shadow testing
B.A/B testing with traffic splitting
C.Canary deployment with weighted production variants
D.Blue/green deployment
AnswerC

Canary deployment shifts a small percentage of endpoint traffic to the new model variant while the remainder serves the existing version, letting you monitor error metrics before full cutover. Weighted production variants provide exactly this gradual traffic routing.

Why this answer

SageMaker real-time endpoints support canary deployments by configuring multiple production variants with weighted traffic distribution. You can assign a small weight (e.g., 5%) to the new model version variant and 95% to the existing one, then monitor CloudWatch metrics for errors before shifting all traffic to the new variant. This matches the requirement for a gradual, monitored rollout.

Exam trap

Candidates often mistakenly choose blue/green deployment because it sounds like a safe rollout, but it lacks the gradual traffic shifting required for monitoring a small percentage first.

How to eliminate wrong answers

Option A is wrong because shadow testing (also called mirroring) sends a copy of live traffic to the new model without affecting the live response, but SageMaker does not natively support shadow testing for real-time endpoints; it is typically used for testing without routing any user-facing traffic. Option B is wrong because A/B testing with traffic splitting is a broader concept that could be implemented via weighted variants, but the specific pattern described in the question (routing a small percentage of traffic to a new version and monitoring before switching fully) is precisely a canary deployment, not just any A/B test. Option D is wrong because blue/green deployment involves switching all traffic at once from the old (blue) to the new (green) environment, which does not allow for a small percentage of traffic to be routed initially for monitoring.

412
MCQmedium

During model training on Amazon SageMaker, the training job fails with a 'ResourceLimitExceeded' error. What is the most likely cause?

A.The algorithm's learning rate is too high
B.The dataset is too large for the instance
C.The training script has a syntax error
D.The account's instance limit for the chosen instance type has been reached
AnswerD

SageMaker training jobs consume ML compute instances from your account's regional quota. When the requested instance type's concurrent limit is already fully consumed, the job cannot provision capacity and fails immediately with ResourceLimitExceeded, rather than a data, IAM, or algorithm error.

Why this answer

The 'ResourceLimitExceeded' error in Amazon SageMaker indicates that the AWS account has reached its service quota for the specified instance type. Each AWS account has default limits on the number of concurrent instances (e.g., ml.p3.2xlarge) that can be used for training jobs. When a training job requests more instances than the account's limit allows, SageMaker throws this error.

This is distinct from dataset size or algorithmic issues.

Exam trap

AWS often tests the distinction between resource-level errors (quotas) versus data-level or code-level errors; the trap here is confusing a 'ResourceLimitExceeded' error with a dataset size issue or a training script bug, leading candidates to pick Option B or C.

How to eliminate wrong answers

Option A is wrong because a high learning rate would cause training divergence or NaN losses, not a 'ResourceLimitExceeded' error, which is an infrastructure quota issue. Option B is wrong because a dataset too large for the instance would result in an out-of-memory (OOM) error or disk full error, not a resource limit exceeded error; SageMaker would still launch the instance but fail during data loading. Option C is wrong because a syntax error in the training script would produce a Python exception or training job failure with an 'AlgorithmError' or 'ClientError', not a resource quota error.

413
MCQeasy

A team wants to track and compare multiple machine learning experiments, including hyperparameters, metrics, and artifacts. They are using Amazon SageMaker. Which AWS service or feature should they use to achieve this?

A.AWS CloudTrail
B.Amazon SageMaker Experiments
C.Amazon SageMaker Model Registry
D.Amazon SageMaker Studio
AnswerB

SageMaker Experiments groups runs into experiment entities, automatically capturing hyperparameters, metrics, and artifacts for each trial. This lets the team track and compare multiple training runs side by side, directly satisfying the stated requirement to log and contrast experiments.

Why this answer

Amazon SageMaker Experiments is the correct service because it is specifically designed to track and compare machine learning experiments, including hyperparameters, metrics, and artifacts. It provides a structured way to log, organize, and analyze multiple runs, enabling teams to identify the best-performing model configurations.

Exam trap

The trap here is that candidates confuse SageMaker Studio (the IDE) with SageMaker Experiments (the tracking service), assuming Studio alone provides experiment tracking, but Studio is merely the interface that can visualize experiment data stored by Experiments.

How to eliminate wrong answers

Option A is wrong because AWS CloudTrail records API activity for auditing and governance, not for tracking ML experiment metadata like hyperparameters or metrics. Option C is wrong because Amazon SageMaker Model Registry is used for cataloging and managing approved model versions, not for tracking the iterative experiments that produce them. Option D is wrong because Amazon SageMaker Studio is an integrated development environment (IDE) for ML workflows; while it can display experiment data, it is not the service that tracks experiments itself.

414
MCQmedium

An organization wants to ensure that only approved model versions can be deployed to production. They use the SageMaker Model Registry to track model versions. How can they enforce that only approved models are deployed?

A.Manually review each model before deployment
B.Use SageMaker Model Monitor to check model quality after deployment
C.Use IAM policies to restrict deployment to only Approved model versions
D.Store model metadata in a DynamoDB table and check it before deployment
AnswerC

IAM policies can be written to allow SageMaker CreateEndpoint only for models with an Approved approval status, which is best practice.

Why this answer

AWS IAM policies can be used to conditionally restrict SageMaker API actions (e.g., CreateEndpointConfig, CreateModel) based on the model version's approval status. By evaluating the `sagemaker:ModelPackageApprovalStatus` condition key in an IAM policy, you can enforce that only model versions with an `Approved` status can be deployed, providing a native, automated, and auditable enforcement mechanism without manual intervention or external dependencies.

Exam trap

The trap here is that candidates confuse SageMaker Model Monitor (post-deployment monitoring) with pre-deployment approval enforcement, or they assume custom external checks (DynamoDB) are necessary when SageMaker provides native IAM-based conditional enforcement.

How to eliminate wrong answers

Option A is wrong because manual review is not a technical enforcement mechanism; it introduces human error, lacks auditability, and does not prevent unauthorized deployments via API or automation. Option B is wrong because SageMaker Model Monitor is a post-deployment tool that detects data drift and quality issues after the model is already serving traffic; it cannot prevent the deployment of unapproved models. Option D is wrong because storing metadata in DynamoDB and checking it before deployment requires custom code, introduces latency, and is not a native SageMaker enforcement mechanism; it also bypasses the built-in approval tracking in the Model Registry.

415
MCQmedium

A machine learning engineer is training a SageMaker job with the TensorFlow estimator and wants to automatically capture model training metadata such as loss curves and accuracy for later comparison, without writing any custom code. Which SageMaker feature should they enable?

A.SageMaker Model Monitor
B.SageMaker Clarify
C.SageMaker Experiments
D.SageMaker Debugger
AnswerC

SageMaker Experiments automatically captures input parameters, metrics, and artifacts from training jobs when the estimator is created within an experiment context. It logs scalar metrics like loss and accuracy without custom code, enabling comparison across runs. Enabling it on the TensorFlow estimator satisfies the requirement for automatic metadata capture and later comparison.

Why this answer

SageMaker Experiments automatically tracks training parameters, metrics, and artifacts when a training job is launched within an experiment context. It requires no custom code to capture loss and accuracy, and it provides a UI and API to compare runs. Debugger, Model Monitor, and Clarify serve different purposes and do not provide automatic scalar metric logging for experiment comparison.

Exam trap

The trap here is confusing SageMaker Debugger with SageMaker Experiments, because both integrate with training jobs but only one automatically logs scalar metrics for run comparison.

416
Multi-Selectmedium

A company uses Amazon SageMaker to deploy a model for real-time inference. They want to perform A/B testing between two model versions. Which TWO actions should the company take to set up A/B testing? (Choose TWO.)

Select 2 answers
A.Create an endpoint configuration with multiple production variants, each with a different model.
B.Use Amazon CloudWatch Evidently to split traffic between models.
C.Set the initial weight of each production variant to the desired traffic split.
D.Enable auto scaling for each production variant individually.
E.Set the second production variant's weight to 0 and update later to 100.
AnswersA, C

Creating an endpoint configuration with multiple production variants lets each model version receive a defined share of inference traffic, satisfying the A/B testing requirement. SageMaker routes requests across variants according to assigned weights, so you can compare live performance. Deploying both models behind one endpoint is the mechanism that enables controlled traffic splitting.

Why this answer

Option A is correct because SageMaker A/B testing is implemented by creating an endpoint configuration that contains multiple production variants, where each variant references a different model (via its ModelName), allowing the endpoint to serve both model versions behind a single endpoint. Option C is correct because each production variant has an InitialVariantWeight, and setting these weights establishes the desired traffic split (e.g., 80/20) so the endpoint routes the corresponding proportion of invocations to each model. Option B is incorrect because CloudWatch Evidently is a feature-flagging/experimentation service for applications, not the mechanism for splitting traffic between SageMaker production variants.

Option D is incorrect because auto scaling adjusts instance counts for capacity, not traffic distribution between model versions, so it does not set up A/B testing. Option E is incorrect because setting a variant's weight to 0 means it receives no traffic, and shifting to 100 later is a blue/green-style cutover rather than a concurrent A/B test with a defined split.

Exam trap

The trap here is that candidates confuse the separate service Amazon CloudWatch Evidently with SageMaker's native traffic splitting, or think that auto scaling or zero-weight strategies are prerequisites for A/B testing.

417
MCQmedium

A machine learning engineer trains a binary classifier in SageMaker and the model outputs class probabilities. The business requires that the model achieve at least 90% recall on the positive class, while keeping precision above 70%. The engineer uses the default threshold of 0.5 when deploying. Which approach should the engineer take to meet these requirements?

A.Perform a threshold analysis on the precision-recall curve and choose a threshold that yields recall ≥ 90% and precision > 70%.
B.Retrain the model with a higher learning rate and re-evaluate at the default threshold.
C.Use SageMaker Clarify to compute bias metrics and adjust the threshold accordingly.
D.Increase the number of epochs and use early stopping to improve both precision and recall.
AnswerA

The precision-recall curve shows the trade-off between precision and recall at various thresholds. By analyzing this curve on a validation set, the engineer can identify a threshold that meets both the recall and precision requirements. This is the standard approach to select an operating point that aligns with business constraints, and it does not require retraining the model.

Why this answer

The precision-recall curve illustrates how precision and recall vary with the decision threshold. To meet a recall target while keeping precision above a minimum, the engineer should evaluate the curve on a validation set and select a threshold that satisfies both constraints. This approach directly addresses the business requirement without retraining or using bias tools.

Exam trap

The trap here is assuming that retraining or changing hyperparameters will automatically meet specific precision and recall targets, rather than adjusting the decision threshold.

418
MCQeasy

A company wants to maintain multiple versions of a trained model in a central repository and track metadata such as training metrics, hyperparameters, and approval status. Which SageMaker feature should they use?

A.SageMaker Pipelines
B.SageMaker Feature Store
C.SageMaker Model Registry
D.SageMaker Experiments
E.SageMaker Studio
AnswerC

SageMaker Model Registry stores model versions centrally and captures metadata including metrics, hyperparameters and approval status, satisfying the versioning and tracking requirement. It also supports model groups and approval workflows, which generic S3 storage cannot provide.

Why this answer

SageMaker Model Registry is the correct choice because it is specifically designed to serve as a central repository for managing multiple versions of trained models, tracking metadata such as training metrics, hyperparameters, and approval status. It integrates with SageMaker Pipelines and Experiments to automate model governance, enabling versioning, approval workflows, and lineage tracking.

Exam trap

The trap here is that candidates often confuse SageMaker Experiments (which tracks training runs) with the Model Registry (which manages model versions and approvals), leading them to select Experiments when the question explicitly asks for a central repository with versioning and approval workflows.

How to eliminate wrong answers

Option A is wrong because SageMaker Pipelines is a workflow orchestration service for building and automating ML pipelines, not a repository for storing model versions and metadata. Option B is wrong because SageMaker Feature Store is designed for storing, sharing, and managing feature data for training and inference, not for tracking model versions or approval status. Option D is wrong because SageMaker Experiments is used for tracking and comparing training runs, including metrics and hyperparameters, but it does not provide a centralized model registry with versioning and approval workflows.

Option E is wrong because SageMaker Studio is an integrated development environment (IDE) for ML, not a dedicated service for model version management and metadata tracking.

419
MCQhard

Refer to the exhibit. A data scientist configured SageMaker Debugger to monitor training for overfitting. However, the rule never triggers even though the model appears to be overfitting. What is the most likely reason?

A.The debug hook is not collecting the validation loss
B.The instance type for the rule is too small
C.The S3 output path is not writable
D.The rule evaluator image is incorrect
AnswerA

SageMaker Debugger's overfitting rule compares training loss against validation loss, so it needs both tensors collected via the debug hook. If the hook only captures training loss, the rule has no validation signal to evaluate and never fires, even when the model genuinely overfits.

Why this answer

SageMaker Debugger monitors training by collecting tensors (e.g., loss, accuracy) via a debug hook. The rule for detecting overfitting typically compares training loss to validation loss. If the hook is not configured to collect validation loss tensors, the rule has no data to evaluate and will never trigger, even if overfitting occurs.

This is the most likely reason because the rule depends on specific tensor names being saved.

Exam trap

AWS often tests the misconception that a SageMaker Debugger rule not triggering is due to infrastructure issues (instance size, permissions) rather than a missing data collection configuration, leading candidates to overlook the debug hook's tensor registration.

How to eliminate wrong answers

Option B is wrong because the instance type for the rule affects only the compute resources for running the rule evaluation, not the collection of tensors; a small instance may slow evaluation but does not prevent the rule from triggering if data is present. Option C is wrong because if the S3 output path were not writable, the training job itself would fail with a permissions error, not silently skip rule triggering. Option D is wrong because the rule evaluator image is managed by SageMaker and is automatically matched to the built-in rule; an incorrect image would cause a runtime error, not a silent failure to trigger.

420
MCQhard

A financial services company uses Amazon SageMaker to deploy a fraud detection model for real-time inference. The model is deployed on an ml.m5.large instance with a SageMaker real-time endpoint. The endpoint has an auto scaling policy configured using a custom scaling policy based on average CPU utilization, with scale out threshold at 70% and scale in threshold at 30%. During a flash sale event, the traffic to the endpoint spikes tenfold within minutes. The endpoint fails to handle the load, resulting in increased latency and timeouts. The data science team needs to improve the scalability of the endpoint to handle sudden traffic spikes. Which solution should the team implement?

A.Implement a SageMaker Model Ensemble with two additional models to balance the load.
B.Replace the custom scaling policy with a target tracking scaling policy based on the number of invocations per instance, with a target value of 1000.
C.Implement a SageMaker Inference Pipeline with a pre-processing step to reduce model input size.
D.Switch to a GPU instance type, such as ml.p3.2xlarge, to increase compute capacity.
AnswerB

Invocation-based target tracking scales on request volume, the metric that actually surges tenfold during flash sales, whereas CPU lags behind sudden concurrency spikes. It satisfies the sudden-traffic requirement by adding instances before latency degrades, rather than reacting to already-saturated compute.

Why this answer

A target tracking scaling policy based on invocations per instance directly ties scaling to the actual workload metric for a real-time endpoint, allowing faster and more accurate scale-out during sudden traffic spikes. Unlike CPU-based custom scaling, invocation-based target tracking reacts to request volume, which is the true driver of load for inference. This improves the endpoint's ability to handle a tenfold spike.

Exam trap

MLA-C01 often tests the misconception that CPU-based scaling is always sufficient, when for inference endpoints invocation-based target tracking is more responsive to traffic spikes.

How to eliminate wrong answers

Option A is wrong because a SageMaker Model Ensemble runs multiple models and does not increase the endpoint's ability to scale; it adds compute overhead and does not address the scaling policy. Option C is wrong because an Inference Pipeline with pre-processing reduces input size but does not improve the endpoint's scaling behavior or capacity during spikes. Option D is wrong because switching to a GPU instance increases per-instance compute but does not solve the scaling policy's inability to react quickly to sudden traffic increases; it also may not be cost-effective.

421
Multi-Selectmedium

A company is building a real-time fraud detection system using Amazon Kinesis Data Streams. The data must be joined with a reference table (e.g., customer profile) that is stored in Amazon DynamoDB and updated frequently. The enriched data will be used for ML predictions. Which THREE AWS services should the company use to build this streaming pipeline? (Select THREE.)

Select 3 answers
A.Amazon Kinesis Data Analytics for Apache Flink
B.Amazon Kinesis Data Firehose
C.Amazon Kinesis Data Streams
D.AWS Glue ETL
E.Amazon DynamoDB
AnswersA, C, E

Kinesis Data Analytics for Apache Flink performs the stateful stream join between incoming transactions and the DynamoDB reference table, then emits enriched records for ML inference. It satisfies the real-time enrichment requirement by running continuous SQL or Flink queries over streaming data.

Why this answer

Amazon Kinesis Data Streams (C) is the ingestion backbone for the real-time fraud events, providing durable, low-latency, ordered streaming with shards and configurable retention so downstream consumers can process records continuously. Amazon Kinesis Data Analytics for Apache Flink (A) is the correct processing engine because it supports stateful stream processing and can perform asynchronous lookups against DynamoDB (via the Flink DynamoDB connector or async I/O) to enrich each event with the frequently updated customer profile before feeding ML predictions. Amazon DynamoDB (E) is the reference table store, offering single-digit-millisecond key-value reads at scale so the Flink job can join each streaming record with current customer profile data.

Amazon Kinesis Data Firehose (B) is not appropriate here because it is a managed delivery service for loading data into destinations like S3, Redshift, or Splunk and does not perform stateful joins or ML enrichment. AWS Glue ETL (D) is a batch/serverless Spark-based ETL service and is not designed for sub-second, continuously running stream joins against a rapidly changing DynamoDB table.

Exam trap

MLA-C01 often tests whether candidates pick Firehose for 'real-time' processing when the requirement involves joins or enrichment — Firehose is delivery, not compute, and cannot join streams to DynamoDB.

422
MCQeasy

A company is using Amazon SageMaker to train a model on sensitive customer data. The security team requires that all data be encrypted in transit and at rest, and that the training job does not have internet access. Which configuration should the team use to meet these requirements?

A.Configure the training job to run in a public subnet with a security group that blocks outbound traffic
B.Configure the training job to run in a private subnet, but disable encryption to reduce latency
C.Configure the training job to run in a private subnet with no internet access, and use a KMS key for encryption
D.Configure the training job to run in a VPC with a NAT gateway, and use default SageMaker encryption
AnswerC

Isolating the training job in a private subnet removes the internet route entirely, satisfying the no-internet-access constraint, while a KMS key encrypts the training data and volumes at rest. Together they meet both the in-transit and at-rest encryption requirements.

Why this answer

Running the SageMaker training job in a private subnet with no internet access ensures the job cannot reach the public internet, satisfying the no-internet-access requirement. Using an AWS KMS key for encryption at rest (for the S3 bucket and EBS volumes) and enforcing encryption in transit (via HTTPS/TLS for SageMaker and S3 endpoints) meets the encryption requirements. SageMaker training jobs in a private subnet use VPC endpoints (e.g., S3 and SageMaker API endpoints) to communicate securely without internet access.

Exam trap

The trap here is that candidates often confuse a private subnet with a NAT gateway as providing no internet access, but a NAT gateway actually enables outbound internet connectivity, which violates the requirement.

How to eliminate wrong answers

Option A is wrong because a public subnet inherently provides internet access via an internet gateway, violating the no-internet-access requirement; blocking outbound traffic with a security group does not prevent the instance from having a public IP or being reachable from the internet. Option B is wrong because disabling encryption violates the requirement that all data be encrypted in transit and at rest; encryption does not inherently increase latency in a meaningful way for SageMaker training jobs. Option D is wrong because a NAT gateway provides outbound internet access for instances in a private subnet, which violates the no-internet-access requirement; default SageMaker encryption uses AWS-managed keys, not a customer-managed KMS key, which may not satisfy the security team's requirement for explicit encryption control.

423
MCQeasy

A machine learning engineer wants to deploy a pre-trained foundation model for text summarization using SageMaker JumpStart. Which of the following is a primary cost consideration when deploying such a model?

A.The cost of fine-tuning the model on custom data
B.The cost of GPU instances required for low-latency inference
C.The cost of data transfer for inference requests
D.The cost of storing the model artifacts in S3
AnswerB

JumpStart foundation models are large and require accelerated compute, so GPU instance hours dominate deployment cost. Low-latency inference demands continuously running GPU capacity rather than serverless or CPU options, making the instance type and count the primary cost consideration for this deployment.

Why this answer

When deploying a pre-trained foundation model via SageMaker JumpStart, the model is already trained, so fine-tuning cost is optional and not primary. The main ongoing cost is the compute instance used for inference, especially GPU instances needed for low-latency, high-throughput text summarization. Data transfer for inference requests is typically negligible compared to compute, and S3 storage for model artifacts is a minor one-time cost.

Exam trap

MLA-C01 often tests the misconception that data transfer or storage costs dominate ML deployment, when in fact compute instances, especially GPUs, are the primary cost driver for inference.

How to eliminate wrong answers

Option A is wrong because fine-tuning is not required for deploying a pre-trained model; it's an optional step that incurs separate training costs. Option C is wrong because data transfer for inference requests is usually small and often free within the same region, not a primary cost driver. Option D is wrong because storing model artifacts in S3 is inexpensive and a one-time cost, not the main ongoing expense.

424
MCQeasy

A data engineer is building a feature store using Amazon SageMaker Feature Store. The team needs to store features that are updated frequently and require low-latency retrieval for real-time inference. Which type of store should the engineer use?

A.Both online and offline store
B.Offline store
C.Online store
D.Amazon DynamoDB directly
AnswerC

The online store serves features at millisecond latency for real-time inference, satisfying the low-latency retrieval constraint. It is backed by a low-latency online database, whereas the offline store holds historical data in Amazon S3 for training and batch scoring, not real-time serving.

Why this answer

The online store in Amazon SageMaker Feature Store is designed for low-latency, real-time retrieval of feature values for inference. It is backed by a low-latency storage layer and supports single-digit millisecond reads, which is exactly what real-time inference requires. The offline store, by contrast, is optimized for batch training and historical lookups, not real-time serving.

Exam trap

MLA-C01 often tests the confusion between online and offline stores; candidates mistakenly select the offline store because it sounds more comprehensive, but the offline store is for batch, not low-latency real-time inference.

How to eliminate wrong answers

Option A is wrong because although both stores are often used together, the question specifically asks for low-latency retrieval for real-time inference, which is served only by the online store; adding the offline store does not address the latency requirement. Option B is wrong because the offline store is an append-only repository in S3/Glue/Athena designed for batch training and historical feature retrieval, and it does not provide low-latency reads. Option D is wrong because using DynamoDB directly bypasses SageMaker Feature Store's managed feature group, online/offline consistency, and feature discovery capabilities; it is not the intended service for this use case.

425
Multi-Selecthard

A healthcare company deploys a model to predict patient readmission risk. The model was trained on historical data and is now showing signs of concept drift. The team needs to implement a monitoring solution that can detect drift and automatically retrain the model when drift is detected. Which THREE steps should the team take to build this solution? (Choose THREE.)

Select 3 answers
A.Deploy SageMaker Model Monitor to track prediction quality over time
B.Disable the existing endpoint to prevent stale predictions during retraining
C.Set up a process to collect ground truth labels from patient outcomes
D.Manually compare the model's predictions against a holdout validation set each week
E.Use AWS Lambda to invoke a SageMaker training job when drift is detected
AnswersA, C, E

SageMaker Model Monitor continuously captures endpoint data and compares it against a baseline, emitting CloudWatch metrics when drift is detected. This satisfies the requirement to detect drift, providing the trigger signal that downstream automation uses to initiate retraining.

Why this answer

Option A is correct because SageMaker Model Monitor is the AWS service designed to continuously capture endpoint data and evaluate it against baselines, emitting CloudWatch metrics and alarms when data quality, model quality, bias, or drift violations occur, which is exactly what detecting concept drift requires. Option C is correct because Model Monitor's model quality monitoring depends on ground truth labels being joined to captured predictions; for readmission risk, those labels come from actual patient outcomes, so a labeling/feedback pipeline is essential to measure real drift. Option E is correct because the automated retraining requirement is met by triggering a SageMaker training job programmatically — for example, a Lambda function invoked by a CloudWatch alarm on the drift metric — which closes the detect-and-retrain loop.

Option B is not appropriate because disabling the endpoint would halt predictions for clinicians and is unnecessary; retraining can occur while the existing endpoint continues serving, with a new model version deployed afterward. Option D is not appropriate because weekly manual comparison against a holdout set is neither automated nor a production drift-detection mechanism, and it would not satisfy the requirement to automatically retrain when drift is detected.

Exam trap

The trap here is that candidates might think disabling the endpoint (Option B) is necessary to prevent stale predictions, but AWS best practice is to keep the endpoint live and use a separate pipeline (e.g., Lambda triggering a training job) to retrain and then update the endpoint without downtime.

426
MCQmedium

A team wants to use SageMaker Clarify to monitor bias in their production model predictions. They have configured a bias drift monitor. What does SageMaker Clarify compare to detect bias drift?

A.Current input data distribution against the training data distribution
B.Current bias metrics against a baseline bias metrics computed from training data
C.Current SHAP feature attributions against baseline SHAP values
D.Current predictions against ground truth labels collected in real-time
AnswerB

SageMaker Clarify's bias drift monitor compares bias metrics computed on current production data against baseline bias metrics derived from the training dataset. This satisfies the stem's requirement to detect drift by quantifying divergence from the original training distribution, flagging when live predictions deviate from the model's established fairness baseline.

Why this answer

SageMaker Clarify bias drift monitoring compares the current bias metrics (e.g., disparate impact) computed on live data against a baseline bias metric computed from the training data. This detects if bias has drifted over time.

Exam trap

MLA-C01 often tests the difference between data drift, bias drift, and feature attribution drift; candidates may confuse bias drift with data drift.

How to eliminate wrong answers

Option A is wrong because comparing input data distributions is for data drift, not bias drift; bias drift focuses on bias metrics. Option C is wrong because SHAP feature attributions are for explainability, not bias drift; comparing SHAP values is for feature attribution drift. Option D is wrong because comparing predictions against ground truth labels is for model quality drift, not bias drift.

427
MCQeasy

A data scientist is using SageMaker built-in XGBoost algorithm for a regression problem. Which metric is most appropriate as the objective metric for hyperparameter tuning?

A.NDCG
B.RMSE
C.AUC
D.F1
AnswerB

RMSE is the standard objective metric for regression with XGBoost, measuring root mean squared prediction error in the target's units. SageMaker hyperparameter tuning minimises it, directly reflecting the regression task's accuracy, unlike classification metrics such as accuracy, F1, or AUC.

Why this answer

RMSE (Root Mean Squared Error) is the most appropriate objective metric for hyperparameter tuning in a regression problem because it directly measures the average magnitude of prediction errors, with larger errors penalized more heavily. SageMaker's built-in XGBoost algorithm supports RMSE as an evaluation metric for regression, and it is commonly used as the objective metric for tuning jobs.

Exam trap

The trap is confusing classification metrics (AUC, F1) with regression metrics; MLA-C01 often tests whether candidates know that RMSE is for regression while AUC and F1 are for classification.

How to eliminate wrong answers

Option A is wrong because NDCG (Normalized Discounted Cumulative Gain) is a ranking metric used for recommendation systems and search relevance, not for regression. Option C is wrong because AUC (Area Under the ROC Curve) is a classification metric that measures the ability to distinguish between classes, not for continuous target prediction. Option D is wrong because F1 score is a classification metric that balances precision and recall, and is not applicable to regression problems.

428
MCQeasy

An ML engineer needs to split a dataset into training, validation, and test sets. The dataset has a time-based column that should not be leaked. Which split method is most appropriate?

A.Stratified split based on target
B.Temporal split based on date
C.Random split with 70/20/10
D.K-fold cross-validation
AnswerB

A temporal split partitions rows by date, so training uses earlier records and validation/test use later ones. This preserves chronological order and prevents future information leaking into training, directly satisfying the stem's constraint that the time-based column must not be leaked. Random or stratified splits would mix periods and leak future data.

Why this answer

A temporal split ensures that the time-based column is not leaked by preserving the chronological order of the data. This method uses the date column to assign earlier records to the training set and later records to the validation and test sets, preventing future information from influencing the model during training.

Exam trap

AWS often tests the concept of data leakage by presenting random or stratified splits as viable options, trapping candidates who overlook the time-based column and assume standard splitting methods are always safe.

How to eliminate wrong answers

Option A is wrong because a stratified split based on the target variable preserves class proportions but does not account for time order, leading to potential data leakage when time-dependent patterns exist. Option C is wrong because a random split ignores the temporal structure entirely, allowing future data points to appear in the training set and causing leakage. Option D is wrong because K-fold cross-validation shuffles data randomly across folds, which breaks the time sequence and introduces leakage; it is unsuitable for time-series or time-sensitive data.

429
Multi-Selectmedium

A data engineer is designing an ETL pipeline using AWS Glue to transform raw data from S3 into a curated set for ML training. The data contains personally identifiable information (PII) that must be masked before being used by data scientists. Which TWO actions should the engineer take? (Choose TWO.)

Select 2 answers
A.Use AWS Glue DataBrew to define PII masking transformations
B.Use Amazon Kinesis Data Firehose to transform data at ingestion
C.Use AWS Glue Data Catalog to automatically mask PII fields
D.Use AWS Glue ETL scripts with PySpark to apply custom masking functions
E.Use AWS Glue Crawler to detect and mask PII automatically
AnswersA, D

DataBrew provides built-in PII masking transformations, such as substitution and hashing, that run as a Glue ETL step before data reaches data scientists. This satisfies the requirement to mask PII in the curated set, keeping sensitive values out of ML training data.

Why this answer

Option A is correct because AWS Glue DataBrew provides purpose-built, no-code visual transformations including PII masking recipes (e.g., replacing, hashing, or redacting sensitive columns) that can be applied within the Glue/DataBrew pipeline before data reaches data scientists. Option D is correct because AWS Glue ETL jobs run on Apache Spark, so the engineer can write PySpark code using functions like sha2(), regexp_replace(), or custom UDFs to deterministically mask PII fields during the transform stage. Option B is not appropriate because Kinesis Data Firehose is a streaming ingestion/delivery service, not the batch ETL transformation layer described here, and it does not provide PII masking logic.

Option C is incorrect because the Glue Data Catalog is a metadata repository storing table and schema definitions; it does not perform data masking on field values. Option E is incorrect because a Glue Crawler only infers schemas and populates the Data Catalog—it detects formats and structures, not PII, and cannot mask data.

Exam trap

MLA-C01 often tests whether candidates confuse Glue Data Catalog and Crawler with data transformation tools — the trap is selecting Data Catalog or Crawler for masking because they sound like they manage data, but they only store metadata and discover schemas, not modify data.

430
MCQmedium

A data scientist is training a binary classification model using Amazon SageMaker. The dataset has a severe class imbalance (95% negative, 5% positive). The model achieves 99% accuracy but fails to identify positive cases correctly. Which action should the data scientist take to improve the model's ability to detect positive cases?

A.Switch to a logistic regression model with balanced class weights.
B.Use accuracy as the evaluation metric and retrain the model.
C.Apply SMOTE (Synthetic Minority Over-sampling Technique) to the training data.
D.Use the F1 score as the evaluation metric and adjust the classification threshold based on the precision-recall curve.
AnswerD

Accuracy is misleading under 95/5 imbalance, since predicting all negatives scores 95%. The F1 score balances precision and recall on the minority class, and moving the threshold along the precision-recall curve trades false positives for higher recall, directly improving positive-case detection.

Why this answer

In a severely imbalanced dataset (95% negative, 5% positive), accuracy is misleading. The F1 score balances precision and recall, and adjusting the classification threshold based on the precision-recall curve allows the model to prioritize recall for the minority class, directly improving detection of positive cases. This approach is recommended in SageMaker when using built-in algorithms or custom models with imbalanced data.

Exam trap

The trap here is that candidates often think oversampling (SMOTE) or changing the model type is the primary fix, but the exam tests understanding that evaluation metrics and threshold tuning are critical for imbalanced classification, not just data preprocessing.

How to eliminate wrong answers

Option A is wrong because switching to logistic regression with balanced class weights may help, but it is not the best action; the question asks for a single action to improve detection, and adjusting the threshold and metric (D) is more direct and effective than changing the model type. Option B is wrong because using accuracy as the evaluation metric will continue to favor the majority class and fail to reflect poor positive detection, reinforcing the original problem. Option C is wrong because applying SMOTE to the training data can introduce synthetic samples, but it does not address the need to evaluate and tune the model's decision threshold; SMOTE alone may not fix the detection issue if the threshold remains at 0.5.

431
MCQhard

A company deploys a model using SageMaker real-time endpoint with auto scaling. They observe that during a traffic spike, the endpoint quickly scales up to 10 instances, but after the spike, it takes a long time to scale down, leading to high costs. The scaling policy is based on a simple average CPU utilization threshold. Which adjustment would optimize the scaling down behavior?

A.Increase the scale-in cooldown period to prevent premature scale-down.
B.Decrease the scale-in cooldown period to allow the endpoint to scale down faster when utilization drops.
C.Use a step scaling policy with a larger step adjustment for scale-in.
D.Change the scaling policy to use memory utilization instead of CPU.
AnswerB

Scale-in cooldown governs how long the endpoint waits after a scale-in before removing more instances. Shortening it lets the endpoint shed surplus instances sooner once CPU utilisation drops, cutting idle capacity costs after the traffic spike.

Why this answer

Decreasing the scale-in cooldown period allows the endpoint to respond more quickly to sustained drops in CPU utilization. By default, SageMaker auto scaling uses cooldown periods to prevent rapid fluctuations; a long scale-in cooldown delays the termination of instances after utilization falls, keeping costs high. Reducing this cooldown lets the endpoint scale down faster when the spike subsides, directly addressing the problem.

Exam trap

The trap here is that candidates often confuse cooldown periods with step adjustments, thinking that larger scale-in steps will speed up the process, when in fact the cooldown period controls the timing of when scaling actions can occur.

How to eliminate wrong answers

Option A is wrong because increasing the scale-in cooldown period would make the problem worse, not better—it would cause the endpoint to wait even longer before scaling down, increasing costs. Option C is wrong because step scaling policies control the magnitude of scaling adjustments (e.g., adding or removing multiple instances at once), but they do not affect the timing or delay of scale-in actions; the cooldown period is the key parameter for timing. Option D is wrong because changing the metric to memory utilization does not address the core issue of slow scale-down timing; the problem is with the cooldown period, not the metric choice.

432
MCQmedium

An MLOps engineer is setting up a SageMaker endpoint for a model that performs inference on large images. The model is containerized and expects input in a specific format. The team wants to preprocess the images (resize and normalize) before passing them to the model. What is the most efficient way to implement this?

A.Configure SageMaker to use a preprocessing container as the first step of an inference pipeline, followed by the model container.
B.Use Amazon API Gateway to perform request transformation before forwarding to the endpoint.
C.Package the preprocessing logic into the same Docker container as the model.
D.Use a Lambda function as a proxy to preprocess requests before calling the SageMaker endpoint.
AnswerA

An inference pipeline chains containers so preprocessing runs inside the same endpoint invocation, resizing and normalising images before the model container receives them. This satisfies the required input format without adding a separate client-side step or extra network hop, keeping latency and operational overhead minimal.

Why this answer

SageMaker Inference Pipelines allow you to chain multiple containers in a serial fashion, where the output of one container becomes the input of the next. By placing a preprocessing container as the first step, you can resize and normalize large images before passing them to the model container, which keeps the model container focused on inference and avoids unnecessary data transfer or custom code. This is the most efficient and natively supported approach within SageMaker for multi-step inference workflows.

Exam trap

The trap here is that candidates often choose Option C (packaging everything into one container) because it seems simpler, but they overlook the fact that SageMaker Inference Pipelines are specifically designed for this exact use case and provide better modularity, maintainability, and efficiency.

How to eliminate wrong answers

Option B is wrong because Amazon API Gateway is designed for request routing and transformation at the HTTP level, not for heavy image preprocessing (e.g., resizing and normalization) — it lacks the computational capability and libraries needed for such tasks, and it would introduce latency without any benefit. Option C is wrong because packaging preprocessing logic into the same container as the model violates the separation of concerns principle and makes the container larger and harder to maintain; it also prevents independent scaling or updating of preprocessing steps. Option D is wrong because using a Lambda function as a proxy adds unnecessary cold-start latency and a 6 MB (or 10 MB via extension) payload limit, which is problematic for large images, and it does not integrate as seamlessly with SageMaker's built-in batching or inference pipeline features.

433
Multi-Selecthard

A data engineer is using AWS Glue to run an ETL job that joins two large datasets and writes the output to S3 for ML training. The job is failing due to out-of-memory errors. Which THREE actions can help resolve this issue? (Select THREE.)

Select 3 answers
A.Filter unnecessary records early in the transformation
B.Increase the number of DPUs for the Glue job
C.Partition the input data on the join keys
D.Switch from Spark to Python shell
E.Use a smaller worker type
AnswersA, B, C

Filtering records before the join shrinks the datasets shuffled across executors, directly lowering the memory each task must hold. This addresses the out-of-memory constraint by reducing data volume at the earliest transformation stage, before the join and write to S3.

Why this answer

Option A is correct because filtering unnecessary records early in the transformation reduces the volume of data shuffled and held in memory during the join, directly lowering memory pressure on the Spark executors. Option B is correct because increasing the number of DPUs for the Glue job adds more workers and memory capacity, allowing Spark to distribute the join workload across more executors and avoid out-of-memory failures. Option C is correct because partitioning the input data on the join keys co-locates matching keys within the same partition, minimizing the shuffle and memory footprint required to perform the join.

Option D is not appropriate because a Python shell job runs single-node, non-distributed Python and cannot handle large-scale joins across two large datasets. Option E is not appropriate because using a smaller worker type reduces the memory and compute resources available per worker, which would worsen the out-of-memory condition.

Exam trap

The trap here is that candidates might think reducing worker size (Option E) saves costs and helps memory, but it actually reduces available memory per worker, making out-of-memory errors more likely.

434
MCQeasy

A healthcare analytics team trains models in Amazon SageMaker and needs an immutable, queryable record of which dataset version and training job produced each registered model version, so an auditor can trace a deployed model back to its inputs months later. Which SageMaker capability should they rely on?

A.SageMaker ML Lineage Tracking, which automatically records entities such as datasets, training jobs, and model package versions and their relationships.
B.SageMaker Model Monitor, which schedules jobs that compare production traffic against a baseline and emit violations to CloudWatch.
C.SageMaker Experiments, which groups training runs into experiments and trials so you can compare their metrics side by side.
D.SageMaker Debugger, which captures tensors and system metrics during training and can halt a job when a rule is triggered.
AnswerA

ML Lineage Tracking creates lineage entities and associations for artifacts, trials, and actions as the workflow runs, so an auditor can traverse from a model package version back to the training job and the input dataset. It is queryable through the SageMaker API and integrates with the model registry, matching the traceability requirement without custom bookkeeping.

Why this answer

ML Lineage Tracking is the SageMaker feature that records artifacts, trials, actions, and their associations automatically as training and registration proceed, producing a queryable graph from a model package version back to the training job and dataset. Other SageMaker capabilities address monitoring, training-run observability, or run comparison, none of which yields the end-to-end provenance record the audit requires.

Exam trap

The trap here is confusing experiment tracking, which compares runs, with lineage tracking, which records the provenance relationships an auditor needs.

435
Multi-Selectmedium

A machine learning engineer is preparing a dataset for training a SageMaker model. The dataset contains missing values in several numerical features. The engineer wants to handle these missing values during the training pipeline. Which two methods are valid ways to handle missing values in SageMaker? (Choose two.)

Select 2 answers
A.Use SageMaker Data Wrangler to impute missing values with the mean or median.
B.Use SageMaker Clarify to replace missing values with the mode of each column.
C.Use SageMaker Debugger to automatically fill missing values during training.
D.Enable the SageMaker built-in XGBoost algorithm's default handling of missing values.
E.Configure SageMaker Automatic Model Tuning to impute missing values via hyperparameter search.
AnswersA, D

SageMaker Data Wrangler provides built-in transformations for handling missing values, including imputation with mean, median, or mode. It allows you to visually inspect and apply these transformations as part of a data preparation flow, which can then be exported to a pipeline. This is a valid and recommended approach for handling missing values before training.

Why this answer

SageMaker Data Wrangler offers built-in transformations to impute missing values, and the built-in XGBoost algorithm natively handles missing values by learning default directions. Both are valid methods for dealing with missing data in a SageMaker training pipeline. The other services are not designed for data preprocessing.

Exam trap

The trap here is confusing monitoring, tuning, or bias detection services with data preprocessing capabilities.

436
MCQmedium

A company is using SageMaker to train a model for image classification. The training dataset contains 100,000 labeled images. The team wants to use a pre-trained model to reduce training time. Which SageMaker feature should they use?

A.SageMaker Debugger
B.SageMaker Model Monitor
C.SageMaker built-in Image Classification algorithm
D.SageMaker JumpStart
AnswerD

SageMaker JumpStart provides pre-trained foundation and task-specific models, including image classification, that can be fine-tuned on your own dataset. This directly satisfies the requirement to start from a pre-trained model and cut training time, rather than building a model from scratch with custom training code.

Why this answer

SageMaker JumpStart provides pre-trained, publicly available foundation models and task-specific models (including image classification) that can be fine-tuned on your dataset, dramatically reducing training time and data requirements. For a team wanting to leverage a pre-trained model for image classification on 100,000 labeled images, JumpStart is the purpose-built feature that offers one-click deployment and fine-tuning of pre-trained models.

Exam trap

The trap is confusing SageMaker's built-in algorithms with pre-trained model hubs — candidates may pick the built-in Image Classification algorithm because it sounds like the 'official' image classification tool, but the question specifically asks for leveraging a pre-trained model, which is JumpStart's core value proposition.

How to eliminate wrong answers

Option A is wrong because SageMaker Debugger is a tool for monitoring and debugging training jobs — it detects issues like vanishing gradients, overfitting, and resource bottlenecks, but it does not provide pre-trained models. Option B is wrong because SageMaker Model Monitor detects data drift and model quality degradation in deployed models — it is a post-deployment monitoring tool, not a training accelerator. Option C is wrong because the SageMaker built-in Image Classification algorithm trains a model from scratch (or with transfer learning mode, but it is not a curated library of pre-trained models like JumpStart) — it does not provide the same breadth of pre-trained model options or the ease of use that JumpStart offers for leveraging pre-trained models.

437
MCQeasy

A company uses SageMaker Neo to compile a trained model for deployment on edge devices. What is the primary benefit of using Neo?

A.It monitors model drift in production
B.It reduces model size and improves inference speed on target hardware
C.It automatically retrains the model on new data
D.It provides a serverless inference endpoint
AnswerB

SageMaker Neo compiles models into optimised executables for specific target hardware, reducing model size and improving inference latency and throughput on edge devices. This satisfies the edge deployment constraint, where resource limits make unoptimised frameworks impractical.

Why this answer

SageMaker Neo compiles trained models into optimized executables for specific target hardware (CPU, GPU, or edge accelerators), producing smaller artifacts and faster inference by leveraging hardware-specific instruction sets. It is designed for edge and constrained deployments where latency, memory, and compute are limited. Neo does not handle monitoring, retraining, or endpoint provisioning.

Exam trap

MLA-C01 often tests confusion between SageMaker features — candidates pick Model Monitor or Serverless Inference when the question is specifically about compiling models for edge hardware with Neo.

How to eliminate wrong answers

Option A is wrong because model drift monitoring is handled by SageMaker Model Monitor, not Neo. Option C is wrong because automatic retraining is a pipeline concern (SageMaker Pipelines, Clarify, or custom MLOps), not a compilation feature. Option D is wrong because serverless inference endpoints are provided by SageMaker Serverless Inference, which is unrelated to Neo's compilation role.

438
Multi-Selectmedium

A company wants to track the lineage of their ML models for reproducibility and auditability. Which THREE services or features should they use together to achieve this? (Choose THREE.)

Select 3 answers
A.Amazon S3 versioning
B.SageMaker Experiments
C.AWS CloudTrail
D.SageMaker ML Lineage Tracking
E.AWS Config
AnswersA, B, D

Amazon S3 versioning preserves every object revision, so each training dataset and model artefact retains an immutable, retrievable history. This satisfies the lineage and auditability constraint by preventing overwrites, letting auditors trace exactly which data version produced a given model. Combined with SageMaker ML Lineage Tracking and Model Registry, it completes the reproducibility requirement.

Why this answer

Amazon S3 versioning (A) is correct because it preserves every version of the datasets and model artifacts stored in S3, so a given training run can be tied to the exact immutable object version used, which is essential for reproducibility and auditability. SageMaker Experiments (B) is correct because it records experiment runs, trial components, parameters, metrics, and input/output artifacts, giving the structured record of each training attempt needed to reproduce results. SageMaker ML Lineage Tracking (D) is correct because it automatically creates and stores entities and relationships (trials, trial components, artifacts, contexts, actions) forming a queryable lineage graph from data through training to the deployed model.

AWS CloudTrail (C) only logs API activity and control-plane events for auditing who did what, not the data/model lineage relationships, so it does not by itself provide reproducibility lineage. AWS Config (E) evaluates and records resource configuration compliance over time, which is unrelated to tracking ML artifact provenance and experiment history.

Exam trap

The trap here is that candidates confuse AWS CloudTrail or AWS Config with lineage tracking because both deal with 'tracking' and 'auditing,' but they operate at the infrastructure/API level, not at the ML experiment and artifact relationship level required for model lineage.

439
MCQeasy

A data science team needs to deploy a trained PyTorch model for real-time inference with sub-100ms latency. The model fits on a single GPU. Which SageMaker inference option is MOST cost-effective while meeting the latency requirement?

A.SageMaker Batch Transform
B.SageMaker real-time endpoint on ml.g4dn.xlarge
C.SageMaker Async Inference
D.SageMaker Serverless Inference
AnswerB

A SageMaker real-time endpoint on ml.g4dn.xlarge provides a persistent, GPU-backed inference host with low single-digit millisecond overhead, meeting sub-100ms latency. Since the model fits one GPU, this single-instance option is more cost-effective than multi-GPU or serverless alternatives.

Why this answer

SageMaker real-time endpoints provide dedicated, persistent instances that can handle synchronous inference with sub-100ms latency. The ml.g4dn.xlarge instance includes a single NVIDIA T4 GPU, which is sufficient for the model size and offers the lowest cost among GPU instances that meet the latency requirement. This option balances performance and cost for real-time, low-latency inference.

Exam trap

The trap here is that candidates often choose SageMaker Serverless Inference for its cost-saving potential, but they overlook the cold start latency and lack of GPU support, which makes it unsuitable for real-time, sub-100ms inference with PyTorch models.

How to eliminate wrong answers

Option A is wrong because SageMaker Batch Transform is designed for asynchronous, offline inference on large datasets, not for real-time sub-100ms latency; it processes data in batches and returns results only after the job completes. Option C is wrong because SageMaker Async Inference queues inference requests and processes them asynchronously, which introduces unpredictable latency and is not suitable for sub-100ms real-time requirements. Option D is wrong because SageMaker Serverless Inference auto-scales from zero and has a cold start latency that can exceed 100ms, especially for GPU-based models, making it unsuitable for strict real-time latency demands.

440
MCQeasy

A company wants to reduce costs for a SageMaker real-time endpoint that has variable traffic. Which feature allows the endpoint to automatically adjust instance count based on demand?

A.SageMaker Savings Plans
B.SageMaker Inference Recommender
C.SageMaker Model Monitor
D.Auto Scaling for SageMaker endpoints
AnswerD

Application Auto Scaling for SageMaker endpoints adjusts the instance count of a production variant in response to CloudWatch metrics such as InvocationsPerInstance, matching capacity to variable demand. This satisfies the requirement to scale automatically while preserving performance during peaks.

Why this answer

Auto Scaling for SageMaker endpoints is the native capability that dynamically adjusts the number of instances behind a real-time endpoint based on CloudWatch metrics such as InvocationsPerInstance or ModelLatency. It uses Application Auto Scaling policies (target tracking or step scaling) to add instances during traffic spikes and remove them during lulls, directly reducing cost for variable workloads. Savings Plans and Inference Recommender do not perform runtime scaling.

Exam trap

MLA-C01 often tests the confusion between cost-optimization features — candidates pick Savings Plans (a billing discount) when the question is actually about dynamic capacity adjustment via autoscaling.

How to eliminate wrong answers

Option A is wrong because SageMaker Savings Plans are a pricing/billing commitment model (1- or 3-year spend commitment) that discounts usage but does not change instance count in response to demand. Option B is wrong because Inference Recommender is a one-time recommendation tool that benchmarks instance types and configurations to suggest the best deployment option — it does not perform ongoing autoscaling. Option C is wrong because Model Monitor detects data drift, bias, and quality issues in production traffic; it has no role in scaling capacity.

441
MCQmedium

A financial services company has a SageMaker real-time endpoint serving a fraud detection model. Compliance requires that all inference requests and responses be logged with the ability to detect anomalous input feature distributions over time. The team wants a managed solution that captures request/response payloads to Amazon S3 and automatically computes statistics and constraints against a baseline. Which combination of SageMaker features should they enable?

A.Configure the endpoint to write inference logs to Amazon CloudWatch Logs and create a custom Lambda function to parse and analyze the logs.
B.Enable SageMaker Model Monitor data capture on the endpoint and schedule a monitoring job using the baseline constraints and statistics.
C.Enable SageMaker Debugger on the endpoint and configure rules to monitor for data drift.
D.Enable AWS CloudTrail data events on the S3 bucket used by the endpoint and configure Amazon CloudWatch Logs metric filters.
AnswerB

SageMaker Model Monitor data capture records request and response payloads to S3, and the monitoring schedule evaluates them against a baseline to detect drift and anomalies. This directly satisfies the compliance need to log inference traffic and detect anomalous feature distributions without custom code.

Why this answer

SageMaker Model Monitor is the managed service for monitoring deployed models. Data capture stores inference request and response data in S3, and monitoring schedules compare that data to a baseline to detect data drift, model quality issues, bias, and feature attribution drift. This provides both the audit trail and the automated anomaly detection required.

Exam trap

The trap here is assuming that CloudWatch Logs or CloudTrail alone can provide managed drift detection, when they only capture logs or API activity and require custom analysis.

442
Multi-Selecthard

An ML engineer is fine-tuning a foundation model using RLHF on SageMaker. Which THREE components are essential for this workflow? (Select THREE.)

Select 3 answers
A.A reward model trained on the preference data
B.A large validation dataset for final evaluation
C.The PPO (Proximal Policy Optimization) algorithm for model updates
D.A preference dataset with human rankings
E.A PEFT technique like LoRA
AnswersA, C, D

RLHF requires a reward model that scores model outputs against learned human preferences, supplying the scalar signal PPO maximises. Without it, no preference-based objective exists, so the fine-tuning loop cannot optimise the foundation model toward preferred responses.

Why this answer

Option A is correct because RLHF requires a reward model that has been trained on human preference data to serve as the scalar reward signal guiding policy optimization. Option C is correct because PPO (Proximal Policy Optimization) is the standard reinforcement learning algorithm used to update the policy (the language model) against the reward model while constraining updates via a KL penalty to the reference model. Option D is correct because a preference dataset containing human rankings (e.g., chosen vs. rejected responses) is the foundational input used to train the reward model in the first place.

Option B is not essential to the RLHF training workflow itself; a validation set is useful for evaluation but is not a required RLHF component. Option E is not essential because PEFT methods like LoRA are an optional efficiency technique for fine-tuning, not a required element of the RLHF pipeline.

Exam trap

MLA-C01 often tests the distinction between RLHF's essential components (preference data, reward model, PPO) and optional optimizations like PEFT or evaluation datasets, causing candidates to select non-essential items.

443
MCQmedium

A data scientist is training a deep learning model on Amazon SageMaker and notices that the training loss decreases but the validation loss starts increasing after a certain number of epochs. The model is likely overfitting. Which SageMaker feature can they use to detect and diagnose this issue during training?

A.SageMaker Model Monitor
B.SageMaker Automatic Model Tuning
C.SageMaker Experiments
D.SageMaker Debugger
AnswerD

SageMaker Debugger captures tensor-level metrics such as training and validation loss throughout the job and applies built-in rules that flag divergence between them. This directly detects the overfitting pattern described, where validation loss rises while training loss keeps falling.

Why this answer

SageMaker Debugger is the correct choice because it provides real-time monitoring of training metrics, including loss values, and can automatically detect anomalies such as overfitting (where training loss decreases but validation loss increases). It allows you to set rules (e.g., `OverfitRule`) that trigger alerts or stop training when overfitting is detected, enabling proactive diagnosis during the training job.

Exam trap

The trap here is that candidates may confuse SageMaker Debugger's real-time training diagnostics with SageMaker Model Monitor's post-deployment monitoring, or assume that hyperparameter tuning (Automatic Model Tuning) inherently addresses overfitting, when in fact it only searches for optimal hyperparameters without detecting the overfitting condition during a specific training run.

How to eliminate wrong answers

Option A is wrong because SageMaker Model Monitor is designed to monitor inference endpoints for data drift and model quality after deployment, not for detecting overfitting during training. Option B is wrong because SageMaker Automatic Model Tuning (hyperparameter tuning) optimizes hyperparameters to improve model performance but does not monitor or diagnose overfitting in real time during a single training run. Option C is wrong because SageMaker Experiments tracks and organizes training runs, metrics, and parameters for comparison, but it does not actively detect or alert on overfitting patterns during training.

444
MCQmedium

A company is building a fraud detection model on an imbalanced dataset (99% legitimate, 1% fraudulent). To improve recall on the minority class, they want to resample data. Which combination of techniques should they use?

A.SMOTE on entire dataset before train/test split
B.Random oversampling of minority class before train/test split
C.Random undersampling of majority class
D.SMOTE on training set only
AnswerD

SMOTE synthesises new minority-class examples by interpolating between existing fraudulent cases, directly raising recall on the 1% class. Applying it only to the training set preserves the genuine 99:1 distribution in validation and test data, preventing the inflated performance estimates that leakage from resampled holdout data would cause.

Why this answer

SMOTE (Synthetic Minority Over-sampling Technique) generates synthetic minority-class samples by interpolating between existing minority instances. Applying SMOTE only to the training set prevents synthetic samples from leaking into the test set, which would inflate evaluation metrics and produce an overly optimistic model. This is the correct resampling approach for improving recall on an imbalanced fraud dataset.

Exam trap

MLA-C01 often tests whether candidates apply resampling before the train/test split — the trap is forgetting that SMOTE or oversampling on the full dataset leaks synthetic information into the test set and invalidates evaluation.

How to eliminate wrong answers

Option A is wrong because applying SMOTE before the train/test split causes data leakage — synthetic samples derived from test-set instances contaminate training, producing misleadingly high evaluation metrics. Option B is wrong because random oversampling before the split duplicates minority instances across train and test, again causing leakage and overfitting. Option C is wrong because random undersampling of the majority class discards potentially useful legitimate-transaction data and, used alone, does not address the recall goal as effectively as SMOTE on the training set.

445
MCQmedium

A company wants to deploy a scikit-learn model to a SageMaker AI real-time endpoint. The model must be loaded from a custom Python module that contains preprocessing logic not present in the built-in scikit-learn container. The team wants to minimize operational overhead and does not need to change system-level libraries. Which approach should the engineer take?

A.Package the preprocessing logic in an inference.py file and pass it as the entry_point to a SageMaker AI framework estimator using the scikit-learn framework.
B.Use the SageMaker AI built-in scikit-learn container without an entry point and rely on the default inference handler.
C.Build a fully custom Docker image with a Bring Your Own Container (BYOC) approach and push it to Amazon ECR.
D.Deploy the model with SageMaker AI batch transform and invoke it from the application on demand.
AnswerA

SageMaker AI framework estimators support an entry_point script that defines model_fn and input_fn or transform_fn. This lets the engineer add custom Python preprocessing while reusing the managed scikit-learn container, which minimizes operational overhead because no Docker image must be built or maintained.

Why this answer

A SageMaker AI framework estimator with an entry_point script lets the team inject custom Python preprocessing while reusing the managed scikit-learn container. This satisfies the functional requirement with far less operational effort than building and maintaining a custom Docker image.

Exam trap

The trap here is defaulting to a fully custom container whenever any custom code is needed, even when an entry point script on a managed framework container is sufficient.

446
MCQmedium

An ML engineer has a real-time SageMaker endpoint serving a fraud-detection model. The team wants to release a new model version to a small percentage of live traffic first, monitor CloudWatch metrics for accuracy regressions, and roll back quickly if performance degrades. They also want the production and candidate variants to share the same endpoint so latency comparisons are apples-to-apples. Which SageMaker deployment strategy should they use?

A.Update the existing endpoint in place by replacing the production variant's model with the new model artifact.
B.Configure the endpoint with two production variants (current and candidate) and set initial variant weights, using CloudWatch alarms to trigger rollback.
C.Create a second endpoint with the new model and use Route 53 weighted routing to split traffic.
D.Deploy the new model as a production variant on the existing endpoint and configure a shadow variant with zero traffic weight.
AnswerB

SageMaker supports multiple production variants on a single endpoint with per-variant traffic weights, so the current model keeps most traffic while the candidate receives a small percentage. Both variants share the endpoint's instances, making latency comparison fair. CloudWatch alarms on model or invocation metrics can invoke automatic rollback via Deployment Guardrails, satisfying the gradual rollout and quick-recovery requirements.

Why this answer

A single SageMaker endpoint can host multiple production variants, each with its own model artifact and a traffic weight that you adjust without redeploying infrastructure. This provides canary-style exposure, apples-to-apples latency on shared instances, and a fast rollback path when CloudWatch alarms fire. It is the native SageMaker mechanism for controlled model releases with live monitoring, matching every stated requirement.

Exam trap

The trap here is assuming that any two-model setup provides safe canary rollout, when separate endpoints or shadow variants cannot send a controlled fraction of live responses to the candidate.

447
MCQeasy

A team wants to fine-tune a pre-trained Hugging Face transformer model for text classification using SageMaker. They have a custom training script. Which SageMaker estimator should they use?

A.SageMaker generic estimator with a custom container
B.SageMaker Hugging Face estimator
C.SageMaker PyTorch estimator
D.SageMaker TensorFlow estimator
AnswerB

The Hugging Face estimator ships pre-built containers with transformers, PyTorch and TensorFlow libraries, so a custom training script written against the Hugging Face API runs without dependency work. It satisfies the requirement to fine-tune a pre-trained transformer for text classification while still accepting the team's own script.

Why this answer

The SageMaker Hugging Face estimator is purpose-built for Hugging Face transformer models: it uses the official Hugging Face Deep Learning Containers, supports the transformers, datasets, and accelerate libraries, and accepts hyperparameters like epoch, learning_rate, and model_name_or_path. It is the natural choice for fine-tuning a pre-trained Hugging Face model with a custom training script.

Exam trap

MLA-C01 often tests whether candidates know that Hugging Face models have a dedicated SageMaker estimator — the trap is picking the PyTorch estimator because transformers are PyTorch-based, but the PyTorch DLC does not ship with Hugging Face libraries by default.

How to eliminate wrong answers

Option A is wrong because a generic estimator with a custom container requires you to build and maintain your own Docker image with all Hugging Face dependencies — unnecessary overhead when an official estimator exists. Option C is wrong because the PyTorch estimator uses the PyTorch DLC, which does not include Hugging Face libraries by default; you would have to install them via a requirements.txt, adding complexity. Option D is wrong because the TensorFlow estimator is for TensorFlow models and does not natively support Hugging Face transformers, which are PyTorch-first.

448
MCQmedium

A financial services company uses SageMaker to train a fraud detection model. They have imbalanced data with 1% fraud. They trained a Gradient Boosting model using SMOTE for oversampling and achieved 99% accuracy on the test set, but the fraud recall is only 10%. The data scientist is concerned about the model's performance. Which change is most likely to improve fraud recall without sacrificing too much precision?

A.Use a different evaluation metric like F1-score during training.
B.Increase the weight of the fraud class in the loss function.
C.Reduce the SMOTE sampling ratio to create more synthetic samples.
D.Use a random undersampling of the majority class.
AnswerB

Class weighting directly penalises misclassified fraud cases during gradient boosting, pushing the model to raise recall on the 1% minority class. Unlike SMOTE, which already oversampled, weighting alters the loss landscape itself, improving fraud detection while retaining most precision.

Why this answer

The model achieves 99% accuracy but only 10% fraud recall because the class imbalance causes the loss function to be dominated by the majority (non-fraud) class. Increasing the weight of the fraud class in the loss function (e.g., via scale_pos_weight in XGBoost or class_weight in scikit-learn) directly penalizes misclassification of fraud cases more heavily, forcing the model to prioritize recall on the minority class. This is a targeted intervention at the learning objective itself, unlike changing the evaluation metric, which only changes how you measure performance without altering what the model learns.

Exam trap

MLA-C01 often tests the misconception that changing the evaluation metric (like switching to F1-score) will improve model performance — it only changes measurement, not the model's learned behavior.

How to eliminate wrong answers

Option A is wrong because switching to F1-score as an evaluation metric only changes how you score the model — it does not change the loss function or gradient updates, so the model will still be biased toward the majority class. Option C is wrong because reducing the SMOTE sampling ratio creates fewer synthetic fraud samples, which would worsen the imbalance and likely decrease recall further. Option D is wrong because random undersampling of the majority class discards potentially useful majority-class information and, while it can help balance, it is a cruder approach than class weighting and can hurt precision more severely by removing informative negatives.

449
MCQmedium

A machine learning team has a model that needs to serve predictions with very low latency (under 10 ms) for a real-time web application. The model is a small ensemble of three neural networks that fits in memory. Which SageMaker inference option is MOST appropriate?

A.SageMaker batch transform
B.SageMaker real-time endpoint
C.SageMaker asynchronous inference
D.SageMaker serverless inference
AnswerB

Real-time endpoints keep the model loaded on persistent instances and return predictions synchronously, avoiding the cold-start and queueing overhead of serverless inference. For a small in-memory ensemble needing sub-10 ms responses, this persistent hosting meets the latency requirement.

Why this answer

SageMaker real-time endpoints are designed for low-latency, synchronous inference, making them the best fit for a model that must serve predictions in under 10 ms. Since the ensemble of three neural networks fits in memory, a real-time endpoint can keep the model loaded and respond to each request with minimal overhead, typically using HTTPS and the SageMaker InvokeEndpoint API.

Exam trap

The trap here is that candidates confuse 'low latency' with 'serverless' or 'asynchronous' options, not realizing that serverless inference has cold starts and asynchronous inference adds queueing delays, both of which break the sub-10 ms requirement.

How to eliminate wrong answers

Option A is wrong because SageMaker batch transform is an asynchronous, offline inference option that processes large datasets in batches and does not provide real-time, low-latency responses. Option C is wrong because SageMaker asynchronous inference is designed for requests with large payloads or long processing times, and it introduces queueing and callback mechanisms that add latency beyond the 10 ms requirement. Option D is wrong because SageMaker serverless inference auto-scales from zero and has a cold-start latency that can exceed 10 ms, making it unsuitable for sub-10 ms real-time predictions.

450
MCQeasy

A company deployed a machine learning model on an Amazon SageMaker real-time endpoint. Over several weeks, they notice that inference latency has been gradually increasing, especially during peak business hours. The model and instance type have remained unchanged. What is the most likely cause of the increased latency?

A.The inference script is not using batch processing.
B.The SageMaker endpoint auto scaling is not configured to scale out quickly enough under increasing traffic.
C.The model size is too large for the instance type.
D.The endpoint has data capture enabled, causing additional overhead.
AnswerB

With the model and instance unchanged, rising peak-hour latency points to insufficient capacity: auto scaling policies or cooldowns react too slowly to traffic growth, so requests queue behind a saturated endpoint instead of additional instances absorbing load.

Why this answer

The gradual increase in latency during peak hours, with no change to the model or instance type, strongly indicates that the endpoint is not scaling out fast enough to handle increased traffic. SageMaker real-time endpoints rely on auto scaling policies to add instances based on metrics like invocation count or CPU utilization; if the scale-out step is too slow or the cooldown period is too long, requests queue up and latency rises. This matches the symptom of latency growing over weeks as traffic patterns evolve, rather than a sudden spike.

Exam trap

The trap here is that candidates may confuse a gradual latency increase with a model size or code issue, but the key clue is the unchanged model and instance type, pointing to a scaling configuration problem rather than a static resource limitation.

How to eliminate wrong answers

Option A is wrong because batch processing is not relevant to a real-time endpoint; SageMaker real-time endpoints process individual requests synchronously, and the inference script's use of batching would not cause gradual latency increases over weeks. Option C is wrong because the model size has remained unchanged, so if it were too large for the instance type, latency would be consistently high from the start, not gradually increasing. Option D is wrong because data capture, when enabled, adds a small, fixed overhead per request (writing to S3), which would cause a constant latency increase, not a gradual one that worsens over weeks.

Page 5

Page 6 of 9

Page 7

All pages