Courseiva

AWS Certified Machine Learning Engineer Associate MLA-C01 (MLA-C01) — Questions 451–525

665 questions total · 9pages · All types, answers revealed

Page 6

Page 7 of 9

Page 8
451
MCQeasy

A data science team deploys a regression model to Amazon SageMaker for real-time inference. After one month, the model's prediction errors increase significantly, but data distributions remain unchanged. Which monitoring approach is MOST suitable for detecting this issue?

A.Set up Amazon SageMaker Model Monitor to track model performance metrics against ground truth labels as they arrive.
B.Use Amazon SageMaker Clarify to monitor feature attribution drift.
C.Enable Amazon CloudWatch to monitor model endpoint latency.
D.Configure Amazon SageMaker Model Monitor to track data drift on the input features.
AnswerA

Because distributions are unchanged, the issue is performance degradation rather than data drift. Model Monitor comparing predictions against ground truth labels detects declining accuracy, satisfying the need to catch error increases when input distributions remain stable.

Why this answer

Amazon SageMaker Model Monitor can be configured to track model performance metrics (e.g., regression error metrics like RMSE or MAE) against ground truth labels as they arrive. Since the question states that data distributions remain unchanged but prediction errors increase, the issue is likely model degradation (e.g., concept drift or model staleness) rather than data drift. Monitoring ground truth labels directly captures this performance degradation, making option A the most suitable approach.

Exam trap

The trap here is that candidates often confuse data drift (changes in input features) with concept drift (changes in the relationship between features and target), and mistakenly choose data drift monitoring (option D) even though the question explicitly states data distributions are unchanged, while the correct approach is to monitor ground truth performance metrics (option A).

How to eliminate wrong answers

Option B is wrong because Amazon SageMaker Clarify is designed for detecting bias and explaining model predictions, not for monitoring model performance degradation over time; it focuses on feature attribution drift, which is a form of explainability monitoring, not a direct measure of prediction error increase. Option C is wrong because Amazon CloudWatch monitoring of endpoint latency tracks infrastructure performance (e.g., response times, invocation counts), not the accuracy or error rate of model predictions; latency issues do not explain increased prediction errors when data distributions are unchanged. Option D is wrong because Amazon SageMaker Model Monitor configured for data drift tracks changes in the input feature distribution, but the question explicitly states that data distributions remain unchanged, so data drift monitoring would not detect the issue; the problem is model performance degradation despite stable input data.

452
MCQmedium

A team needs to deploy a new model version to production while minimizing risk. They want to route 5% of live traffic to the new model and 95% to the current model, and then gradually increase the new model's traffic. Which SageMaker deployment pattern should they use?

A.Shadow testing
B.Blue/green deployment
C.A/B testing with production variants
D.Canary deployment using production variants
AnswerD

Canary deployment using production variants lets you split live traffic across multiple model variants behind one endpoint, initially weighting 5% to the new version and 95% to the current one. You then shift the traffic distribution gradually, satisfying the low-risk, incremental rollout constraint.

Why this answer

Canary deployment using production variants in Amazon SageMaker allows you to route a small percentage of live traffic (e.g., 5%) to the new model version while the rest goes to the existing model. You can then gradually increase the traffic to the new model as confidence grows, minimizing risk. This pattern is specifically designed for progressive rollouts with the ability to roll back if issues arise.

Exam trap

The trap is confusing canary deployment with A/B testing; both use production variants, but canary is for gradual traffic shifting to minimize risk, while A/B testing is for comparing model performance with a fixed split.

How to eliminate wrong answers

Option A is wrong because shadow testing mirrors traffic to the new model without affecting production responses, so it does not route live traffic to the new model. Option B is wrong because blue/green deployment shifts all traffic at once after testing, not gradually. Option C is wrong because A/B testing with production variants is used to compare model performance by splitting traffic, but it typically involves a fixed split for experimentation, not a gradual increase for safe deployment.

453
Multi-Selecteasy

A machine learning engineer is setting up an Amazon SageMaker notebook instance. The instance needs to access a private S3 bucket that contains training data. The notebook instance is in a VPC. Which combination of steps will grant access to the S3 bucket? (Choose TWO.)

Select 2 answers
A.Create a VPC endpoint for S3 in the same VPC and subnet.
B.Assign a public IP address to the notebook instance.
C.Set up a NAT gateway in the public subnet.
D.Create an IAM role with S3 access permissions and attach it to the notebook instance.
E.Attach an internet gateway to the VPC.
AnswersA, D

A gateway VPC endpoint for S3 lets the notebook instance reach the bucket over the private AWS network, satisfying the stem's requirement that traffic stay inside the VPC. It also removes the need for a NAT gateway or internet gateway, since the endpoint routes S3 traffic directly.

Why this answer

Option A is correct because a VPC gateway endpoint for S3 lets instances in the VPC reach S3 privately over the AWS network without needing internet access, NAT, or public IPs, which is exactly what a notebook instance in a VPC requires to reach a private S3 bucket. Option D is correct because S3 access is authorized through IAM: the notebook instance must assume an IAM role whose policy grants the required s3 actions (e.g., s3:GetObject, s3:ListBucket) on the bucket and its objects, and that role must be attached to the instance. Option B is not needed and does not by itself grant S3 permissions, since a public IP only provides internet routing and S3 access still requires IAM authorization.

Option C is unnecessary because a NAT gateway only enables outbound internet traffic and would not be required when using an S3 VPC endpoint. Option E is also unnecessary, as an internet gateway only provides VPC internet connectivity and does not authorize or privately route S3 access.

454
MCQhard

An ML platform team must orchestrate a workflow that trains a model, evaluates it against a baseline, and only registers the model if evaluation passes. If evaluation fails, the workflow must notify the data science channel and stop without registering. The team wants the orchestration logic to be expressed as code, versioned in git, and integrated with SageMaker training jobs and Model Registry. Which approach best fits these requirements?

A.Author a SageMaker Pipeline with a Condition step that compares evaluation metrics to the baseline and branches to a RegisterModel step or a Fail step.
B.Write a Python script on an EC2 instance that calls the SageMaker APIs sequentially and uses if/else statements to decide whether to register.
C.Use AWS Glue workflows to chain the training and evaluation jobs and branch on the evaluation result.
D.Define an AWS Step Functions state machine with a Choice state that inspects evaluation output and calls SageMaker APIs directly.
AnswerA

SageMaker Pipelines are defined as code, can be versioned in git, and natively integrate training jobs, processing jobs, and the Model Registry. A Condition step compares the evaluation metric to the baseline and routes execution to either RegisterModel or a terminal Fail step, enforcing the gate automatically. This expresses the entire orchestration logic declaratively and reproducibly, matching all stated requirements.

Why this answer

SageMaker Pipelines express orchestration as versionable code and include a Condition step for metric-based branching, plus native RegisterModel and Fail steps. This enforces the evaluation gate before registration without custom glue, and the pipeline definition can live in git and be executed reproducibly. It best satisfies the code-as-orchestration, integration, and versioning requirements.

Exam trap

The trap here is assuming any general-purpose orchestrator with an if/else can substitute for native SageMaker pipeline gating and Model Registry integration.

455
MCQmedium

A company uses Amazon SageMaker Ground Truth to label images for object detection. They want to minimize labeling costs while maintaining high accuracy. Which feature should they enable?

A.Active learning to automatically select samples for labeling
B.Use of mechanical turk for all labeling
C.Pre-built annotation workflows for bounding boxes
D.Automated data labeling with AWS Lambda
AnswerA

Active learning uses model uncertainty sampling to prioritise only the most informative unlabelled images for human annotation, cutting the total number of labels purchased. This directly satisfies the cost-minimisation constraint while preserving model accuracy, since ambiguous samples still receive human review.

Why this answer

SageMaker Ground Truth's active learning automatically selects the most informative unlabeled samples for human labeling, reducing the total number of labels needed while maintaining model accuracy. This directly lowers labeling costs because fewer human annotations are required for the same or better model performance.

Exam trap

MLA-C01 often tests whether candidates confuse 'automated labeling' (model pre-labels, humans verify) with 'active learning' (model selects which samples humans label) — the trap is picking automated labeling when the question asks about minimizing cost via sample selection.

How to eliminate wrong answers

Option B is wrong because using Mechanical Turk for all labeling maximizes human labeling volume and cost without intelligent sample selection — it does not minimize cost. Option C is wrong because pre-built bounding box workflows are just the annotation UI; they do not reduce the number of samples needing labels. Option D is wrong because 'automated data labeling with AWS Lambda' is not a Ground Truth cost-optimization feature — Ground Truth's automated labeling uses active learning plus a model, not Lambda.

456
MCQmedium

A trained model needs to be deployed for real-time inference with low latency. Which AWS service is best suited for this?

A.SageMaker Batch Transform
B.SageMaker endpoints
C.SageMaker Hyperparameter Tuning
D.AWS Lambda with model packaged
AnswerB

SageMaker endpoints host trained models behind a persistent HTTPS endpoint, delivering the low-latency, real-time inference the stem demands. Unlike batch transform, which processes datasets asynchronously, endpoints keep compute provisioned and respond per-request, satisfying the low-latency constraint directly.

Why this answer

SageMaker endpoints are designed for real-time inference by provisioning persistent, auto-scaled HTTPS endpoints that return predictions with millisecond latency. They support automatic scaling, A/B testing, and can be deployed behind a VPC for low-latency access, making them the ideal choice for serving a trained model in production.

Exam trap

AWS often tests the distinction between batch and real-time inference, and the trap here is that candidates confuse SageMaker Batch Transform (which processes data in bulk) with a real-time serving solution, or they overestimate Lambda's ability to handle large model payloads and sustained low-latency requests.

How to eliminate wrong answers

Option A is wrong because SageMaker Batch Transform is an asynchronous, batch-processing service that processes large datasets in chunks and returns results to S3, not suitable for real-time, low-latency inference. Option C is wrong because SageMaker Hyperparameter Tuning is a model training optimization process that searches for optimal hyperparameters, not a deployment or inference service. Option D is wrong because AWS Lambda has a maximum execution timeout of 15 minutes and a payload limit of 6 MB, making it impractical for hosting large models or handling sustained real-time inference with low latency; it is better suited for lightweight, event-driven tasks.

457
MCQmedium

A machine learning engineer is preparing a training dataset stored in Amazon S3 for a SageMaker training job. The data is in CSV format, and the engineer wants to ensure that the training job can access the data efficiently and securely. The S3 bucket is in the same AWS Region as the SageMaker training job. The engineer needs to provide the training job with the necessary permissions to read the data. Which of the following is the MOST secure and appropriate way to grant the training job access to the S3 bucket?

A.Store AWS credentials (access key and secret key) in the training script and use them to access S3 directly.
B.Create an IAM role with a policy that allows s3:GetObject on the specific S3 bucket and attach it to the SageMaker training job.
C.Make the S3 bucket public and allow the training job to read the data without authentication.
D.Generate a presigned URL for each object in the S3 bucket and pass the URLs as hyperparameters to the training job.
AnswerB

This approach follows the principle of least privilege by granting only the necessary s3:GetObject permission on the specific bucket. The IAM role is assumed by the SageMaker training job, allowing secure access without embedding credentials. It is the recommended practice for SageMaker training jobs to access S3 data.

Why this answer

The correct approach is to create an IAM role with a policy that grants s3:GetObject permission on the specific S3 bucket and attach it to the SageMaker training job. This follows the principle of least privilege and allows the training job to access the data securely without embedding credentials. It is the standard method for granting SageMaker training jobs access to S3 data.

Exam trap

The trap here is assuming that presigned URLs or embedded credentials are acceptable for long-running training jobs, when they introduce security and reliability risks.

458
MCQhard

A data science team is using Amazon SageMaker Pipelines to orchestrate a multi-step workflow that includes data preprocessing, training, and model evaluation. They want to reuse the preprocessed data across multiple pipeline executions without re-running the preprocessing step if the source data hasn't changed. What should they configure?

A.Use SageMaker Training steps with checkpointing
B.Use SageMaker Processing steps with caching
C.Use SageMaker Feature Store to store the preprocessed features
D.Use SageMaker Data Wrangler for the preprocessing
AnswerB

SageMaker Pipelines caches step outputs keyed on the step's code, inputs, and parameters, so unchanged source data lets downstream executions skip preprocessing and reuse the stored artefacts. This directly satisfies the requirement to avoid re-running preprocessing across executions, while training and evaluation still run normally.

Why this answer

SageMaker Processing steps support caching, which allows the pipeline to skip re-execution of the preprocessing step if the input data and pipeline parameters have not changed. This is achieved by configuring a `CacheConfig` with a caching key based on the input data source and step parameters, ensuring that the preprocessed data is reused across multiple pipeline executions without redundant computation.

Exam trap

The trap here is that candidates may confuse checkpointing (for training resumption) with caching (for step reuse), or assume that Feature Store or Data Wrangler inherently provide caching, when in fact only Processing steps with explicit CacheConfig enable this behavior in SageMaker Pipelines.

How to eliminate wrong answers

Option A is wrong because SageMaker Training steps with checkpointing are designed to save intermediate model state during training (e.g., for resuming from failures), not to cache or reuse preprocessed data across pipeline executions. Option C is wrong because SageMaker Feature Store is a managed repository for storing, sharing, and managing features for ML models, but it does not automatically cache the output of a preprocessing step; it requires explicit feature ingestion and retrieval, which adds complexity and does not directly address skipping the preprocessing step based on unchanged source data. Option D is wrong because SageMaker Data Wrangler is a visual interface for data preparation and feature engineering, but it does not provide built-in caching for pipeline steps; it can be used within a Processing step, but the caching behavior is a property of the Processing step itself, not of Data Wrangler.

459
MCQmedium

A data scientist is training a large model on SageMaker and wants to reduce training time by using multiple GPUs. The model is small enough to fit on a single GPU but training is slow. Which SageMaker feature should be used?

A.Data parallelism using SageMaker's Distributed Data Parallel
B.Use a larger instance with more vCPUs
C.Model parallelism using SageMaker's Model Parallel
D.Use Elastic Inference
AnswerA

Data parallelism replicates the small model across multiple GPUs, each processing a shard of the batch, then synchronises gradients via SageMaker's Distributed Data Parallel. This cuts training time by using several GPUs concurrently, matching the requirement.

Why this answer

SageMaker's Distributed Data Parallel (DDP) is the correct choice because it splits the mini-batch across multiple GPUs, allowing each GPU to hold a copy of the model and process a subset of the data simultaneously. This reduces training time for models that fit on a single GPU by leveraging data parallelism, where gradients are synchronized across GPUs after each step.

Exam trap

The trap here is that candidates confuse model parallelism (for large models) with data parallelism (for slow training of small models), or mistakenly think Elastic Inference can accelerate training when it is strictly for inference latency reduction.

How to eliminate wrong answers

Option B is wrong because using a larger instance with more vCPUs does not directly accelerate GPU-bound training; the bottleneck is GPU compute, not CPU cores. Option C is wrong because model parallelism is designed for models that are too large to fit on a single GPU, partitioning layers across devices, which adds communication overhead and is unnecessary when the model fits on one GPU. Option D is wrong because Elastic Inference attaches a separate accelerator for inference only, not for training, and cannot be used to speed up training loops.

460
MCQeasy

A data scientist is using SageMaker built-in XGBoost algorithm for a binary classification task. Which objective metric is MOST appropriate for SageMaker Automatic Model Tuning to maximize?

A.validation:mae
B.validation:rmse
C.validation:ndcg
D.validation:auc
AnswerD

For binary classification, `validation:auc` maximises the area under the ROC curve, giving a threshold-independent measure of class separation. This satisfies the stem's requirement for the most appropriate tuning objective, since AUC handles imbalanced binary labels better than accuracy and is natively supported by SageMaker Automatic Model Tuning.

Why this answer

For binary classification, AUC (Area Under the ROC Curve) is the standard evaluation metric because it measures the model's ability to discriminate between the two classes across all classification thresholds, independent of class balance. SageMaker's built-in XGBoost exposes validation:auc as the objective metric for binary classification tuning. MAE and RMSE are regression metrics, and NDCG is a ranking metric, so none of them fit binary classification.

Exam trap

The trap is mixing regression metrics (MAE, RMSE) and ranking metrics (NDCG) into a binary classification question — candidates who do not map metric families to task types will pick a plausible-sounding but wrong objective.

How to eliminate wrong answers

Option A is wrong because validation:mae (mean absolute error) is a regression metric measuring average absolute prediction error, which is meaningless for binary class labels. Option B is wrong because validation:rmse (root mean squared error) is also a regression metric, penalizing large errors in continuous predictions, not classification correctness. Option C is wrong because validation:ndcg (normalized discounted cumulative gain) is a ranking-quality metric used for learning-to-rank tasks, not binary classification.

461
MCQmedium

A machine learning engineer is training a deep learning model on SageMaker and notices that the training loss decreases rapidly in the first few epochs but then plateaus. The validation loss starts increasing after 10 epochs. Which action should the engineer take to improve generalization?

A.Add more layers to the model
B.Use early stopping with validation loss monitoring
C.Increase the learning rate
D.Decrease the batch size
AnswerB

Early stopping halts training once validation loss stops improving, directly countering the divergence seen after epoch 10 where training loss keeps falling but validation loss rises. This regularisation technique restores the best checkpoint, preventing overfitting and improving generalisation without altering the model architecture or dataset.

Why this answer

Early stopping is the correct action because the validation loss increasing after 10 epochs while training loss continues to decrease is a classic sign of overfitting. By monitoring validation loss and halting training when it stops improving (e.g., using a patience parameter), the engineer prevents the model from memorizing noise in the training data, thereby improving generalization. SageMaker's built-in training job features or the `EarlyStopping` callback in frameworks like TensorFlow or PyTorch can implement this directly.

Exam trap

AWS often tests the distinction between underfitting and overfitting symptoms, and the trap here is that candidates mistake a plateauing training loss for a need to increase model complexity or learning rate, when the rising validation loss clearly signals overfitting that early stopping can mitigate.

How to eliminate wrong answers

Option A is wrong because adding more layers increases model capacity, which would exacerbate overfitting and likely cause validation loss to rise even sooner, not improve generalization. Option C is wrong because increasing the learning rate would make training more unstable, potentially causing the loss to diverge or oscillate, and would not address the overfitting indicated by the rising validation loss. Option D is wrong because decreasing the batch size introduces more noise into gradient estimates, which can sometimes help escape local minima but does not directly prevent overfitting; it may even slow convergence and does not target the core issue of validation loss increasing.

462
MCQhard

A machine learning engineer manages a SageMaker Model Monitor schedule for a real-time endpoint. The monitor's baseline was computed from a training dataset with a categorical feature named region. In production, a new category value appears that was never seen in training, and the monitor begins reporting violations. The engineer wants the monitor to flag only the appearance of unknown categories without treating normal distribution shifts in known categories as violations. Which approach should the engineer take?

A.Recreate the baseline with a larger sample that includes the new category, then set the monitor's comparison threshold to zero for all features.
B.Disable the constraint checks for the feature and rely solely on the monitor's distribution comparison to detect the new category.
C.Enable the monitor's constraint on the categorical feature so that only values present in the baseline are allowed, and keep distribution comparison thresholds relaxed for that feature.
D.Increase the monitor's sampling percentage to one hundred percent and lower the KMS key rotation period so the monitor can read all incoming records.
AnswerC

Model Monitor emits constraints from the baseline that list allowed categorical values, so a value absent from the baseline violates the constraint and is flagged as an unknown category. Keeping the distribution comparison threshold relaxed for that feature prevents normal frequency shifts among known categories from generating violations, which is exactly the separation the engineer is asking for.

Why this answer

Model Monitor separates two kinds of checks. Constraints, derived from the baseline statistics, enumerate allowed values for categorical features, so a value never seen in training violates the constraint and is flagged as unknown. Distribution comparison measures frequency drift among values that exist in the baseline.

Keeping constraints enabled while relaxing distribution thresholds isolates unknown-category detection from ordinary frequency shifts.

Exam trap

The trap here is assuming that distribution comparison automatically detects brand-new categorical values, when the explicit allowed-value check that catches them is the constraint check.

463
MCQeasy

An ML engineer needs to compile a trained TensorFlow model to run efficiently on a target edge device with an ARM CPU. Which AWS service should they use?

A.SageMaker Debugger
B.AWS Inferentia
C.SageMaker Neo
D.Amazon Elastic Inference
AnswerC

SageMaker Neo compiles trained models for specific target hardware, including ARM CPUs, producing optimised executables that run efficiently on edge devices. It directly satisfies the stem's requirement to compile a TensorFlow model for an ARM-based edge target.

Why this answer

SageMaker Neo compiles trained models from frameworks like TensorFlow, PyTorch, and MXNet into optimized executables for specific target hardware, including ARM CPUs, Intel, and AWS Inferentia. It performs graph-level optimizations and generates a runtime that runs efficiently on the edge device.

Exam trap

MLA-C01 often tests the distinction between hardware accelerators (Inferentia, Elastic Inference) and the compilation/optimization service (Neo) — candidates pick the chip when the question asks for a service to compile a model.

How to eliminate wrong answers

Option A is wrong because SageMaker Debugger is a training-time tool for inspecting tensors, gradients, and metrics — it does not compile or optimize models for edge deployment. Option B is wrong because AWS Inferentia is a custom inference chip, not a compilation service; it is a hardware target, not a tool for compiling TensorFlow models. Option D is wrong because Amazon Elastic Inference attaches GPU-powered inference acceleration to EC2/SageMaker instances in the cloud, not to ARM edge devices, and it is being deprecated.

464
MCQeasy

A data scientist wants to track the training and validation accuracy of a SageMaker training job over time. They need to visualize these metrics in Amazon CloudWatch. Which action should they take?

A.Configure the training job to write metrics directly to Amazon CloudWatch using the AWS SDK for Python (Boto3) in the training script.
B.Use the SageMaker estimator's metric_definitions parameter to specify regex patterns that extract metrics from the training logs.
C.Write the metrics to a file in the output S3 bucket and create a CloudWatch dashboard from the file.
D.Enable SageMaker Debugger and configure rules to emit metrics to CloudWatch.
AnswerB

SageMaker automatically streams training logs to CloudWatch Logs. By defining metric_definitions with regex patterns, SageMaker extracts the specified metrics from the logs and publishes them as CloudWatch metrics. This enables real-time monitoring and visualization without custom code.

Why this answer

The metric_definitions parameter in the SageMaker estimator allows you to define regex patterns that extract metric values from the training logs. SageMaker then automatically publishes these as CloudWatch metrics, enabling visualization. This is the recommended and simplest approach for tracking standard training metrics.

Exam trap

The trap here is overcomplicating the solution by considering custom code or other SageMaker features, when the built-in metric extraction is sufficient.

465
Multi-Selecteasy

A company wants to monitor its Amazon SageMaker real-time endpoint for data quality issues. Which TWO actions should the company take?

Select 2 answers
A.Create a baseline from the training data to compare against live data.
B.Use SageMaker Debugger to analyze training jobs.
C.Set up an AWS Lambda function to preprocess incoming requests.
D.Configure Amazon S3 bucket notifications for model artifacts.
E.Enable data capture on the SageMaker endpoint.
AnswersA, E

A baseline computed from training data defines expected feature distributions, enabling SageMaker Model Monitor to detect drift and data quality violations on live traffic. Without this reference, the endpoint has no statistical standard against which to flag anomalies.

Why this answer

Option A is correct because SageMaker Model Monitor requires a baseline (statistics and constraints) generated from the training dataset, which is then compared against the live inference data to detect data quality drift. Option E is correct because Model Monitor relies on data capture, which must be enabled on the real-time endpoint to record the request and response payloads into an Amazon S3 location for monitoring. Option B is incorrect because SageMaker Debugger analyzes training jobs for issues like vanishing gradients or overfitting, not live endpoint data quality.

Option C is incorrect because a Lambda function for preprocessing requests does not provide data quality monitoring capabilities. Option D is incorrect because S3 bucket notifications on model artifacts only signal object events and do not evaluate inference data quality.

Exam trap

The trap here is that candidates confuse SageMaker Debugger (training-time debugging) with SageMaker Model Monitor (post-deployment data quality), leading them to select Option B instead of recognizing that data capture and baseline creation are the two required actions for endpoint monitoring.

466
MCQmedium

A company uses SageMaker JumpStart to deploy a foundation model for a summarization task. They want to minimize costs while still meeting a latency requirement of under 2 seconds. Which option should they consider?

A.Use SageMaker Inference Recommender to select the cheapest instance that meets latency
B.Deploy the model on a serverless endpoint
C.Enable auto-scaling to handle variable traffic
D.Use the largest GPU instance to ensure fast inference
AnswerA

SageMaker Inference Recommender runs automatic load tests across instance types and returns the cheapest instance satisfying the sub-2-second latency constraint, directly optimising the cost-versus-latency trade-off. Manual instance selection risks over-provisioning or breaching latency, so this satisfies both the cost-minimisation and latency requirements in the stem.

Why this answer

SageMaker Inference Recommender runs load tests against your model on various instance types and provides latency and cost metrics. By selecting the cheapest instance that still meets the sub-2-second latency requirement, you directly minimize cost while satisfying the performance constraint. This is the most systematic and cost-effective approach for this scenario.

Exam trap

A common misconception is that serverless endpoints are always the cheapest option, but for latency-sensitive workloads with large models, the cold-start overhead and lack of guaranteed compute resources make them unsuitable. Inference Recommender is the correct tool for cost-latency trade-off analysis.

How to eliminate wrong answers

Option B is wrong because serverless endpoints have a cold-start latency that can exceed 2 seconds, especially for large foundation models, and they do not guarantee consistent sub-2-second inference under variable traffic. Option C is wrong because auto-scaling handles variable traffic but does not reduce per-invocation cost or latency; it only adjusts capacity, and the chosen instance type still determines base latency and cost. Option D is wrong because using the largest GPU instance is unnecessarily expensive and may provide excess compute capacity that is not needed to meet a 2-second latency requirement, violating the cost-minimization goal.

467
MCQmedium

A financial services company ingests transaction data from multiple sources into an S3 data lake. They want to use AWS Glue to catalog this data and make it queryable by Amazon Athena. The data schema changes frequently as new sources are added. Which AWS Glue feature should they enable to automatically detect and update the schema?

A.AWS Glue DataBrew
B.AWS Glue crawlers with schema update policy set to 'UPDATE'
C.AWS Glue ETL job scheduled to run daily
D.Manual schema definition in the AWS Glue Data Catalog
AnswerB

Glue crawlers with the UPDATE policy re-run classification against S3 and amend the Data Catalog tables in place, adding new columns without dropping existing ones. This satisfies the frequently changing schema constraint, keeping Athena queries working as new sources introduce fields.

Why this answer

AWS Glue crawlers scan data in S3, infer the schema, and write table definitions to the Glue Data Catalog. By setting the crawler's schema update policy to 'UPDATE', the crawler will modify existing table definitions in the catalog when it detects new columns or changed types, keeping Athena queries working as sources evolve. This is the purpose-built feature for automatically detecting and propagating schema changes.

Exam trap

The trap is confusing data preparation (DataBrew) or ETL (Glue jobs) with cataloging — only crawlers with an UPDATE schema policy automatically detect and propagate schema changes to the Data Catalog.

How to eliminate wrong answers

Option A is wrong because AWS Glue DataBrew is a visual data-preparation tool for profiling and transforming datasets; it does not catalog schemas or update the Data Catalog. Option C is wrong because an ETL job transforms and moves data — it does not perform schema discovery or catalog maintenance, and scheduling it daily does not address schema drift. Option D is wrong because manual schema definition is the opposite of automation; it requires a human to update the catalog every time a source changes, which does not scale with frequently changing schemas.

468
Multi-Selecthard

An ML team uses SageMaker to deploy a model for real-time inference. They want to monitor and improve cost efficiency. Which THREE actions should they take? (Select THREE.)

Select 3 answers
A.Use SageMaker Inference Recommender to find the optimal instance type and count
B.Enable auto-scaling to adjust the number of instances based on demand
C.Create a CloudWatch dashboard to monitor endpoint latency
D.Use SageMaker Managed Spot Training for endpoint instances
E.Purchase SageMaker Savings Plans for a discounted rate
AnswersA, B, E

Inference Recommender benchmarks candidate instance types and counts against the model's actual traffic and latency requirements, recommending the most cost-effective configuration. This directly addresses right-sizing, which is the primary lever for reducing real-time inference cost before scaling or commitment discounts.

Why this answer

Option A is correct because SageMaker Inference Recommender runs load tests and benchmarks to recommend the optimal instance type and instance count for a real-time endpoint, directly improving cost efficiency by avoiding over-provisioning. Option B is correct because configuring auto-scaling on a SageMaker endpoint adjusts the number of instances to match traffic demand, so the team pays only for capacity actually needed during peaks and troughs. Option E is correct because SageMaker Savings Plans offer discounted pricing (up to 64% off) in exchange for a committed hourly spend, reducing the cost of steady-state real-time inference workloads.

Option C is not a cost-efficiency action; a CloudWatch dashboard for latency is an observability tool and does not by itself reduce spend. Option D is incorrect because Managed Spot Training applies to training jobs, not to real-time inference endpoint instances, which cannot use spot capacity for persistent endpoints.

Exam trap

The trap here is that candidates confuse monitoring (Option C) with cost optimization, or they mistakenly apply Spot Training (Option D) to inference endpoints, not realizing that Spot instances are only supported for training and not for real-time inference due to interruption risk.

469
MCQeasy

A data science team deploys a machine learning model to a SageMaker endpoint for real-time inference. They need to monitor the model for feature distribution drift over time to ensure the model's predictions remain accurate. Which AWS service should they use?

A.Amazon CloudWatch Evidently
B.AWS Glue DataBrew
C.SageMaker Clarify
D.SageMaker Model Monitor
E.SageMaker Debugger
AnswerD

SageMaker Model Monitor directly satisfies the requirement to detect feature distribution drift on a live endpoint. It captures real-time inference request data, computes baseline statistics, and compares them against production traffic using configurable drift thresholds, emitting CloudWatch metrics and alerts when divergence is detected.

Why this answer

SageMaker Model Monitor is the correct service because it is specifically designed to continuously monitor machine learning models deployed to SageMaker endpoints for data quality issues, including feature distribution drift. It automatically captures inference data, computes statistics against a baseline, and triggers alerts when drift is detected, ensuring the model's predictions remain accurate over time.

Exam trap

The trap here is confusing SageMaker Model Monitor with SageMaker Clarify or Debugger, as candidates often misattribute drift monitoring to Clarify's bias detection or Debugger's training-time analysis, but only Model Monitor handles post-deployment feature drift.

How to eliminate wrong answers

Option A is wrong because Amazon CloudWatch Evidently is a feature flag and A/B testing service, not designed for monitoring feature distribution drift in ML models. Option B is wrong because AWS Glue DataBrew is a visual data preparation tool for cleaning and normalizing data, not for monitoring model drift. Option C is wrong because SageMaker Clarify is used for bias detection and explainability of model predictions, not for continuous drift monitoring.

Option E is wrong because SageMaker Debugger is used for debugging training jobs by monitoring system and model metrics during training, not for monitoring inference data drift post-deployment.

470
MCQeasy

A machine learning engineer needs to split a dataset for binary classification where the positive class represents only 2% of the data. Which data splitting strategy ensures that both training and test sets maintain the same class proportion as the original dataset?

A.Stratified sampling based on the target variable
B.Time-series split respecting the timestamp order
C.Simple random split with a 80/20 ratio
D.K-fold cross-validation without stratification
AnswerA

Stratified sampling partitions rows so each split preserves the original 2% positive-class prevalence, preventing a random split from yielding a test set with zero or distorted minority cases. This directly satisfies the stem's constraint that training and test sets mirror the dataset's class proportion, which matters acutely at severe imbalance.

Why this answer

Stratified sampling explicitly preserves the class distribution of the original dataset in each split. For a binary classification problem with a 2% positive class, stratification ensures that both the training and test sets contain approximately 2% positive examples, preventing a scenario where random chance could yield a test set with zero or very few positives. This is critical for imbalanced datasets because it maintains representativeness and allows reliable evaluation of metrics like recall or F1-score for the minority class.

Exam trap

MLA-C01 often tests the misconception that any random split preserves class proportions, but with imbalanced data, random splits can easily distort the minority class distribution, making stratified sampling the only reliable choice.

How to eliminate wrong answers

Option B is wrong because time-series split respects temporal order and is used for time-dependent data, not for preserving class proportions; it would not guarantee the same positive rate in each fold. Option C is wrong because a simple random split does not control for class distribution, so with a 2% positive rate, random variation could easily produce splits with significantly different proportions, especially with small datasets. Option D is wrong because k-fold cross-validation without stratification also fails to preserve class proportions; each fold may have varying numbers of positive examples, leading to unreliable performance estimates for imbalanced classification.

471
Multi-Selectmedium

A data scientist is training a binary classification model using Amazon SageMaker. The dataset is highly imbalanced (95% negative class, 5% positive class). The model is evaluated on a held-out test set, and the F1 score is 0.12. The data scientist wants to improve the F1 score. Which two actions should the data scientist take? (Choose two.)

Select 2 answers
A.Reduce the model complexity by decreasing the number of layers in a deep neural network.
B.Apply SMOTE (Synthetic Minority Oversampling Technique) to the training data using a preprocessing script in SageMaker Processing.
C.Increase the decision threshold to reduce false positives.
D.Use recall as the primary evaluation metric instead of F1.
E.Set the `scale_pos_weight` parameter in the SageMaker XGBoost estimator to the ratio of negative to positive samples.
AnswersB, E

SMOTE synthesises new minority-class examples by interpolating between existing positive samples and their nearest neighbours, rebalancing the training distribution. With only 5% positives, the model otherwise predicts the majority class, so oversampling directly lifts recall and therefore F1 on the held-out test set.

Why this answer

Option B is correct because SMOTE generates synthetic examples of the minority (positive) class, rebalancing the 95/5 training distribution so the model learns the minority class decision boundary instead of being dominated by the negative class, which directly improves precision and recall (and thus F1) on the held-out test set; this can be implemented via a SageMaker Processing job with an imbalanced-learn script. Option E is correct because setting scale_pos_weight in the SageMaker XGBoost estimator to the ratio of negative to positive samples (roughly 95/5 = 19) upweights the minority class in the gradient boosting loss, making the model penalize minority-class misclassifications more heavily and improving F1 without altering the data. Option A is not appropriate because reducing neural network depth addresses overfitting, not class imbalance, and could even underfit and lower F1.

Option C is wrong because raising the decision threshold reduces recall and typically lowers F1 in a heavily imbalanced setting. Option D is wrong because switching the evaluation metric to recall does not improve the model's F1 score; it only changes what is reported.

Exam trap

The trap here is that candidates often confuse threshold tuning with addressing imbalance directly, not realizing that adjusting the threshold without rebalancing the data or weighting classes typically fails to improve F1 score because it does not change the underlying model's learned distribution.

472
MCQmedium

A company runs a real-time fraud detection model on a SageMaker endpoint. The model is updated weekly, and each update must be validated against live traffic without affecting existing predictions. The team wants to compare the new model's performance against the current model using a small percentage of incoming requests, while ensuring that the current model continues to serve the majority of traffic. Which SageMaker deployment strategy should they use?

A.A/B testing with production variants, where the new model variant receives a small portion of traffic.
B.Canary deployment, where the new model is deployed to a separate endpoint and traffic is gradually shifted using a load balancer.
C.Blue/green deployment, where traffic is shifted all at once from the old model to the new model after validation.
D.Shadow testing, where the new model receives a copy of live traffic but its predictions are not returned to the caller.
AnswerA

A/B testing with production variants allows you to deploy multiple models to the same endpoint and distribute traffic between them. By assigning a small weight to the new variant, you can evaluate its performance on live traffic while the existing model handles the rest. This meets the requirement of validating the new model without impacting the majority of predictions.

Why this answer

SageMaker production variants enable A/B testing by allowing multiple models on a single endpoint with configurable traffic weights. This lets the team direct a small percentage of live traffic to the new model while the current model serves the rest, facilitating performance comparison without disrupting service. Other strategies either do not return predictions or switch all traffic at once.

Exam trap

The trap here is confusing shadow testing with A/B testing, as shadow testing does not return predictions to the caller and thus cannot be used for live performance validation.

473
MCQeasy

A hospital's ML team deploys a diagnostic model to a SageMaker real-time endpoint. Compliance requires that every inference request and response be recorded for later auditing, and the records must be retrievable months later. The team needs a low-effort way to capture this data. Which approach should they use?

A.Turn on AWS CloudTrail data events for the endpoint and deliver them to an S3 bucket.
B.Enable data capture on the endpoint with a sampling percentage of 100 and an S3 destination, and apply an S3 lifecycle policy to retain the objects for the required period.
C.Enable Amazon CloudWatch metrics on the endpoint and export the metrics to Amazon S3 for archival.
D.Add application code that writes each request and response to Amazon CloudWatch Logs and set the log group retention to the audit period.
AnswerB

Data capture records request and response payloads to Amazon S3 automatically, and setting the sampling percentage to 100 captures every invocation, satisfying an audit requirement for complete records. Pointing capture at an S3 prefix and applying a lifecycle policy retains the objects for the mandated duration. This requires minimal application changes and is the purpose-built feature for this need.

Why this answer

SageMaker data capture is the purpose-built feature for recording inference request and response payloads to Amazon S3, and a 100 percent sampling rate ensures every invocation is stored for audit. Pairing the capture destination with an S3 lifecycle policy retains records for the required period at low cost. Metrics, CloudTrail, and application-level logging do not capture full payloads as directly.

Exam trap

The trap here is confusing request-level API auditing via CloudTrail or numeric monitoring via CloudWatch metrics with full payload capture, which only data capture provides.

474
MCQeasy

A company is using SageMaker Pipelines to automate a multi-step ML workflow. The pipeline includes data preprocessing, training, and model evaluation. The team wants to ensure that if the evaluation step fails, the pipeline stops and sends an alert to the operations team. Which SageMaker Pipelines feature should they use?

A.Configure an Amazon CloudWatch Events rule to monitor the pipeline execution status and stop it if the evaluation step fails
B.Register the model in the Model Registry only if evaluation passes, and configure the pipeline to stop if registration fails
C.Add a Lambda step after the evaluation step that checks the evaluation metrics and sends an SNS notification if the metrics are below a threshold
D.Use a Condition step to check the evaluation result and route to a Fail step if the result indicates failure
AnswerD

A Condition step evaluates the evaluation step's output against a threshold and, on failure, routes execution to a Fail step that terminates the pipeline and triggers the alert. This satisfies the stem's requirement to stop and notify operations.

Why this answer

SageMaker Pipelines provides a built-in Condition step that evaluates a boolean expression (e.g., checking if evaluation metrics meet a threshold) and then routes execution to different steps. If the condition fails, you can direct the pipeline to a Fail step, which immediately stops the pipeline and marks it as failed. This is the native, event-driven way to halt a pipeline based on step output without relying on external services.

Exam trap

The trap here is that candidates often confuse external monitoring (CloudWatch) or post-step actions (Lambda) with native pipeline control flow, missing that SageMaker Pipelines has a dedicated Condition step for conditional branching and halting execution.

How to eliminate wrong answers

Option A is wrong because CloudWatch Events rules can monitor pipeline state changes but cannot stop a running pipeline; they can only trigger notifications or invoke other actions after the fact. Option B is wrong because registering a model in the Model Registry is an optional downstream step, not a mechanism to stop the pipeline; if registration fails, the pipeline would still continue to subsequent steps unless explicitly handled. Option C is wrong because a Lambda step can send SNS notifications but does not have the ability to halt the pipeline execution; it would only alert after the step completes, not prevent further steps from running.

475
Multi-Selecthard

Which THREE steps should be taken to optimize a large-scale distributed training job on SageMaker? (Choose 3.)

Select 3 answers
A.Attach multiple EBS volumes with throughput provisioning.
B.Use GPU instances with high bandwidth and memory (e.g., ml.p4d.24xlarge).
C.Enable batch transform for offline inference after training.
D.Use Elastic Fabric Adapter (EFA) for low-latency inter-node communication.
E.Select the appropriate distributed training strategy (e.g., Horovod, SageMaker data parallel, or model parallel).
AnswersB, D, E

High-bandwidth GPU instances such as ml.p4d.24xlarge provide the inter-node communication throughput and per-device memory that distributed training demands, directly addressing the stem's large-scale constraint. Gradient synchronisation across many workers is bandwidth-bound, so faster interconnect and larger memory reduce communication stalls and prevent out-of-memory failures during training.

Why this answer

Option B is correct because large-scale distributed training is compute- and memory-intensive, so GPU instances like ml.p4d.24xlarge provide high-bandwidth networking (up to 400 Gbps) and large GPU memory (A100 40 GB each) that accelerate training and reduce communication bottlenecks. Option D is correct because Elastic Fabric Adapter (EFA) enables low-latency, high-throughput inter-node communication using OS bypass, which is essential for scaling distributed training across many instances. Option E is correct because choosing the right distributed training strategy — Horovod for data parallelism, SageMaker distributed data parallel for optimized AllReduce, or model parallel for models too large for a single GPU — directly determines training efficiency and scalability.

Option A is not appropriate because attaching multiple EBS volumes with provisioned throughput addresses storage I/O, not the inter-node GPU communication and compute parallelism that dominate large-scale distributed training. Option C is not appropriate because batch transform is an inference-time feature and does not optimize the training job itself.

Exam trap

The trap here is that candidates confuse storage optimization (EBS) or inference features (batch transform) with training optimization, failing to recognize that distributed training performance hinges on compute, memory, and inter-node communication, not disk I/O or post-training steps.

476
MCQhard

A data science team uses SageMaker Pipelines to orchestrate their ML workflow. They noticed that even when source data hasn't changed, the pipeline re-runs all steps, wasting compute time. What should they enable to avoid redundant runs?

A.Enable pipeline caching by setting the CacheConfig property for each step
B.Configure the pipeline to run on a schedule instead of on-demand
C.Use the Parameter step to pass previous execution ID
D.Use Lambda step to check data changes before running
AnswerA

CacheConfig on each pipeline step lets SageMaker Pipelines skip re-execution when inputs, code and parameters are unchanged, reusing prior outputs. Enabling it satisfies the stem's requirement to avoid redundant runs and wasted compute when source data has not changed.

Why this answer

SageMaker Pipelines supports step caching via the `CacheConfig` property. When enabled, the pipeline checks if the step's inputs (including source data, parameters, and code) have changed since the last successful run. If no changes are detected, the step is skipped and the previous output is reused, eliminating redundant compute.

Exam trap

The trap here is that candidates may think caching requires external logic (like a Lambda step) or scheduling, when SageMaker Pipelines has a native `CacheConfig` property that directly addresses redundant runs with minimal configuration.

How to eliminate wrong answers

Option B is wrong because scheduling the pipeline does not prevent redundant runs; it only triggers execution at fixed intervals, which could still re-run all steps even when data hasn't changed. Option C is wrong because passing a previous execution ID via a Parameter step does not enable caching; it merely provides a reference but does not automatically skip unchanged steps. Option D is wrong because using a Lambda step to check data changes before running adds custom logic but is not a built-in mechanism for step-level caching; SageMaker Pipelines already provides `CacheConfig` for this purpose, making a Lambda workaround unnecessary and less efficient.

477
MCQmedium

A company is building a time series forecasting model using SageMaker DeepAR. The raw data is a CSV with columns: timestamp, item_id, and value. What is the correct data format required for DeepAR training?

A.JSON Lines files with 'start', 'target', and optional fields per time series
B.A wide-format CSV where each column is a different time series
C.Parquet files with a schema containing timestamp, item_id, and value
D.A single CSV file with columns: timestamp, item_id, value
AnswerA

DeepAR consumes JSON Lines, where each line holds one time series with a 'start' timestamp and a 'target' array of values. The stem's CSV columns must therefore be pivoted into that per-series structure; CSV is not accepted directly by the algorithm.

Why this answer

DeepAR requires time series data to be provided in JSON Lines format, where each line represents a single time series with a 'start' timestamp (in ISO 8601 format), a 'target' array of values, and optional fields like 'cat' for categorical features. This structured format allows DeepAR to handle variable-length sequences and missing values natively, which is not possible with simple CSV or wide-format data.

Exam trap

The trap here is that candidates assume DeepAR can accept raw CSV data like other SageMaker built-in algorithms (e.g., XGBoost), but DeepAR is a specialized time series algorithm that requires a specific JSON Lines structure with 'start' and 'target' fields, not a simple tabular format.

How to eliminate wrong answers

Option B is wrong because wide-format CSV (each column as a separate time series) is not supported by DeepAR; it expects each time series to be a separate JSON object, not columns. Option C is wrong because Parquet files are not a native input format for DeepAR; the built-in algorithm specifically requires JSON Lines or RecordIO-protobuf format. Option D is wrong because a single CSV with timestamp, item_id, and value columns does not provide the 'start' and 'target' structure DeepAR needs; it would require significant preprocessing to group by item_id and convert to the required JSON Lines format.

478
MCQeasy

A data scientist is preparing a dataset for a machine learning model that predicts customer churn. The dataset contains a column 'CustomerID' that is a unique identifier. What should the data scientist do with this column before training the model?

A.Keep the column as a feature because it uniquely identifies each customer.
B.Use the column as the target variable.
C.Remove the column from the feature set.
D.Encode the column using one-hot encoding.
AnswerC

CustomerID is a unique identifier carrying no generalisable signal; each value appears once, so the model would memorise rather than learn. Removing it from the feature set prevents noise and spurious splits, satisfying the requirement to prepare the churn dataset correctly.

Why this answer

'CustomerID' is a unique identifier with no predictive power for churn. Including it as a feature would cause the model to memorize individual customers rather than learn generalizable patterns, leading to overfitting and poor performance on unseen data. In machine learning, such columns should be removed during data preparation to ensure the model learns from meaningful features.

Exam trap

The trap here is that candidates may think unique identifiers are useful for tracking or that they can be encoded as categorical features, but the exam tests the principle that identifiers with no predictive relationship to the target must be removed to avoid overfitting and data leakage.

How to eliminate wrong answers

Option A is wrong because keeping 'CustomerID' as a feature introduces a high-cardinality categorical variable with no correlation to the target, which can cause overfitting and degrade model generalization. Option B is wrong because the target variable for churn prediction should be a binary or categorical label indicating churn status, not a unique identifier that has no relationship to the outcome. Option D is wrong because one-hot encoding a unique identifier like 'CustomerID' would create thousands of sparse binary columns, dramatically increasing dimensionality without adding any predictive value, and is computationally wasteful.

479
MCQmedium

A team uses SageMaker Pipelines to automate retraining. They want to skip the training step if the data has not changed since the last run. Which feature should they enable?

A.Parameterized executions
B.Lineage tracking
C.Step caching
D.Condition step with a custom check
AnswerC

Step caching stores the step's inputs and outputs, so when the data and hyperparameters match a previous run, SageMaker reuses the cached artefacts and skips execution entirely. This directly satisfies the requirement to bypass retraining when the dataset is unchanged, avoiding redundant compute costs.

Why this answer

Step caching in SageMaker Pipelines allows you to reuse the output from a previous execution of a step if its input data and configuration parameters have not changed. By enabling caching on the training step, the pipeline automatically skips re-executing that step when the data is identical, saving time and cost. This directly addresses the requirement to skip retraining when data has not changed.

Exam trap

Candidates often confuse step caching (automatic, built-in) with a Condition step (manual, custom logic), leading them to overthink and choose the more complex option D when the simpler caching feature is the correct answer.

How to eliminate wrong answers

Option A is wrong because parameterized executions allow you to pass different parameters into a pipeline run, but they do not automatically skip steps based on data changes; they simply enable dynamic input values. Option B is wrong because lineage tracking records the relationships between artifacts and steps for governance and reproducibility, but it does not provide any mechanism to skip step execution. Option D is wrong because while a Condition step can branch pipeline execution based on a custom check, it requires you to implement the logic to compare data versions manually, whereas step caching provides built-in, automatic detection of unchanged inputs.

480
Multi-Selecthard

A data scientist is using SageMaker Experiments to track multiple training runs. They want to compare runs based on the objective metric and visualize performance. Which THREE steps should they perform? (Choose THREE.)

Select 3 answers
A.Deploy the best model to an endpoint
B.Use SageMaker Studio Experiments UI to list and compare trials
C.Log hyperparameters and metrics using the SageMaker SDK
D.Create a SageMaker Experiment
E.Enable SageMaker Model Monitor for each run
AnswersB, C, D

The SageMaker Studio Experiments UI lists trials and plots objective metrics side by side, enabling direct comparison and visualisation of performance across runs. This satisfies the requirement to compare runs by objective metric and visualise their results.

Why this answer

Option D is correct because a SageMaker Experiment is the top-level container that groups related trials (training runs) so they can be tracked and compared together. Option C is correct because the SageMaker SDK (e.g., Run, Trial, Tracker, or the log_parameters/log_metric calls) is how hyperparameters and objective metrics get recorded for each run, which is required before any comparison can be made. Option B is correct because the SageMaker Studio Experiments UI lets the data scientist list trials and compare them visually by objective metric, directly satisfying the stated goal of comparing runs and visualizing performance.

Option A is not required because deploying the best model to an endpoint is an inference/hosting step, not part of tracking or comparing experiments. Option E is not required because Model Monitor is used for detecting drift and data quality issues on deployed endpoints, not for logging or comparing training-run metrics.

Exam trap

MLA-C01 often tests whether candidates conflate Experiment tracking with deployment or monitoring, causing them to select Model Monitor or endpoint deployment as part of the comparison workflow.

481
MCQeasy

A company wants to deploy a machine learning model that was trained on-premises using TensorFlow. The model is a TensorFlow SavedModel. The company uses AWS and wants to minimize operational overhead. Which deployment option meets these requirements?

A.Deploy the model on Amazon ECS using a custom Docker image.
B.Deploy the model as an AWS Lambda function with the TensorFlow runtime.
C.Deploy the model using Amazon SageMaker Studio.
D.Deploy the model using Amazon SageMaker with a TensorFlow inference container.
AnswerD

SageMaker's TensorFlow inference container natively loads TensorFlow SavedModel artefacts, so the on-premises-trained model deploys without custom serving code or infrastructure management. This satisfies the minimal-operational-overhead constraint, unlike self-managed options such as EC2 or ECS that require patching and scaling work.

Why this answer

Amazon SageMaker provides a fully managed TensorFlow inference container that directly supports TensorFlow SavedModel format, enabling deployment without any custom infrastructure management. This minimizes operational overhead compared to self-managed options like ECS or Lambda, as SageMaker handles scaling, load balancing, and model updates automatically.

Exam trap

AWS often tests the distinction between SageMaker Studio (an IDE) and SageMaker hosting (deployment endpoints), leading candidates to mistakenly select Studio as a deployment option when it is only for development and experimentation.

How to eliminate wrong answers

Option A is wrong because deploying on Amazon ECS with a custom Docker image requires you to build, maintain, and scale the container infrastructure yourself, increasing operational overhead. Option B is wrong because AWS Lambda has a maximum deployment package size limit (250 MB unzipped) and a 15-minute timeout, making it unsuitable for large TensorFlow SavedModels or inference requests that require significant compute. Option C is wrong because Amazon SageMaker Studio is an integrated development environment (IDE) for building, training, and debugging models, not a deployment target; the actual deployment would still require creating an endpoint, which is covered by Option D.

482
MCQhard

A machine learning engineer deploys a new model version to a SageMaker endpoint with production variants. They want to gradually shift traffic from the old model to the new model, monitoring for errors, and automatically roll back if the error rate exceeds 5%. Which deployment pattern should they use?

A.Canary deployment with CloudWatch alarms
B.A/B testing with traffic splitting
C.Blue/green deployment
D.Shadow testing
AnswerA

A canary deployment routes a small percentage of traffic to the new variant while CloudWatch alarms watch the error rate, triggering automatic rollback when it exceeds 5%. It satisfies the stem's gradual-shift and auto-rollback constraints, unlike all-at-once or shadow patterns.

Why this answer

Canary deployment with CloudWatch alarms is correct because SageMaker production variants allow you to split traffic between the old and new model (e.g., 90/10), and CloudWatch alarms can monitor the new variant's error rate and trigger automatic rollback when it exceeds 5%. This pattern is purpose-built for gradual, monitored traffic shifting with automated safety nets.

Exam trap

MLA-C01 often tests the distinction between canary (gradual shift + auto rollback) and A/B testing (statistical comparison) — candidates pick A/B testing because both involve traffic splitting, but only canary is designed for progressive rollout with automated rollback on error thresholds.

How to eliminate wrong answers

Option B is wrong because A/B testing with traffic splitting is designed to compare model performance statistically (e.g., conversion rates) rather than to gradually shift all traffic with automatic rollback on error thresholds. Option C is wrong because blue/green deployment shifts traffic all at once (or in a single cutover) between two identical environments, which does not provide the gradual, monitored ramp-up the scenario requires. Option D is wrong because shadow testing mirrors live traffic to the new model without serving its responses to users, so it cannot be used to gradually shift production traffic or trigger rollback based on user-facing error rates.

483
MCQeasy

A data scientist wants to quickly build a binary classification model without writing any code. Which SageMaker feature is MOST suitable?

A.SageMaker Debugger
B.SageMaker Model Monitor
C.SageMaker Ground Truth
D.SageMaker Autopilot
AnswerD

SageMaker Autopilot automates model selection, feature engineering and hyperparameter tuning, then generates candidate models with full visibility, satisfying the no-code requirement. It handles binary classification directly, so the data scientist obtains a deployable model without authoring training scripts.

Why this answer

SageMaker Autopilot is a no-code AutoML feature that automatically explores data, selects algorithms, tunes hyperparameters, and builds the best model for a given tabular dataset, including binary classification. It requires no model-building code from the user, making it the most suitable choice. The other options are operational/monitoring or labeling tools, not model-building features.

Exam trap

The trap is that several SageMaker features have 'model' in their name (Model Monitor, Debugger), so candidates associate them with model building — but only Autopilot actually creates models without code.

How to eliminate wrong answers

Option A is wrong because SageMaker Debugger is a training-time debugging and profiling tool that inspects tensors and metrics during training — it does not build models. Option B is wrong because SageMaker Model Monitor detects data drift and quality issues on deployed models; it is a post-deployment monitoring service, not a model builder. Option C is wrong because SageMaker Ground Truth is a data labeling service for creating training datasets, not a model-building feature.

484
MCQmedium

A data scientist needs to ingest streaming clickstream data from a website into an S3 data lake for ML training. The data must be processed in near real-time and partitioned by hour. Which AWS service combination should be used?

A.Amazon Kinesis Data Firehose with S3 as destination and dynamic partitioning enabled
B.Amazon S3 Transfer Acceleration with direct uploads from the website
C.Amazon Kinesis Data Streams with a custom consumer writing to S3
D.AWS Glue ETL job reading from Kinesis Data Streams and writing to S3
AnswerA

Kinesis Data Firehose delivers streaming data to S3 and supports dynamic partitioning, which derives partition keys from incoming record fields. This satisfies the hourly partitioning requirement without custom consumers, while buffering provides the near real-time ingestion the clickstream pipeline needs.

Why this answer

Amazon Kinesis Data Firehose can directly stream data to S3 with dynamic partitioning, enabling automatic hourly partitioning for near real-time ingestion. Option A is correct. Option B (S3 Transfer Acceleration) is for speeding up large uploads over long distances, not for streaming or partitioning.

Option C requires a custom consumer to write to S3, adding complexity without automatic partitioning. Option D uses a batch-oriented ETL job, not designed for near real-time streaming.

485
MCQeasy

A machine learning engineer needs to split a dataset into training, validation, and test sets for a SageMaker training job. The dataset is stored in Amazon S3 as a single CSV file. The engineer wants to ensure that the splits are reproducible and that the test set is never used during training or hyperparameter tuning. Which approach should the engineer use?

A.Manually split the CSV file using a Python script with a fixed random seed, write the three splits to separate S3 prefixes, and use the training and validation prefixes in the training job, keeping the test prefix for final evaluation.
B.Use the SageMaker training job's built-in data splitting feature by specifying a validation split percentage in the hyperparameters.
C.Use SageMaker Data Wrangler to split the data into three parts and export them to S3, then use all three parts in the training job with different channels.
D.Use Amazon Athena to run a query that randomly assigns rows to three splits, and store the results in S3. Then use all three splits in the training job.
AnswerA

Using a Python script with a fixed random seed ensures reproducibility. Writing the splits to separate S3 prefixes allows the training job to access only the training and validation data, while the test set remains untouched for final evaluation. This meets all requirements.

Why this answer

To ensure reproducibility and prevent data leakage, the engineer should split the data with a fixed random seed and store each split in separate S3 locations. The training job should only consume the training and validation sets, while the test set is reserved for final model evaluation. This is a standard best practice for ML workflows.

Exam trap

The trap here is assuming that SageMaker training jobs automatically split data or that using all splits in training is acceptable.

486
MCQmedium

A machine learning team stores training data in an Amazon S3 bucket and wants to catalog it so that Amazon Athena and Amazon SageMaker Feature Store can discover the schema. The data is partitioned by year, month, and day in Hive-style prefixes, and new partitions are added daily. A data engineer must ensure new partitions are automatically discoverable without manual intervention. Which solution meets these requirements?

A.Enable S3 Inventory and point Athena at the inventory report to enumerate partitions.
B.Run MSCK REPAIR TABLE on an Athena table each time new data arrives.
C.Register the S3 prefix as a SageMaker Feature Store offline store and let it infer the partitions.
D.Create an AWS Glue crawler with a daily schedule that points at the S3 prefix and updates the Data Catalog.
AnswerD

A scheduled Glue crawler scans the S3 location, detects the Hive-style year/month/day prefixes as partition columns, and updates the Data Catalog tables automatically each day. This gives Athena and other integrated services an up-to-date schema and partition list with no manual steps, satisfying the automation requirement.

Why this answer

A scheduled AWS Glue crawler is the standard mechanism for automatically discovering new Hive-style partitions in S3 and refreshing the Data Catalog. Because it runs on a daily schedule, new year/month/day partitions appear in the catalog without any manual command, and Athena and other catalog-aware services can query them immediately.

Exam trap

The trap here is reaching for Athena's partition repair command, which works only on demand and still leaves the catalog stale between runs.

487
MCQmedium

A company has 200 small PyTorch models that are each used infrequently but need to be available for real-time inference. To minimize costs, they want to host all models on a single endpoint. Which SageMaker feature should they use?

A.Multi-model endpoint (MME)
B.Multi-container endpoint
C.Batch Transform job
D.Asynchronous inference endpoint
AnswerA

Multi-model endpoint hosts many models on one endpoint, loading each from S3 on invocation and caching it in memory. Because the PyTorch models are infrequent, this satisfies the single-endpoint and cost-minimisation constraints without provisioning per-model hosting.

Why this answer

SageMaker multi-model endpoints (MME) let a single endpoint host hundreds or thousands of models behind one container, loading each model into memory or disk on demand and unloading idle ones. This is purpose-built for the scenario of many small, infrequently used models that must still serve real-time inference, because you pay for one endpoint's worth of instances rather than one endpoint per model. The SageMaker SDK/API passes a TargetModel parameter on each InvokeEndpoint call so the endpoint knows which model to load and serve.

Exam trap

MLA-C01 often tests the confusion between multi-model endpoints (many models, one container, dynamic load) and multi-container endpoints (few containers, different frameworks, static), so candidates who see 'multiple models' and pick multi-container get it wrong.

How to eliminate wrong answers

Option B (Multi-container endpoint) is wrong because multi-container endpoints are designed to run a small number of containers (typically up to 15) that each host a different framework or pipeline stage, and they do not dynamically load/unload hundreds of models — they are for heterogeneous inference pipelines, not model fleets. Option C (Batch Transform) is wrong because it is an offline, batch-scoring service that spins up a job, processes a dataset, and tears down; it cannot provide the always-available real-time inference the company requires. Option D (Asynchronous inference endpoint) is wrong because async inference is for large payloads/long processing times with queued requests, not for hosting many small models on one endpoint — it still requires one endpoint per model or a single model per endpoint.

488
Multi-Selectmedium

An organization wants to automate ML retraining using an event-driven architecture. Which THREE services should they combine? (Select THREE.)

Select 3 answers
A.SageMaker (training jobs or pipelines)
B.Amazon EventBridge
C.AWS Lambda
D.AWS Glue
E.Amazon CloudWatch Logs
AnswersA, B, C

SageMaker training jobs or pipelines execute the actual retraining computation when triggered. Combined with event detection and orchestration services, it forms the event-driven retraining chain, satisfying the stem's automation requirement. It performs model training rather than storing data or routing events.

Why this answer

Amazon EventBridge [CORRECT] is the event-driven backbone: it receives events (e.g., new data in S3, a schedule, or a custom PutEvents call) and routes them to targets via rules, which is exactly what an event-driven retraining architecture requires. AWS Lambda [CORRECT] acts as the lightweight, serverless compute glue that responds to those EventBridge events and invokes the training workflow, so no servers need to be managed. SageMaker (training jobs or pipelines) [CORRECT] performs the actual model retraining and orchestration, since SageMaker Pipelines can be triggered programmatically to run training jobs and register updated models.

AWS Glue is a data-integration/ETL service and, while useful for preparing data, it is not required to deliver the event-driven retraining trigger or the training itself. Amazon CloudWatch Logs is for storing and monitoring log data and does not provide event routing or ML training capability.

Exam trap

The trap here is that candidates often confuse AWS Glue as a compute trigger for ML retraining, but Glue is designed for batch ETL and lacks the event-driven, low-latency invocation capabilities required for this architecture.

489
Multi-Selectmedium

A machine learning engineer is building a real-time fraud detection pipeline using Amazon Kinesis Data Streams. The data must be prepared (e.g., feature engineering, normalization) before being fed into a SageMaker endpoint. Which TWO steps should the engineer implement to ensure low-latency data preparation?

Select 2 answers
A.Use AWS Lambda functions to apply feature transformations on each record as it arrives.
B.Use SageMaker batch transform jobs scheduled every hour to process the streaming data.
C.Use AWS Glue ETL jobs running on a recurring schedule to transform the data.
D.Use Amazon Kinesis Data Analytics to perform SQL-based transformations on the stream.
E.Use SageMaker Processing jobs to read from Kinesis and write transformed data to S3.
AnswersA, D

Lambda applies feature engineering and normalisation to each record as it arrives from Kinesis, before the SageMaker endpoint is invoked. This keeps transformations in-stream at per-record latency, avoiding batch reprocessing and satisfying the low-latency requirement of the fraud pipeline.

Why this answer

Option A is correct because AWS Lambda can be invoked per record or micro-batch directly from Kinesis Data Streams as an event source, applying feature engineering and normalization inline with minimal latency and no server management. Option D is correct because Amazon Kinesis Data Analytics (Kinesis Data Analytics for SQL/Flink) performs continuous SQL-based transformations on the stream itself, enabling real-time normalization and feature computation before records reach the SageMaker endpoint. Option B is not suitable because SageMaker batch transform jobs run on a schedule against stored data, introducing hourly latency and requiring data to be landed first.

Option C is not suitable because AWS Glue ETL jobs on a recurring schedule are batch-oriented and cannot transform records in real time. Option E is not suitable because SageMaker Processing jobs are batch jobs that read from storage (typically S3) and write results to S3, not a low-latency streaming transformation path from Kinesis.

Exam trap

MLA-C01 often tests real-time vs batch confusion, so the trap is selecting scheduled batch services (Glue, batch transform, Processing jobs) because they sound like 'data preparation' without regard to latency.

490
MCQeasy

A company wants to deploy a trained XGBoost model for batch inference on a large dataset stored in S3. The inference job should be cost-effective and does not require real-time responses. Which SageMaker inference option should they use?

A.SageMaker Batch Transform
B.SageMaker real-time endpoint
C.SageMaker Asynchronous Inference
D.SageMaker Serverless Inference
AnswerA

SageMaker Batch Transform runs inference over large S3 datasets on managed instances that terminate when the job completes, avoiding the cost of a persistently running endpoint. This satisfies the stem's cost-effectiveness requirement and its lack of any real-time response need.

Why this answer

SageMaker Batch Transform is designed for offline, high-throughput inference on large datasets stored in S3, with no persistent endpoint and no real-time requirement. It is the most cost-effective option for this scenario because you pay only for the duration of the batch job.

Exam trap

MLA-C01 often tests the cost/latency trade-off between Batch Transform and Asynchronous Inference — candidates pick Asynchronous because it sounds 'batch-like,' but Asynchronous Inference is for near-real-time large-payload requests, not bulk offline scoring.

How to eliminate wrong answers

Option B is wrong because real-time endpoints are always-on and billed continuously, making them cost-inefficient for non-real-time batch workloads. Option C is wrong because Asynchronous Inference is designed for large payloads or long processing times with near-real-time queuing, not for bulk offline scoring of an entire S3 dataset. Option D is wrong because Serverless Inference is for intermittent, unpredictable real-time requests with cold-start tolerance, not for large-scale batch jobs.

491
MCQeasy

A machine learning engineer needs to deploy a new version of a model gradually, initially sending 5% of traffic to the new version and 95% to the current version, while monitoring for errors. Which deployment pattern should they use?

A.Blue/green deployment
B.Canary deployment
C.Shadow testing
D.Rolling deployment
AnswerB

A canary deployment shifts a small percentage of traffic, here 5%, to the new model version while the remainder stays on the current version, allowing error monitoring before full rollout. This matches the gradual, low-risk traffic-splitting requirement in the stem.

Why this answer

Canary deployment is the correct pattern because it allows the ML engineer to route a small percentage of traffic (e.g., 5%) to the new model version while keeping the majority (95%) on the current version. This enables gradual rollout with real-time monitoring for errors, and if issues are detected, traffic can be instantly shifted back to the stable version.

Exam trap

Candidates often confuse canary deployment with blue/green deployment, thinking both involve gradual traffic shifting. However, blue/green is an all-or-nothing switch between environments, while canary allows incremental percentage-based routing (e.g., 5%) with real-time monitoring and automated rollback.

How to eliminate wrong answers

Option A is wrong because blue/green deployment involves switching all traffic at once from the current environment (blue) to the new environment (green), which does not support gradual traffic shifting or incremental error monitoring. Option C is wrong because shadow testing sends a copy of live traffic to the new model without affecting user-facing responses, but it does not route actual user traffic to the new version, so it cannot be used to gradually shift real traffic percentages. Option D is wrong because rolling deployment updates instances incrementally (e.g., replacing pods one by one), but it does not provide fine-grained traffic splitting like 5% vs 95% and lacks the instant rollback capability of canary deployments.

492
MCQmedium

A media company uses SageMaker endpoints to serve a model that predicts video engagement. They have two production variants: Variant A (ml.c5.large) for regular traffic and Variant B (ml.c5.xlarge) for burst traffic. They use weighted routing (90% to A, 10% to B). Recently, during peak hours, Variant A's latency increase causes many requests to time out. The metrics show that both variants are under similar CPU load, but the number of concurrent requests to Variant A is very high. The team wants to ensure that burst traffic is handled properly without manual intervention. What should they do?

A.Increase the traffic weight to Variant B to 70% and reduce Variant A to 30%.
B.Configure Application Auto Scaling for each variant with a target tracking scaling policy based on the number of concurrent requests per instance.
C.Set a CloudWatch alarm on Variant A's p99 latency and trigger a step scaling policy to add instances.
D.Create a separate endpoint for burst traffic and route peak traffic to it via DNS.
AnswerB

Concurrent-request target tracking scales each variant independently on the actual bottleneck, since CPU load is similar but Variant A's request concurrency is high. This removes manual intervention and lets Variant B absorb burst traffic as routing shifts.

Why this answer

Application Auto Scaling with a target tracking policy based on concurrent requests per instance automatically adds instances when concurrency rises, directly addressing the high concurrent request load on Variant A. This is the managed, no-manual-intervention solution that scales each variant independently based on the actual bottleneck metric.

Exam trap

MLA-C01 often tests the difference between reactive alarm-based scaling and proactive target tracking — candidates pick CloudWatch alarms because they sound more controlled, missing that target tracking is the managed, automatic solution for concurrency-based scaling.

How to eliminate wrong answers

Option A is wrong because shifting traffic weights to Variant B is a manual, static change that does not scale automatically and may overload Variant B during peak hours. Option C is wrong because a CloudWatch alarm with step scaling is reactive and requires manual configuration of thresholds and steps, and it does not scale based on the actual concurrency metric as cleanly as target tracking. Option D is wrong because creating a separate endpoint and routing via DNS adds complexity, does not automatically scale, and DNS-based routing is not ideal for real-time burst handling.

493
Multi-Selecthard

A retail company runs a SageMaker real-time endpoint serving a demand forecasting model. Security policy requires that all inference requests travel over the AWS private network and never traverse the public internet, and that the endpoint cannot be invoked from outside the company VPC. The endpoint already uses a customer-managed KMS key for volume encryption. Which TWO configurations should the engineer apply to meet these requirements? (Choose two.)

Select 2 answers
A.Configure the endpoint with EnableNetworkIsolation set to true so the container cannot make outbound network calls.
B.Enable request and response data capture on the endpoint and store the captures in an S3 bucket with a bucket policy that allows only the VPC endpoint.
C.Set the endpoint's network access type to VPC-only by attaching the appropriate VPC configuration so the endpoint is reachable only from the specified subnets and security groups.
D.Attach an IAM resource policy to the endpoint that denies all principals except the account root, and enable AWS CloudTrail data events on the endpoint.
E.Create an interface VPC endpoint (AWS PrivateLink) for the SageMaker runtime in the VPC and invoke the endpoint through that endpoint.
AnswersC, E

Configuring the endpoint with a VPC configuration and restricting network access to VPC-only ensures the endpoint is accessible only through the specified subnets and security groups inside the VPC. This removes public invocation paths. Together with a PrivateLink interface endpoint for the runtime API, all inference traffic stays private and external invocation is blocked, meeting the stated policy.

Why this answer

Keeping inference traffic on the AWS private network and blocking external invocation requires two complementary controls. An interface VPC endpoint for the SageMaker runtime lets InvokeEndpoint calls resolve to a private IP inside the VPC, and configuring the endpoint for VPC-only network access restricts reachability to the specified subnets and security groups. IAM policies, network isolation, and data capture do not change the request path or prevent public invocation.

Exam trap

The trap here is treating container network isolation or an IAM resource policy as sufficient for private-only invocation, when the request path and endpoint reachability are governed by PrivateLink and the endpoint's VPC network access configuration.

494
MCQhard

A company deploys a SageMaker model using AWS KMS for encryption at rest. They have a compliance requirement to rotate the KMS key every year without causing downtime for the inference endpoint. Which approach should they take?

A.Use AWS Certificate Manager (ACM) for encryption
B.Create a new KMS key and update the endpoint configuration
C.Manually rotate the key by recreating the endpoint
D.Enable automatic key rotation on the existing KMS key
AnswerD

Automatic key rotation generates new backing key material under the same KMS key ID, so existing ciphertext and the endpoint's references remain valid. This satisfies the annual rotation requirement without re-encrypting data or redeploying, avoiding inference downtime.

Why this answer

AWS KMS supports automatic key rotation, which creates new backing keys annually while retaining the same key ID and metadata. This ensures that the SageMaker endpoint continues to use the same KMS key alias and configuration, so no endpoint update or downtime is required. Automatic rotation satisfies the compliance requirement without any manual intervention or endpoint recreation.

Exam trap

The trap here is that candidates may think rotating a KMS key requires creating a new key and updating the resource (Option B), or that manual recreation is necessary (Option C), when in fact AWS KMS automatic key rotation handles the rotation seamlessly without any endpoint modification or downtime.

How to eliminate wrong answers

Option A is wrong because AWS Certificate Manager (ACM) is for managing SSL/TLS certificates, not for encryption at rest of SageMaker model data; it does not provide KMS key rotation capabilities. Option B is wrong because creating a new KMS key and updating the endpoint configuration would require a deployment update, which can cause a brief interruption or require a rolling update, and it does not leverage the simpler automatic rotation mechanism. Option C is wrong because manually rotating the key by recreating the endpoint would cause downtime during the recreation process, violating the no-downtime requirement.

495
Multi-Selecteasy

Which TWO data storage options are commonly used by Amazon SageMaker Feature Store for offline and online storage?

Select 2 answers
A.Amazon Redshift
B.Amazon RDS
C.Amazon ElastiCache
D.Amazon S3
E.Amazon DynamoDB
AnswersD, E

Amazon S3 serves as the offline store, holding historical feature data in Parquet for training and batch retrieval, while the online store provides low-latency serving. S3 satisfies the offline storage requirement of SageMaker Feature Store.

Why this answer

Amazon S3 (option D) is correct because SageMaker Feature Store uses it as the underlying offline store, where feature data is written in Parquet format for historical/batch retrieval and training. Amazon DynamoDB (option E) is correct because it serves as the online store, providing low-latency, high-throughput reads of the latest feature values for real-time inference. Amazon Redshift (A), Amazon RDS (B), and Amazon ElastiCache (C) are not the managed storage backends used by Feature Store for its online and offline stores, even though they are valid AWS data services in other contexts.

Exam trap

The trap here is that candidates often confuse Amazon ElastiCache (a caching layer) with the primary online storage service, or assume Amazon Redshift is used for offline storage due to its analytical capabilities, but SageMaker Feature Store specifically integrates DynamoDB for online and S3 for offline storage as first-class options.

496
MCQhard

A data scientist is trying to create a SageMaker endpoint configuration with 6 instances of ml.c5.large for a production variant. The creation fails with the error shown in the exhibit. Which action should the data scientist take to resolve this issue?

A.Create two separate endpoint configurations, each with 3 instances, and distribute traffic between them.
B.Request a service quota increase for ml.c5.large for real-time endpoints from the AWS Service Quotas console.
C.Use a different instance type, such as ml.m5.large, which has a higher limit.
D.Delete unused endpoints to free up resources.
AnswerB

The endpoint creation fails because the account's service quota for ml.c5.large instances on real-time endpoints is below the six requested. Raising that quota through the AWS Service Quotas console permits the production variant to allocate all six instances.

Why this answer

The error indicates that the requested number of instances exceeds the service quota for ml.c5.large for real-time endpoints. AWS enforces default limits on instance counts per instance type per region. Requesting a quota increase via the Service Quotas console is the correct action to raise the limit and allow the deployment of 6 instances.

Exam trap

The trap here is that candidates may confuse service quotas with resource availability, thinking that deleting unused endpoints or splitting configurations will free up capacity, when in fact the quota is a hard limit that must be explicitly increased.

How to eliminate wrong answers

Option A is wrong because creating two separate endpoint configurations does not bypass the service quota; the total instance count across all endpoints still counts against the same quota. Option C is wrong because using a different instance type like ml.m5.large does not inherently have a higher limit; each instance type has its own default quota, and the limit for ml.m5.large may also be insufficient or unknown without checking. Option D is wrong because deleting unused endpoints does not increase the quota for ml.c5.large; it only frees up currently used instances, but the quota itself remains unchanged.

497
Multi-Selecthard

A data scientist is training a model with SageMaker and needs to reduce the cost of a long-running training job that can tolerate interruptions. The job uses a custom training script and reads data from Amazon S3. The data scientist wants the job to resume from the last saved state if the underlying compute is reclaimed. (Choose two.)

Select 2 answers
A.Set the estimator's enable_spot_training parameter to true and rely on SageMaker to automatically restart the job from the beginning after an interruption.
B.Increase the volume_size parameter so that the container's local disk can hold the entire dataset and all checkpoints during training.
C.Configure the estimator to use managed spot training and set max_wait time greater than max_run time.
D.Enable network isolation on the estimator to prevent the training container from accessing the internet during spot interruptions.
E.Set the checkpoint_s3_uri parameter on the estimator to an Amazon S3 location where the training script saves checkpoints.
AnswersC, E

Managed spot training uses Amazon EC2 Spot Instances, which are cheaper but can be interrupted. Setting max_wait greater than max_run allows the job to wait for capacity and to resume after an interruption within the overall wait window. This directly addresses cost reduction while tolerating interruptions, and it is a required configuration for the resume behavior to be useful.

Why this answer

Managed spot training lowers cost by using Spot Instances, and checkpointing to Amazon S3 lets a restarted job resume from the last saved state. Setting max_wait above max_run gives the job time to wait for capacity and to recover from interruptions. Network isolation, larger volumes, and restarting from scratch do not support the resume requirement and can increase cost or break access.

Exam trap

The trap here is assuming spot training alone preserves progress, when checkpointing to Amazon S3 is what enables resuming after an interruption.

498
MCQmedium

A machine learning engineer observes that model performance on a SageMaker endpoint has degraded over the past week. Ground truth labels are available with a 2-day delay. The engineer wants to automatically trigger a retraining pipeline when prediction quality drops below an acceptable threshold. Which approach is most appropriate?

A.Use SageMaker Model Monitor - Model Quality Monitor with ground truth, create a CloudWatch alarm on the metric, and trigger an AWS Lambda function to start retraining
B.Manually evaluate the model weekly and retrain as needed
C.Use SageMaker Model Monitor - Data Quality Monitor to detect drift, then trigger retraining
D.Use SageMaker Clarify to monitor bias drift and trigger retraining
AnswerA

Model Quality Monitor ingests the delayed ground truth from S3, computes quality metrics, and publishes them to CloudWatch. An alarm on the threshold invokes Lambda, which starts the retraining pipeline, fully automating detection and remediation without manual intervention.

Why this answer

SageMaker Model Monitor's Model Quality Monitor is specifically designed to compare model predictions against ground truth labels (available with a 2-day delay) and track metrics like accuracy, precision, recall, or F1 score. You can configure a CloudWatch alarm on a metric such as 'accuracy' dropping below a threshold, which triggers an AWS Lambda function to start the retraining pipeline. This automates the detection of prediction quality degradation and the retraining response without manual intervention.

Exam trap

The trap here is that candidates confuse Data Quality Monitor (which monitors input data drift) with Model Quality Monitor (which monitors prediction accuracy against ground truth), leading them to choose Option C incorrectly.

How to eliminate wrong answers

Option B is wrong because manually evaluating the model weekly is not automated and does not meet the requirement to automatically trigger retraining when prediction quality drops; it introduces latency and human error. Option C is wrong because Data Quality Monitor detects drift in input data distribution (e.g., feature skew), not in prediction quality against ground truth labels, so it cannot directly measure model performance degradation. Option D is wrong because SageMaker Clarify is used for bias detection and explainability, not for monitoring prediction quality or triggering retraining based on performance metrics.

499
MCQeasy

A team wants to monitor the number of requests and latency of their SageMaker endpoint using a unified dashboard. Which AWS service should they use to create a custom dashboard with these metrics?

A.Amazon CloudWatch Dashboards
B.AWS CloudTrail
C.AWS Config
D.SageMaker Studio
AnswerA

CloudWatch Dashboards natively aggregate endpoint metrics such as Invocations and ModelLatency into a single custom view, satisfying the unified dashboard requirement without extra tooling. SageMaker automatically publishes these metrics to CloudWatch, so no custom instrumentation is needed.

Why this answer

Amazon CloudWatch Dashboards is the native AWS service for building custom, unified dashboards from CloudWatch metrics. SageMaker automatically publishes endpoint metrics such as Invocations, ModelLatency, OverheadLatency, and Invocation4XXErrors to CloudWatch, so a CloudWatch dashboard can display requests and latency side by side.

Exam trap

MLA-C01 often tests the confusion between CloudWatch (metrics/monitoring) and CloudTrail (API auditing) — candidates pick CloudTrail when the question asks for a metrics dashboard.

How to eliminate wrong answers

Option B is wrong because AWS CloudTrail records API activity (who called what, when) for auditing — it does not store or visualize performance metrics like latency or request counts. Option C is wrong because AWS Config tracks resource configuration changes and compliance, not runtime performance metrics. Option D is wrong because SageMaker Studio is an IDE/notebook environment for building and training models; it does not provide a metrics dashboard service.

500
MCQhard

A team uses SageMaker real-time endpoints for inference. They want to deploy a new model version and compare its performance with the current version under live traffic without affecting user experience. Which method should they use?

A.A/B testing with production variant traffic splitting
B.Batch transform on a holdout test set
C.Blue/green deployment
D.Shadow testing with SageMaker
AnswerD

Shadow testing deploys the new model variant alongside the production variant, mirroring live traffic to it without returning its responses to users. This satisfies the stem's requirement to compare performance under live traffic while leaving user experience unaffected.

Why this answer

Shadow testing with SageMaker deploys the new model alongside the production model and mirrors live traffic to it without returning its responses to users. This lets the team compare performance under real traffic with zero impact on user experience.

Exam trap

MLA-C01 often tests the difference between shadow testing (no user impact, mirrored traffic) and A/B testing (real user impact, split traffic) — candidates pick A/B testing because both compare models under live traffic, but only shadow testing guarantees zero user exposure.

How to eliminate wrong answers

Option A is wrong because A/B testing with production variant traffic splitting sends a portion of real user traffic to the new model and returns its responses, which does affect user experience. Option B is wrong because Batch Transform on a holdout set evaluates offline on historical data, not live traffic, so it does not reflect real production conditions. Option C is wrong because blue/green deployment shifts live traffic to the new model, which directly affects users and is not a silent comparison.

501
MCQhard

A company is deploying a large language model (LLM) to a SageMaker endpoint. They want to minimize inference latency and cost by using GPU acceleration and model parallelism. The model is too large to fit on a single GPU. Which SageMaker feature should they use?

A.SageMaker Elastic Inference
B.SageMaker distributed data parallel library
C.SageMaker model parallelism library
D.SageMaker Inference Recommender
AnswerC

The SageMaker model parallelism library enables training and inference of large models that cannot fit on a single GPU by partitioning the model across multiple GPUs. It supports tensor parallelism and pipeline parallelism, and it can be used for inference to reduce latency and cost. This directly addresses the need for model parallelism with GPU acceleration.

Why this answer

The SageMaker model parallelism library is designed to partition large models across multiple GPUs, enabling inference for models that exceed a single GPU's memory. It supports various parallelism strategies and can reduce latency and cost. The other options are either for training, deprecated, or only provide recommendations, not the required model parallelism.

Exam trap

The trap here is confusing the distributed data parallel library with model parallelism; data parallel is for training and does not split the model.

502
MCQeasy

A data engineer stores raw ML training data in Amazon S3 and needs to catalog the schema, track partition changes, and make the data queryable by Amazon Athena without running ETL. Which AWS service should the engineer use?

A.Amazon SageMaker Feature Store with an offline store backed by S3.
B.AWS Glue ETL job with a Python shell script.
C.AWS Glue Data Catalog with an AWS Glue crawler.
D.Amazon Redshift Spectrum with an external schema.
AnswerC

A Glue crawler inspects data in S3, infers schema and partitions, and registers tables in the Glue Data Catalog, which Athena uses as its metastore for querying S3 data in place. This satisfies cataloging, partition tracking, and query access without moving or transforming the data, matching the stated requirement exactly.

Why this answer

The Glue Data Catalog is the central metadata repository that stores table definitions, schemas, and partition information for data in S3, and Athena natively uses it as its metastore. A crawler automates schema inference and partition discovery, so no ETL is required. The other services either transform data, serve features, or consume the catalog rather than create it.

Exam trap

The trap here is confusing a service that consumes the Glue Data Catalog, such as Athena or Redshift Spectrum, with the service that actually builds and maintains the catalog entries.

503
MCQeasy

A data engineer needs to convert a JSON dataset to Parquet format for efficient querying with Amazon Athena. The JSON files are in an S3 bucket. Which service can perform this conversion with minimal coding?

A.Amazon SageMaker Processing
B.Amazon EMR
C.AWS Lambda
D.AWS Glue Studio with a visual job
AnswerD

AWS Glue Studio's visual job editor generates the PySpark ETL code that reads JSON from S3, applies a schema, and writes Parquet back to S3, satisfying the minimal-coding constraint. Its built-in transforms and crawler-derived schemas remove hand-written conversion logic, and the output is directly queryable by Athena.

Why this answer

AWS Glue Studio with a visual job is the correct choice because it provides a no-code, drag-and-drop interface to create ETL jobs that can read JSON from S3 and write it as Parquet, with built-in schema inference and transformation capabilities. This minimizes coding effort while leveraging Glue's serverless Spark engine for efficient conversion, making it ideal for preparing data for Athena queries.

Exam trap

The trap here is that candidates often confuse AWS Glue Studio with AWS Glue DataBrew or assume that any AWS service with 'processing' in its name (like SageMaker Processing) is suitable for simple ETL tasks, overlooking the specific no-code visual job capability of Glue Studio.

How to eliminate wrong answers

Option A is wrong because Amazon SageMaker Processing is designed for data preprocessing and model training workflows within the ML pipeline, not for simple file format conversion; it requires writing custom processing scripts and managing infrastructure, which adds unnecessary complexity. Option B is wrong because Amazon EMR is a managed Hadoop/Spark cluster that can perform the conversion, but it requires provisioning and configuring a cluster, writing Spark or Hive code, and managing lifecycle, which is far more coding and operational overhead than a visual job. Option C is wrong because AWS Lambda has a maximum execution time of 15 minutes and a deployment package size limit, making it impractical for converting large JSON datasets to Parquet; it also requires custom Python code with libraries like PyArrow or Pandas, which is not minimal coding.

504
MCQmedium

A company wants to automatically trigger a retraining pipeline when concept drift is detected in their deployed model. Which combination of services should they use?

A.SageMaker Model Monitor → Lambda
B.CloudWatch Events → SageMaker Training Job
C.SageMaker Model Monitor → CloudWatch Alarm → SNS → Lambda
D.SageMaker Clarify → SNS → Step Functions
AnswerC

SageMaker Model Monitor detects drift and emits metrics; CloudWatch alarms on those thresholds, SNS fans out the notification, and Lambda invokes the retraining pipeline. This chain satisfies the requirement to trigger retraining automatically upon concept drift detection without manual intervention.

Why this answer

SageMaker Model Monitor detects concept drift by analyzing model predictions against a baseline, then publishes metrics to CloudWatch. A CloudWatch Alarm triggers when drift exceeds a threshold, sending a notification via SNS to invoke a Lambda function, which starts the retraining pipeline. This end-to-end integration ensures automated, event-driven retraining without manual intervention.

Exam trap

A common exam trap is the distinction between monitoring services (Model Monitor for drift vs. Clarify for bias) and the correct event chain (Model Monitor → CloudWatch → SNS → Lambda) versus incomplete chains like direct Lambda invocation or using the wrong service for drift detection.

How to eliminate wrong answers

Option A is wrong because SageMaker Model Monitor alone cannot directly invoke Lambda; it requires CloudWatch Alarms and SNS to bridge the monitoring output to Lambda execution. Option B is wrong because CloudWatch Events (now EventBridge) can trigger SageMaker Training Jobs, but it lacks the concept drift detection capability provided by Model Monitor, so it cannot determine when retraining is needed. Option D is wrong because SageMaker Clarify is designed for bias detection and explainability, not concept drift monitoring; using SNS and Step Functions without drift detection would not trigger retraining based on model performance degradation.

505
MCQmedium

A company uses SageMaker endpoints with auto-scaling. The endpoint is experiencing high latency during peak hours. The metrics show CPU utilization is low but memory is high. What is the most likely cause?

A.The model is not optimized for inference, causing memory leaks.
B.The auto-scaling policy is based on CPU utilization, which does not trigger scaling.
C.The instance type has insufficient network bandwidth.
D.The endpoint is deployed in a VPC without a NAT gateway.
AnswerB

Memory-bound inference keeps CPU low, so a target-tracking policy on CPU utilisation never breaches its threshold and no additional instances are provisioned. Scaling must instead track a memory metric such as `MemoryUtilization`, or invoke-based concurrency, to satisfy the peak-hour latency constraint.

Why this answer

The auto-scaling policy is based on CPU utilization, which remains low during the memory-bound issue. Since the scaling trigger is not met, the endpoint does not add more instances to handle the increased load, leading to high latency. Memory pressure without CPU spikes indicates the bottleneck is memory, not compute, so a CPU-based metric fails to scale appropriately.

Exam trap

The trap here is that candidates assume high latency always means CPU is the bottleneck, but the exam tests understanding that auto-scaling must be based on the correct metric; memory pressure can cause latency without CPU spikes, and a CPU-based policy will fail to scale.

How to eliminate wrong answers

Option A is wrong because a memory leak would cause memory to increase over time, not specifically during peak hours, and would likely degrade performance gradually rather than cause latency spikes tied to load. Option C is wrong because insufficient network bandwidth would manifest as network-related errors or timeouts, not high memory utilization with low CPU; network metrics would show saturation. Option D is wrong because a VPC without a NAT gateway affects outbound internet access, not inbound inference requests to the endpoint; SageMaker endpoints in a VPC can receive traffic via VPC endpoints or public endpoints without a NAT gateway.

506
Multi-Selecteasy

Which TWO actions are recommended best practices when preparing training data for a machine learning model in AWS? (Choose two.)

Select 2 answers
A.Remove all outliers from the dataset.
B.Train the model on the entire dataset to maximize data usage.
C.Check for and handle missing values appropriately.
D.Split the data into training, validation, and test sets.
E.Always normalize all features to a [0,1] range.
AnswersC, D

Missing values bias model training and can cause failures in algorithms that reject nulls. Detecting and handling them—via imputation, removal, or indicator flags—preserves data integrity, satisfying the best-practise requirement for robust training data preparation in AWS pipelines such as SageMaker Data Wrangler or Glue.

Why this answer

Option C is correct because missing values can bias or break training algorithms, so best practice is to detect them (e.g., with pandas isnull() or Amazon SageMaker Data Wrangler) and handle them via imputation, removal, or model-native handling. Option D is correct because splitting data into training, validation, and test sets lets you fit parameters, tune hyperparameters, and estimate generalization performance on unseen data, avoiding overfitting and data leakage. Option A is not recommended because not all outliers are errors; blindly removing them can discard legitimate signal and distort the distribution.

Option B is wrong because training on the entire dataset leaves no held-out data for validation or unbiased testing. Option E is wrong because normalization is not always required and [0,1] scaling is only one of several techniques (e.g., standardization, log transforms) chosen based on the algorithm and feature distribution.

Exam trap

The trap here is that candidates assume all outliers must be removed (Option A) or that normalization is always required (Option E), but the exam tests nuanced understanding that these steps depend on the algorithm and data characteristics, not blanket rules.

507
MCQmedium

A team is using SageMaker Pipelines to automate retraining and deployment. They want to trigger the pipeline automatically when new training data is available in an S3 bucket. Which approach should they use?

A.Create an Amazon EventBridge rule that triggers the pipeline execution on S3 PutObject events
B.Register the pipeline as a model package in SageMaker Model Registry
C.Configure a cron job to run the pipeline every hour
D.Use AWS Step Functions to poll the S3 bucket and start the pipeline when a new object appears
AnswerA

EventBridge natively consumes S3 PutObject events and can invoke SageMaker Pipelines directly as a target, satisfying the automatic trigger requirement without polling or custom glue code. The event-driven rule fires the pipeline execution the moment new training data lands in the bucket.

Why this answer

Amazon EventBridge can directly capture S3 PutObject events and invoke a SageMaker Pipeline execution as a target. This provides a fully event-driven, serverless integration without polling or manual intervention, aligning with best practices for automating ML workflows when new data arrives.

Exam trap

The trap here is that candidates may overcomplicate the solution by choosing Step Functions (Option D) for orchestration, not realizing that EventBridge provides a simpler, event-driven trigger without the need for polling or additional state machines.

How to eliminate wrong answers

Option B is wrong because registering a pipeline as a model package in SageMaker Model Registry is for versioning and managing trained models, not for triggering pipeline executions based on S3 events. Option C is wrong because a cron job runs on a fixed schedule, which is inefficient and may miss data arrivals or run unnecessarily, whereas the requirement is to trigger only when new data appears. Option D is wrong because using AWS Step Functions to poll S3 introduces latency, cost, and complexity compared to the native event-driven approach with EventBridge, which reacts instantly to S3 events.

508
MCQhard

A machine learning team runs a SageMaker AI Pipeline that trains a model and registers it in the SageMaker AI Model Registry. A separate deployment process must promote the model to production only after a human reviewer approves the model version. The team wants to automate the promotion so that approval in the Model Registry triggers the deployment without manual intervention. Which combination of steps should the engineer implement?

A.Use a SageMaker AI Projects template that automatically deploys every model version as soon as it is registered.
B.Schedule an AWS Lambda function every minute to call DescribeModelPackage and deploy any approved model version.
C.Add a ConditionStep to the existing pipeline that checks the model version status and deploys the model.
D.Configure an Amazon EventBridge rule for the SageMaker AI Model Registry model version state change to Approved, and use it to start a deployment workflow.
AnswerD

The SageMaker AI Model Registry emits state change events when a model version moves to Approved. An EventBridge rule can match that event and invoke a target such as AWS Step Functions or AWS CodePipeline to run the deployment. This automates promotion only after the human approval step, satisfying the requirement without polling.

Why this answer

Model approval in the SageMaker AI Model Registry changes the model version state and emits an event. Capturing that event with Amazon EventBridge and routing it to a deployment workflow links the human approval gate to automated promotion, which is exactly the required behavior.

Exam trap

The trap here is trying to enforce a post-pipeline human approval inside the pipeline itself instead of reacting to the Model Registry state change event.

509
MCQeasy

Refer to the exhibit. A user is unable to invoke a SageMaker endpoint. The IAM policy shown is attached to the user. Which permission is missing to allow invocation?

A.sagemaker:InvokeEndpoint
B.sagemaker:DescribeEndpoint
C.sagemaker:CreateEndpoint
D.sagemaker:ListEndpoints
AnswerA

The IAM policy lacks the sagemaker:InvokeEndpoint action, which is the specific permission required to call a SageMaker endpoint's inference API. Without it, every InvokeEndpoint request is denied regardless of any resource-level grants, so adding this action to the policy resolves the user's failure.

Why this answer

To invoke a SageMaker endpoint, the user needs the `sagemaker:InvokeEndpoint` permission. The IAM policy shown lacks this action, which is required for making real-time inference requests to the endpoint. Without it, any attempt to call the endpoint via the SDK or CLI will fail with an access denied error.

Exam trap

AWS often tests the distinction between read-only permissions (like `DescribeEndpoint` or `ListEndpoints`) and the specific action required to perform an operation, leading candidates to confuse metadata access with actual invocation capability.

How to eliminate wrong answers

Option B is wrong because `sagemaker:DescribeEndpoint` only allows retrieving metadata about an endpoint, not invoking it for inference. Option C is wrong because `sagemaker:CreateEndpoint` is for creating new endpoints, not for sending inference requests to an existing one. Option D is wrong because `sagemaker:ListEndpoints` only lists endpoints in the account, which does not grant the ability to invoke them.

510
MCQhard

A company uses Amazon SageMaker Feature Store to store features for a real-time recommendation model. The feature data is updated continuously, and the model must use the most recent feature values for each user at inference time. Which type of Feature Store should the company use for serving features to the model?

A.Offline store
B.Point-in-time queries
C.Online store
D.Both online and offline store
AnswerC

An Amazon SageMaker Feature Store online store serves features with low-latency access for real-time inference, satisfying the requirement that the recommendation model must use the most recent feature values per user at inference time. Unlike the offline store, which is optimised for batch retrieval and historical analysis, the online store provides a synchronised, low-latency key-value lookup for continuously updated feature data.

Why this answer

The SageMaker Feature Store online store is optimized for low-latency, real-time reads of the latest feature values, which is exactly what a real-time recommendation model needs at inference time. It stores the most recent record for each entity (e.g., user) in a low-latency store backed by a fast database. Offline store is for batch/historical access, not real-time serving.

Exam trap

MLA-C01 often tests the online vs. offline store distinction by embedding a real-time serving requirement, luring candidates toward 'both' or 'point-in-time queries' when only the online store satisfies low-latency inference.

How to eliminate wrong answers

Option A is wrong because the offline store is backed by S3 and is designed for batch training and historical analysis, not low-latency real-time inference. Option B is wrong because point-in-time queries are a feature of the offline store used to reconstruct feature values as of a past timestamp for training to avoid data leakage — they are not a serving mechanism. Option D is wrong because while many teams create both stores (online for serving, offline for training), the question asks specifically which store serves features to the model at inference time; only the online store does that.

511
Multi-Selectmedium

A machine learning team needs to monitor a deployed model for both data drift and concept drift. Which TWO approaches should they implement? (Select TWO.)

Select 2 answers
A.Set up SageMaker Model Monitor for data quality monitoring
B.Use SageMaker Clarify for bias monitoring
C.Configure CloudWatch Logs Insights to query inference logs
D.Set up SageMaker Model Monitor for model quality monitoring
E.Enable SageMaker Debugger during inference
AnswersA, D

Data quality monitoring compares live inference inputs against the training baseline, computing per-feature distribution statistics to detect data drift — changes in input feature distributions. This directly satisfies the stem's data drift requirement, while concept drift needs the separate model quality monitor.

Why this answer

Option A is correct because SageMaker Model Monitor's data quality monitoring detects data drift by comparing the statistical properties of incoming inference requests against the baseline statistics captured from the training dataset, alerting when feature distributions shift. Option D is correct because Model Monitor's model quality monitoring evaluates concept drift by comparing predicted values against actual ground-truth labels, tracking metrics such as accuracy, precision, and recall over time to detect degradation in the model's real-world performance. Option B is not correct because SageMaker Clarify is used for bias detection and explainability (e.g., SHAP values), not for detecting data or concept drift.

Option C is not correct because CloudWatch Logs Insights merely queries and analyzes log data; it does not provide built-in drift detection statistics or baseline comparisons. Option E is not correct because SageMaker Debugger is designed to debug training jobs by capturing tensors and monitoring training metrics, not to monitor deployed models for drift.

Exam trap

The trap is assuming a single monitoring approach covers both drift types; candidates often pick Clarify or Debugger because they sound like monitoring tools, but only Model Monitor's data quality and model quality modes address data and concept drift respectively.

512
MCQmedium

A machine learning engineer is deploying a model to a SageMaker endpoint that must handle occasional large payloads up to 1 GB. The inference time can take up to 10 minutes. The team wants to minimize cost and avoid idle compute. Which deployment option is most appropriate?

A.SageMaker real-time endpoint with automatic scaling
B.SageMaker Serverless Inference
C.SageMaker Asynchronous Inference
D.SageMaker batch transform
AnswerC

Asynchronous Inference supports payloads up to 1 GB and allows inference times up to 15 minutes. It queues requests and processes them asynchronously, returning results via Amazon S3. This makes it ideal for large payloads and long-running inference, and it can scale to zero when idle, minimizing cost.

Why this answer

Asynchronous Inference is designed for large payloads (up to 1 GB) and long inference times (up to 15 minutes). It queues requests and returns results via S3, and it can scale to zero when idle, reducing cost. Real-time and serverless inference have payload and timeout limits, and batch transform is for offline batch processing, not on-demand requests.

Exam trap

The trap here is assuming that real-time endpoints can handle any payload size, but they are limited to 6 MB and are cost-inefficient for long-running jobs.

513
MCQeasy

An ML engineer runs the CLI command shown in the exhibit. However, the training job fails immediately with an error: 'Unable to assume role'. What is the most likely cause?

A.The IAM role 'SageMakerExecutionRole' does not have permission to create the training job.
B.The training image in ECR does not exist.
C.The S3 bucket 'my-bucket' does not exist.
D.The IAM role's trust policy does not grant SageMaker permission to assume the role.
AnswerD

SageMaker assumes the execution role via AWS Security Token Service, so the role's trust policy must list sagemaker.amazonaws.com as a principal. Without that sts:AssumeRole grant, the training job cannot obtain temporary credentials and fails immediately with 'Unable to assume role'.

Why this answer

The 'Unable to assume role' error indicates that SageMaker cannot assume the IAM role specified in the CLI command. This is a trust policy issue: the role's trust policy must include SageMaker as a trusted service (i.e., `"Service": "sagemaker.amazonaws.com"`). Without this, SageMaker is not authorized to assume the role, regardless of the role's permissions.

Exam trap

AWS often tests the distinction between IAM role permissions (what the role can do) and trust policies (who can assume the role), leading candidates to mistakenly select a permission-related option when the error is about trust.

How to eliminate wrong answers

Option A is wrong because the error is about assuming the role, not about the role's permissions to create the training job; permission errors would appear as 'AccessDenied' or similar, not 'Unable to assume role'. Option B is wrong because a missing ECR image would cause an error like 'Image not found' or 'RepositoryNotFoundException', not an assume role error. Option C is wrong because a non-existent S3 bucket would result in an error like 'NoSuchBucket' or 'AccessDenied' when SageMaker tries to access it, not an assume role failure.

514
Multi-Selecteasy

A company is using Amazon SageMaker to deploy a model for real-time inference. The model requires access to a private S3 bucket that contains reference data. The company wants to ensure that the endpoint can access the S3 bucket without using a public internet connection. Which TWO actions should they take? (Select TWO.)

Select 2 answers
A.Configure the endpoint's security group to allow outbound traffic to the S3 bucket's IP range.
B.Attach the endpoint to a VPC that has a VPC endpoint for S3.
C.Ensure the SageMaker execution role has an IAM policy that grants s3:GetObject access to the bucket.
D.Attach the endpoint to a VPC with an internet gateway and route the S3 traffic through the internet gateway.
E.Attach the endpoint to a VPC with a NAT gateway to route traffic to S3.
AnswersB, C

A VPC endpoint for S3 (gateway endpoint) routes traffic from the endpoint's subnets to S3 over the AWS private network, keeping requests off the public internet. This satisfies the stem's constraint of private bucket access without public connectivity, since the endpoint must reside in a VPC to use it.

Why this answer

Option B is correct because attaching the SageMaker endpoint to a VPC that has an S3 gateway VPC endpoint (or interface endpoint) allows traffic to reach the private S3 bucket through the AWS private network instead of the public internet, satisfying the no-public-internet requirement. Option C is correct because the SageMaker execution role must have an IAM policy granting s3:GetObject on the bucket, otherwise the endpoint cannot read the reference data regardless of network path. Option A is not correct because security groups cannot whitelist an S3 bucket's IP range reliably (S3 uses dynamic AWS IP ranges and gateway endpoints don't use those IPs), and network access alone doesn't grant authorization.

Option D is not correct because routing S3 traffic through an internet gateway uses the public internet, violating the requirement. Option E is not correct because a NAT gateway also routes traffic over the public internet, so it does not meet the private-access requirement.

Exam trap

The trap here is that candidates often confuse VPC endpoints (which keep traffic private) with NAT gateways or internet gateways (which route traffic over the public internet), and they may overlook the mandatory IAM permissions required for S3 access even when using a VPC endpoint.

515
MCQmedium

A data engineer is using Amazon SageMaker Ground Truth to create a labeled dataset for an object detection task. The dataset contains millions of images, and the labeling budget is limited. Which approach can reduce labeling costs while maintaining high model accuracy?

A.Enable active learning in Ground Truth to automatically select a subset of images for human labeling
B.Label 100% of the images using a pre-built worker template to ensure accuracy
C.Use Amazon SageMaker Data Wrangler to annotate images
D.Use automated labeling with a pre-trained model for all images and skip human review
AnswerA

Active learning uses model uncertainty sampling to route only the most informative images to human labellers, leaving confident predictions auto-labelled. This cuts the volume of paid human annotations across millions of images, satisfying the stem's limited-budget constraint while preserving accuracy on the object detection task.

Why this answer

Ground Truth active learning (automated data labeling) uses a machine learning model to identify the most informative unlabeled images — those where the model is least confident — and sends only those to human labelers. This reduces the number of human annotations required while maintaining model accuracy, directly addressing the limited labeling budget for millions of images. The other options either label everything (expensive), use the wrong tool, or skip human review entirely (risking accuracy).

Exam trap

MLA-C01 often tests the difference between active learning (selective human labeling to reduce cost) and automated labeling without review (which risks accuracy), causing candidates to choose full automation as a cost-saving measure.

How to eliminate wrong answers

Option B is wrong because labeling 100% of millions of images with human workers maximizes cost, which contradicts the goal of reducing labeling costs. Option C is wrong because SageMaker Data Wrangler is a data preparation and feature engineering tool, not an image annotation service — Ground Truth is the correct service for labeling. Option D is wrong because using automated labeling for all images and skipping human review removes the quality control that ensures high model accuracy; Ground Truth's automated labeling still requires a human-verified test set and typically a review step.

516
MCQmedium

A company wants to deploy a PyTorch model on SageMaker using the NVIDIA Triton Inference Server for GPU acceleration. They have an existing Triton configuration. Which approach should they take?

A.Use SageMaker Neo to compile the model for Triton
B.Package Triton as a custom container and use SageMaker batch transform
C.Use the SageMaker Triton Inference Server container from the Deep Learning Containers
D.Use the standard SageMaker PyTorch container and install Triton at runtime
AnswerC

The SageMaker Triton Deep Learning Container ships NVIDIA Triton Inference Server preinstalled and configured, so the existing Triton model configuration can be deployed directly. This satisfies the GPU acceleration requirement without building or maintaining a custom container image.

Why this answer

AWS provides a pre-built SageMaker Triton Inference Server container as part of the Deep Learning Containers (DLCs), which is optimized for GPU acceleration and supports the existing Triton configuration without modification. This container integrates directly with SageMaker hosting endpoints, enabling seamless deployment of PyTorch models with Triton's features like dynamic batching and model concurrency.

Exam trap

The trap here is that candidates may assume SageMaker Neo is a universal compilation tool for any inference server, but Neo is specifically for hardware-specific optimization and does not support Triton's runtime environment, leading them to incorrectly select Option A.

How to eliminate wrong answers

Option A is wrong because SageMaker Neo compiles models for specific hardware targets (e.g., Intel, ARM) and does not support compilation for the NVIDIA Triton Inference Server; Neo is designed for edge devices and does not integrate with Triton's serving architecture. Option B is wrong because while packaging Triton as a custom container is possible, using SageMaker batch transform is not the recommended approach for real-time inference with GPU acceleration; batch transform is for offline, asynchronous processing, not for low-latency serving. Option D is wrong because installing Triton at runtime on the standard PyTorch container is inefficient and error-prone; it adds startup latency, may cause dependency conflicts, and bypasses the pre-optimized, tested Triton container that AWS provides.

517
MCQmedium

A team uses AWS Auto Scaling for a SageMaker real-time endpoint. They notice that when scaling in, the latest instance is always terminated first, causing disruption to recent requests. How can they configure the scaling policy to terminate the oldest instance first?

A.Configure the termination policy as 'OldestInstance'
B.No action needed; this is the default behavior
C.Use lifecycle hooks
D.Use AWS CloudFormation to manage the endpoint
AnswerA

You can set the termination policy to 'OldestInstance' in the scaling policy configuration.

Why this answer

AWS Auto Scaling for SageMaker endpoints supports a termination policy of 'OldestInstance', which explicitly instructs the scaling process to terminate the instance that has been running the longest. By default, Auto Scaling terminates the newest instance (the default termination policy), which can disrupt recent requests. Configuring the termination policy to 'OldestInstance' ensures that the oldest, most stable instance is removed first, minimizing disruption to in-flight requests.

Exam trap

The trap here is that candidates assume the default termination policy is 'OldestInstance' or that lifecycle hooks can influence instance selection, when in fact the default is 'NewestInstance' and lifecycle hooks only add a delay, not a selection rule.

How to eliminate wrong answers

Option B is wrong because the default behavior of AWS Auto Scaling is to terminate the newest instance first (the 'Default' termination policy), not the oldest, so action is needed to change this. Option C is wrong because lifecycle hooks are used to perform custom actions (e.g., draining connections) before an instance is terminated or launched, but they do not control which instance is selected for termination; they only add a pause in the lifecycle. Option D is wrong because AWS CloudFormation is an infrastructure-as-code service for provisioning resources, not a mechanism to configure the termination policy of an Auto Scaling group; the termination policy must be set directly on the Auto Scaling group or via the SageMaker endpoint configuration.

518
MCQmedium

A team is using Amazon SageMaker to train a neural network. They want to minimize training time while effectively exploring the hyperparameter space. Which approach should they use?

A.Random search
B.Bayesian optimization
C.Grid search
D.Manual tuning
AnswerB

Bayesian optimisation builds a probabilistic surrogate model of the objective and selects each hyperparameter combination based on prior results, converging on strong configurations in fewer training jobs than grid or random search, minimising total training time.

Why this answer

Bayesian optimization is the correct approach because it builds a probabilistic model of the objective function and uses it to select the most promising hyperparameters to evaluate next, balancing exploration and exploitation. This method converges to optimal hyperparameters in fewer iterations than random or grid search, significantly reducing training time for expensive neural network models.

Exam trap

AWS often tests the misconception that random search is always the best for hyperparameter tuning, but the question explicitly asks to minimize training time, which favors Bayesian optimization's efficient use of prior evaluations.

How to eliminate wrong answers

Option A is wrong because random search, while better than grid search for high-dimensional spaces, does not use past evaluation results to inform future trials, leading to more wasted iterations and longer training time. Option C is wrong because grid search exhaustively evaluates all combinations of a predefined set of hyperparameter values, which is computationally prohibitive for neural networks with many hyperparameters and does not scale efficiently. Option D is wrong because manual tuning relies on human intuition and trial-and-error, which is slow, non-reproducible, and cannot systematically explore the hyperparameter space to minimize training time.

519
MCQhard

A company needs to update a model in production without any downtime. They currently have a single real-time endpoint serving traffic. Which approach allows them to deploy a new model version and switch traffic gradually while being able to roll back quickly?

A.Use a canary deployment by creating a new production variant with the new model and shifting traffic incrementally
B.Use a multi-model endpoint and replace the model file
C.Stop the endpoint, update the model, and restart the endpoint
D.Update the existing endpoint's model directly using UpdateEndpoint
AnswerA

A canary deployment adds a new production variant holding the new model, then shifts traffic incrementally via variant weights. This satisfies the stem's no-downtime and gradual-shift requirements, and weights can be reverted instantly to roll back.

Why this answer

A canary deployment creates a new production variant on the existing endpoint with the new model and shifts a small percentage of traffic to it, allowing gradual validation and quick rollback by shifting traffic back to the original variant. This avoids downtime and provides controlled risk.

Exam trap

The trap is thinking 'update the endpoint' is sufficient — candidates must recognize that in-place updates cause downtime and lack gradual traffic control, whereas canary variants provide safe, reversible rollouts.

How to eliminate wrong answers

Option B is wrong because a multi-model endpoint hosts multiple models behind one endpoint but does not provide traffic-shifting or gradual rollout semantics for a new version. Option C is wrong because stopping the endpoint causes downtime, violating the no-downtime requirement. Option D is wrong because updating the existing endpoint's model directly replaces the model in place with no gradual traffic shift and no easy rollback path.

520
MCQeasy

An ML engineer has trained a model and stored the model artifacts in an Amazon S3 bucket in the same AWS Region as the planned SageMaker AI endpoint. During endpoint creation, the engineer must specify the S3 location of the model artifacts. Which permission must the SageMaker AI execution role have for the endpoint to load the model successfully?

A.s3:PutObject on the model artifact prefix in the bucket.
B.s3:DeleteObject on the model artifact prefix in the bucket.
C.s3:GetObject on the model artifact objects in the bucket.
D.s3:ListAllMyBuckets at the account level.
AnswerC

The SageMaker AI execution role is assumed by the hosting infrastructure to download model artifacts from Amazon S3 at container startup. Without s3:GetObject on the specific artifact objects, the container cannot retrieve the model and endpoint creation or invocation fails, so this permission is required.

Why this answer

The SageMaker AI execution role must be able to read the model artifacts from Amazon S3 when the container starts. Granting s3:GetObject on the artifact objects provides exactly that read access and follows least privilege.

Exam trap

The trap here is granting broad Amazon S3 permissions such as listing all buckets instead of the specific read access the endpoint actually needs.

521
MCQeasy

An ML team wants to use Amazon SageMaker Ground Truth to create a labeled dataset for a multi-class image classification task. They have a large set of unlabeled images and want to minimize labeling costs while maintaining high accuracy. Which Ground Truth feature should they enable?

A.Active learning
B.Annotation consolidation
C.Data labeling workforce management
D.Consolidated labeling
AnswerA

Active learning automatically selects the most informative unlabeled images for human labelling, so annotators spend effort only on samples that most improve the model. This directly minimises labelling cost while sustaining high accuracy, satisfying the stem's requirement to cut expense without sacrificing quality on the multi-class image task.

Why this answer

Active learning in SageMaker Ground Truth automatically selects the most informative unlabeled images for human labeling, reducing the total number of labels needed while maintaining model accuracy. By iteratively training a model on a small labeled subset and then using that model to identify uncertain predictions, the system focuses labeling effort on the data that will most improve the model, directly minimizing labeling costs.

Exam trap

The trap here is that candidates may confuse 'annotation consolidation' (a post-labeling quality step) with a cost-reduction feature, or think that workforce management alone reduces costs, when in fact active learning is the specific feature designed to minimize the number of labels required.

How to eliminate wrong answers

Option B (Annotation consolidation) is wrong because it refers to combining multiple annotations for the same data point to produce a ground truth label, which does not reduce the number of labels needed. Option C (Data labeling workforce management) is wrong because it involves managing human labelers (e.g., public, private, or vendor workforces) but does not inherently reduce labeling volume or cost. Option D (Consolidated labeling) is not a distinct SageMaker Ground Truth feature; it is a generic term that might be confused with annotation consolidation, and it does not address cost minimization through selective labeling.

522
MCQmedium

A data scientist deploys a model and wants to monitor the endpoint's invocation latency. They notice that the CloudWatch metric 'ModelLatency' is high, but 'OverheadLatency' is low. Which statement correctly interprets these metrics?

A.The SageMaker overhead is causing the delay; check endpoint configuration
B.The model inference time is the bottleneck; consider optimizing the model or using a faster instance type
C.The endpoint is overloaded; increase the number of instances
D.The network latency is high; move the endpoint closer to clients
AnswerB

ModelLatency measures time spent inside the model container performing inference, while OverheadLatency covers SageMaker platform overhead. High ModelLatency with low OverheadLatency isolates the model itself as the bottleneck, so optimising the model or using a faster instance helps.

Why this answer

The 'ModelLatency' metric measures the time taken by the SageMaker model container to process a single request, including inference and any preprocessing/postprocessing within the container. 'OverheadLatency' measures the time spent on SageMaker infrastructure (e.g., network I/O, request queuing, and response handling). When ModelLatency is high and OverheadLatency is low, the bottleneck is clearly the model inference time itself, not the infrastructure overhead. Therefore, optimizing the model (e.g., quantization, pruning) or upgrading to a faster instance type (e.g., GPU vs.

CPU) is the correct remediation.

Exam trap

The trap here is that candidates confuse 'ModelLatency' with overall endpoint latency and assume any high latency is due to infrastructure or scaling issues, when in fact the metric explicitly isolates the model's own inference time from overhead.

How to eliminate wrong answers

Option A is wrong because high ModelLatency with low OverheadLatency indicates the delay is inside the model container, not in SageMaker's infrastructure overhead; checking endpoint configuration would not address the model's own inference time. Option C is wrong because endpoint overload typically manifests as increased OverheadLatency (due to request queuing) or increased Invocations and 5xx errors, not as isolated high ModelLatency with low OverheadLatency. Option D is wrong because network latency is captured within OverheadLatency, not ModelLatency; moving the endpoint closer to clients would reduce OverheadLatency but would not affect the model's inference computation time.

523
MCQeasy

A data engineer is preparing a large dataset of 10 TB for ML training on Amazon SageMaker. The data is stored in Amazon S3 as CSV files. To reduce training time and cost, the engineer wants to use a columnar format that is optimized for analytical queries. Which format should the engineer convert the data to?

A.XML
B.Parquet
C.ORC
D.JSON Lines
AnswerB

Parquet stores data column-wise with compression and encoding, so SageMaker reads only the columns each training job needs rather than scanning every field. That columnar layout and smaller footprint cut I/O and cost across the 10 TB CSV dataset, unlike row-based formats such as CSV or JSON.

Why this answer

Parquet is a columnar storage format that is highly optimized for analytical queries and is natively supported by Amazon SageMaker for efficient data loading. By converting the 10 TB of CSV data to Parquet, the data engineer can reduce I/O and storage costs because columnar formats allow SageMaker to read only the columns needed for training, rather than scanning entire rows. This directly addresses the goal of reducing training time and cost for ML workloads.

Exam trap

AWS often tests the distinction between columnar formats (Parquet vs. ORC) by making both appear correct, but the trap here is that ORC is tightly coupled with Hive and less commonly used with SageMaker, while Parquet is the de facto standard for AWS-native ML and analytics services.

How to eliminate wrong answers

Option A (XML) is wrong because XML is a verbose, row-oriented text format that is not optimized for analytical queries; it would increase storage size and I/O overhead, making training slower and more expensive. Option C (ORC) is also a columnar format optimized for analytical queries, but it is primarily designed for and tightly integrated with the Apache Hive ecosystem, whereas Parquet is the more universally supported and recommended format for Amazon SageMaker and AWS analytics services. Option D (JSON Lines) is wrong because it is a row-oriented, text-based format that lacks the compression and columnar pruning benefits of Parquet, leading to higher storage costs and slower data access for ML training.

524
MCQeasy

A small team runs a SageMaker real-time endpoint in production. They want a low-effort way to know when the endpoint's invocations are failing so they can react quickly, and they want the alert delivered to their on-call channel. Which approach requires the least custom code?

A.Use AWS CloudTrail to capture InvokeEndpoint API calls and create an Amazon EventBridge rule that notifies the on-call channel when errors occur.
B.Create an Amazon CloudWatch alarm on the endpoint's ModelLatency and Invocation4XXErrors metrics, and route the alarm to Amazon SNS with an email or chat subscription.
C.Enable SageMaker Model Monitor on the endpoint and configure the monitor to raise an alarm when constraint violations are detected.
D.Subscribe an AWS Lambda function to the endpoint's CloudWatch log group and have the function parse logs and publish to Amazon SNS when errors appear.
AnswerB

SageMaker automatically publishes endpoint metrics such as Invocations, Invocation4XXErrors, Invocation5XXErrors, and ModelLatency to CloudWatch without any instrumentation. Alarms on those metrics publish to an SNS topic that the on-call channel subscribes to, giving failure notification with essentially no custom code or additional services.

Why this answer

SageMaker endpoints emit CloudWatch metrics automatically, including invocation counts and 4XX and 5XX error metrics, so a CloudWatch alarm wired to an SNS topic gives failure notification without any custom instrumentation. Model Monitor and log-parsing pipelines solve different problems and require substantially more configuration.

Exam trap

The trap here is reaching for CloudTrail or Model Monitor to detect invocation failures, when the endpoint's built-in CloudWatch error metrics already expose them directly.

525
MCQeasy

A data scientist is using SageMaker to train a linear regression model. After training, they evaluate the model on the test set and get an R² of 0.95. However, when they deploy the model to a SageMaker endpoint and run predictions on new data, the predictions are far off. What is the most likely cause?

A.The endpoint is using a different inference script.
B.The test set is not representative of the production data distribution.
C.The model was trained with a wrong algorithm.
D.The model is overfitting the training data.
AnswerB

Strong test-set R² with poor endpoint predictions indicates distribution shift: the test set does not represent production data, so offline metrics are misleading. The model itself is sound; the evaluation data differs from live traffic, breaking generalisation.

Why this answer

A high R² of 0.95 on the test set indicates the model fits the test data well, but if the test set was drawn from the same distribution as the training data and does not reflect the real-world production data, the model will fail to generalize. In SageMaker, the endpoint serves predictions on live data that may have different statistical properties, leading to poor performance despite high test-set metrics. This is a classic case of dataset shift, not a model training or deployment configuration issue.

Exam trap

The trap here is that candidates confuse high test-set R² with model generalization, overlooking that the test set itself may be non-representative of production data, which is a core concept in the MLA-C01 exam under 'Model Evaluation and Validation'.

How to eliminate wrong answers

Option A is wrong because a different inference script would cause runtime errors or incorrect preprocessing, not systematically poor predictions on new data; SageMaker endpoints use the same inference code as the training container unless explicitly changed. Option C is wrong because using a wrong algorithm would typically result in poor training metrics (e.g., low R² on the test set), not a high R² of 0.95; the model converged well on the given data. Option D is wrong because overfitting would manifest as a large gap between training and test set performance (e.g., R² near 1.0 on training but much lower on test), but here the test R² is 0.95, suggesting the model generalizes to the test set; the issue is with production data differing from the test set.

Page 6

Page 7 of 9

Page 8

All pages