Courseiva

AWS Certified Machine Learning Engineer Associate MLA-C01 (MLA-C01) — Questions 226–300

665 questions total · 9pages · All types, answers revealed

Page 3

Page 4 of 9

Page 5
226
Multi-Selectmedium

A data scientist is building a text classification model using a pre-trained BERT model from the Hugging Face library on SageMaker. The scientist wants to fine-tune the model on a custom dataset. Which TWO steps are necessary to set up the fine-tuning job? (Select TWO.)

Select 2 answers
A.Use the HuggingFace estimator provided by SageMaker
B.Enable SageMaker Clarify for explainability during training
C.Build a custom Docker container with PyTorch and Transformers
D.Specify the PyTorch framework version and Transformers version in the estimator
E.Use SageMaker Processing to preprocess the data in parallel
AnswersA, D

The HuggingFace estimator is the SageMaker-provided class that packages the training container, script and hyperparameters for Hugging Face models. It satisfies the requirement to set up a fine-tuning job for the pre-trained BERT model on a custom dataset.

Why this answer

Option A is correct because the HuggingFace estimator from the SageMaker Python SDK is the purpose-built, supported interface for launching Hugging Face training jobs on SageMaker; it handles the training container, script entry point, and hyperparameters needed to fine-tune a pre-trained BERT model. Option D is correct because the HuggingFace estimator lets you pin the exact framework and library versions via the transformers_version and pytorch_version (or tensorflow_version) arguments, which is essential for reproducibility and compatibility with the pre-trained BERT checkpoint. Option B is not needed because SageMaker Clarify provides bias and explainability analysis, not a prerequisite for fine-tuning.

Option C is unnecessary since the HuggingFace estimator already supplies a managed container with PyTorch and Transformers, so building a custom Docker image is only required for unsupported dependencies. Option E is not required because data preprocessing can be done inside the training script or beforehand; SageMaker Processing is optional and not a mandatory setup step for the fine-tuning job.

Exam trap

AWS often tests the misconception that custom Docker containers are required for any non-standard framework, but the HuggingFace estimator eliminates that need by providing a managed environment with version control.

227
MCQeasy

A data science team has trained a PyTorch model for real-time inference and needs to deploy it on AWS with GPU acceleration while minimizing cold-start latency. Which SageMaker inference option should they choose?

A.Serverless inference
B.Batch transform
C.Asynchronous inference endpoint
D.Real-time endpoint with ml.g4dn instance
AnswerD

A real-time endpoint keeps the model loaded on a persistent ml.g4dn GPU instance, so inference requests avoid the container initialisation delay that serverless inference incurs. This satisfies the GPU acceleration and minimal cold-start latency constraints simultaneously.

Why this answer

Real-time endpoints with GPU instances (e.g., ml.g4dn) provide low latency and support GPU acceleration, suitable for interactive inference. Serverless inference does not support GPU instances, asynchronous inference is for non-real-time, and batch transform is for offline predictions.

228
MCQmedium

A team wants to use a custom PyTorch training script in SageMaker. They need to install additional Python packages not included in the base PyTorch container. Which approach should they take?

A.Use SageMaker Script Mode with a custom Dockerfile
B.Build a custom container with Docker
C.Install packages using a lifecycle configuration
D.Use the SageMaker PyTorch estimator with a requirements.txt file
AnswerD

The PyTorch estimator accepts a requirements.txt file, which SageMaker installs into the container before training begins. This adds the extra Python packages without building a custom image, satisfying the need for dependencies absent from the base container.

Why this answer

The SageMaker PyTorch estimator supports a 'requirements.txt' file in the source directory (specified via 'source_dir'), which SageMaker automatically installs into the training container before the script runs. This is the simplest, AWS-recommended way to add Python packages without building a custom image. It preserves the managed PyTorch container's optimizations while adding only the extra dependencies.

Exam trap

MLA-C01 often tests the boundary between Script Mode with requirements.txt (simple Python deps) and custom containers (OS-level or framework changes) — candidates over-engineer by choosing custom Docker builds when requirements.txt suffices.

How to eliminate wrong answers

Option A is wrong because SageMaker Script Mode does not accept a custom Dockerfile directly — Script Mode uses a pre-built AWS container, and while you can extend it, the Dockerfile approach is the same as building a custom container (Option B), not a distinct Script Mode feature. Option B is wrong because building a custom container with Docker is a heavier, more maintenance-intensive approach that is only necessary when you need OS-level packages, custom CUDA libraries, or a non-PyTorch framework — for simple Python package additions, requirements.txt is preferred. Option C is wrong because lifecycle configurations apply to SageMaker notebook instances (and some Studio apps), not to training jobs — they run shell scripts at notebook start/stop and have no effect on the training container.

229
Multi-Selectmedium

Which TWO SageMaker Pipelines steps are essential for automating a complete ML workflow from data processing to model deployment? (Choose 2.)

Select 2 answers
A.A TuningStep for hyperparameter tuning.
B.A ProcessingStep to run data preprocessing and feature engineering.
C.A TransformStep for batch inference on the training data.
D.A CreateModelStep (or RegisterModelStep) to register or deploy the trained model.
E.A ConditionStep to decide whether to train a model based on data quality.
AnswersB, D

A ProcessingStep executes the data preprocessing and feature engineering container, transforming raw input into the training-ready dataset. This satisfies the stem's requirement to automate the workflow's data processing stage within SageMaker Pipelines, feeding curated features downstream to training.

Why this answer

Option B (ProcessingStep) is correct because it is the pipeline step that runs a SageMaker Processing job to execute data preprocessing, feature engineering, and dataset splitting, which is the required starting point for an end-to-end workflow from raw data to a deployable model. Option D (CreateModelStep or RegisterModelStep) is correct because it produces the SageMaker model artifact needed to complete the workflow: CreateModelStep packages the model for deployment, while RegisterModelStep records it in the Model Registry for governed deployment, thereby covering the model deployment end of the pipeline. Option A (TuningStep) is not essential here because hyperparameter tuning is an optional optimization stage, not a required step for a complete processing-to-deployment workflow.

Option C (TransformStep) is not essential because batch inference on training data is an evaluation/inference activity, not a required stage for automating training and deployment. Option E (ConditionStep) is not essential because conditional branching on data quality is an optional control-flow enhancement rather than a mandatory step in a basic end-to-end pipeline.

Exam trap

AWS often tests the misconception that hyperparameter tuning or conditional logic are mandatory for a complete ML workflow, when in fact the minimal essential steps are data processing and model creation/deployment.

230
MCQeasy

A data scientist is working with a dataset that contains missing values in several numeric features. The data scientist wants to impute the missing values with the median of each feature. Which Amazon SageMaker Data Wrangler transformation should be used?

A.Replace missing with constant
B.Custom transform with Python
C.Drop missing rows
D.Handle missing values (with median strategy)
AnswerD

The Handle missing values transformation with the median strategy computes each numeric feature's median and substitutes it for nulls, satisfying the requirement to impute with the median rather than mean or mode. It operates directly on the dataset within Data Wrangler's flow.

Why this answer

Amazon SageMaker Data Wrangler includes a built-in 'Handle missing values' transformation that supports imputation with the median strategy. This directly matches the requirement to replace missing numeric values with the median of each feature without writing custom code.

Exam trap

The trap here is that candidates may confuse the 'Replace missing with constant' option (which uses a fixed value) with the median strategy, or they may overcomplicate the solution by choosing a custom Python transform when a built-in option exists.

How to eliminate wrong answers

Option A is wrong because 'Replace missing with constant' imputes a user-specified constant value (e.g., 0 or a fixed number), not the median of the feature. Option B is wrong because 'Custom transform with Python' would require writing custom Python code to compute and apply the median, which is unnecessary when a built-in transformation exists. Option C is wrong because 'Drop missing rows' removes entire rows with missing values, discarding potentially valuable data instead of imputing the missing values.

231
MCQhard

A team is deploying a TensorFlow model on a SageMaker real-time endpoint with automatic scaling. They set the scaling policy to target an average CPU utilization of 50%. However, during traffic spikes, the endpoint experiences high latency and 503 errors. The instance type is ml.c5.large. What should the team do to resolve this while minimizing cost?

A.Pre-warm the endpoint by keeping a fixed number of additional instances
B.Increase the scale-in cooldown period to avoid frequent downsizing
C.Change the instance type to a larger one like ml.c5.xlarge to handle the spikes
D.Add a scaling policy based on the number of concurrent requests per instance
AnswerD

Concurrent-requests scaling tracks actual demand rather than CPU, which lags behind request bursts on ml.c5.large. It adds capacity before latency and 503 errors occur, satisfying the responsiveness constraint while avoiding the cost of permanently larger instances.

Why this answer

Scaling based on CPU utilization alone is often insufficient for inference workloads where latency is the primary concern. By adding a scaling policy based on the number of concurrent requests per instance, the team can proactively scale out before CPU saturation occurs, reducing latency and eliminating 503 errors. SageMaker's automatic scaling supports multiple target tracking metrics, and using concurrent requests per instance aligns more closely with the actual demand on the model serving container.

Exam trap

The trap here is that candidates assume larger instances (Option C) are the only way to handle spikes, but the exam tests understanding that scaling policies based on the right metric (concurrent requests) can be more cost-effective and responsive than simply scaling up instance size.

How to eliminate wrong answers

Option A is wrong because pre-warming with a fixed number of additional instances increases cost without adapting to variable traffic patterns, and it does not address the root cause of scaling delays during spikes. Option B is wrong because increasing the scale-in cooldown period only delays instance termination, which does not help during rapid traffic increases; it may even worsen resource waste. Option C is wrong because moving to a larger instance type (ml.c5.xlarge) increases cost per instance and still relies on CPU-based scaling, which may still lag behind sudden spikes; it does not solve the fundamental issue of scaling responsiveness.

232
MCQhard

A company deploys a machine learning model as a SageMaker real-time endpoint. They need to implement a mechanism to automatically roll back to the previous model version if performance degrades after a deployment. Which approach should they use?

A.Manually update the endpoint to point to the previous model version
B.Configure the SageMaker endpoint deployment with traffic shifting and set up CloudWatch alarms to trigger automatic rollback
C.Create multiple endpoints and use Amazon Route 53 weighted routing to shift traffic
D.Use AWS CodeDeploy with Amazon EC2 instances behind an Elastic Load Balancer
AnswerB

Traffic shifting with CloudWatch alarms enables automatic rollback: alarms on endpoint metrics trigger the deployment to revert traffic to the previous model version. This satisfies the requirement for automatic rollback on performance degradation without manual intervention.

Why this answer

SageMaker endpoints support deployment with traffic shifting (e.g., canary or linear patterns) via the 'DeploymentConfig' parameter, and you can attach CloudWatch alarms to the endpoint's variant metrics. If the alarm triggers (e.g., due to increased error rate or latency), SageMaker automatically rolls back the traffic to the previous model version, ensuring minimal manual intervention and fast recovery.

Exam trap

The trap here is that candidates often confuse manual rollback (Option A) as acceptable automation, or they overcomplicate the solution with external services like Route 53 (Option C) or CodeDeploy (Option D), missing that SageMaker's native deployment configuration with CloudWatch alarms provides a fully automated, integrated rollback mechanism.

How to eliminate wrong answers

Option A is wrong because manual rollback is not automated and introduces human delay and error risk, failing the requirement for an automatic mechanism. Option C is wrong because using multiple endpoints with Route 53 weighted routing does not provide native integration with SageMaker's deployment monitoring or automatic rollback; it requires custom health check logic and does not leverage SageMaker's built-in traffic shifting and alarm-based rollback. Option D is wrong because AWS CodeDeploy with EC2 instances behind an ELB is designed for traditional application deployments, not for SageMaker endpoints; SageMaker endpoints are managed services that do not use EC2 instances or ELBs directly, and this approach would bypass SageMaker's native deployment capabilities.

233
MCQmedium

After deploying a model to a SageMaker endpoint, the operations team notices high inference latency. They suspect it is due to insufficient instance capacity. Which first step should they take to diagnose the issue?

A.Check AWS CloudTrail logs for API errors.
B.Use Amazon SageMaker Debugger to analyze inference performance.
C.Review Amazon CloudWatch metrics for the endpoint, such as CPUUtilization and Invocations.
D.Retrain the model with more training data.
AnswerC

CloudWatch publishes per-endpoint CPUUtilization and Invocations, letting the team confirm whether the instance is saturated before changing anything. These metrics directly test the capacity hypothesis, distinguishing genuine under-provisioning from model or payload latency without redeploying.

Why this answer

Amazon CloudWatch metrics for a SageMaker endpoint, such as `CPUUtilization`, `MemoryUtilization`, and `Invocations`, directly indicate whether the instance is overloaded. High `CPUUtilization` combined with a high `Invocations` count and increased latency strongly suggests insufficient instance capacity. This is the standard first diagnostic step for capacity-related performance issues.

Exam trap

The trap here is that candidates confuse SageMaker Debugger (for training debugging) with inference monitoring tools, or they assume CloudTrail can provide performance metrics, when in fact CloudWatch is the correct service for real-time endpoint health and capacity diagnostics.

How to eliminate wrong answers

Option A is wrong because AWS CloudTrail logs record API calls (e.g., CreateEndpoint, InvokeEndpoint) and are used for auditing and security, not for real-time performance metrics like latency or CPU load. Option B is wrong because Amazon SageMaker Debugger is designed for monitoring training jobs (e.g., gradient norms, loss convergence) and does not provide inference-time performance metrics for deployed endpoints. Option D is wrong because retraining the model with more data addresses model accuracy, not inference latency caused by insufficient instance capacity; it would not reduce the time taken to process each request on an already overloaded instance.

234
MCQhard

A machine learning engineer is using SageMaker to train a model with the built-in LightGBM algorithm. The engineer wants to use early stopping to prevent overfitting. The training job is configured with a validation dataset. Which hyperparameter should be set to enable early stopping?

A.early_stopping_rounds
B.num_iterations
C.early_stopping
D.num_boost_round
AnswerA

Setting early_stopping_rounds halts training once the validation metric fails to improve for that many consecutive boosting rounds, retaining the best iteration. This directly satisfies the requirement to prevent overfitting using the configured validation dataset in the built-in LightGBM algorithm.

Why this answer

In SageMaker's built-in LightGBM algorithm, the hyperparameter `early_stopping_rounds` controls early stopping. When a validation dataset is provided, training will stop if the evaluation metric does not improve for the specified number of consecutive rounds, preventing overfitting.

Exam trap

The trap here is that candidates confuse the generic concept of early stopping with the exact hyperparameter name used by SageMaker's built-in LightGBM, often selecting `early_stopping` (which is not a valid parameter) instead of the precise `early_stopping_rounds`.

How to eliminate wrong answers

Option B is wrong because `num_iterations` sets the total number of boosting iterations, not the early stopping behavior; it defines the maximum number of rounds, not a stopping criterion. Option C is wrong because `early_stopping` is not a valid hyperparameter in SageMaker's LightGBM implementation; the correct parameter name is `early_stopping_rounds`. Option D is wrong because `num_boost_round` is an alias for `num_iterations` in some frameworks but is not the hyperparameter used for early stopping in SageMaker's LightGBM.

235
MCQeasy

A company wants to automate its machine learning pipeline using AWS CodePipeline and Amazon SageMaker. The pipeline should train a model, evaluate it, and if the evaluation passes, register the model in the SageMaker Model Registry. Which service should the company use to orchestrate the training and evaluation steps?

A.AWS CodePipeline
B.AWS Glue Workflows
C.AWS Step Functions
D.Amazon SageMaker Pipelines
AnswerD

Amazon SageMaker Pipelines provides native orchestration of training, evaluation and conditional model-registration steps, satisfying the requirement to register only when evaluation passes. Its ConditionStep gates registration on the evaluation metric, while CodePipeline triggers the pipeline rather than orchestrating ML steps. This integrates directly with the SageMaker Model Registry.

Why this answer

Amazon SageMaker Pipelines is the correct choice because it is a purpose-built, fully managed service for creating end-to-end machine learning workflows directly within the SageMaker ecosystem. It natively integrates with SageMaker training jobs, processing jobs for evaluation, and the Model Registry for conditional registration, allowing the entire pipeline—train, evaluate, and conditionally register—to be defined as a directed acyclic graph (DAG) of steps without needing to stitch together separate services.

Exam trap

The trap here is that candidates may confuse AWS Step Functions (a general-purpose orchestrator) with SageMaker Pipelines (a specialized ML orchestrator), overlooking that SageMaker Pipelines provides built-in SageMaker step types and native Model Registry integration, which Step Functions lacks without custom Lambda functions.

How to eliminate wrong answers

Option A is wrong because AWS CodePipeline is a CI/CD service designed for software delivery pipelines (e.g., building, testing, deploying applications), not for orchestrating ML training and evaluation steps that require direct integration with SageMaker resources like training jobs or the Model Registry. Option B is wrong because AWS Glue Workflows are used for orchestrating ETL (extract, transform, load) jobs and data preparation tasks within AWS Glue, not for managing ML training or model evaluation workflows. Option C is wrong because while AWS Step Functions can orchestrate SageMaker API calls, it requires custom integration code and does not provide native, declarative support for SageMaker-specific steps like training, tuning, or model registration, making it less efficient and more error-prone than SageMaker Pipelines for this use case.

236
MCQmedium

A team receives alerts that their SageMaker endpoint latency has increased significantly. They check CloudWatch metrics and see Invocations rising, but ModelLatency remains stable. Which metric should they investigate to find the source of the increased latency?

A.OverheadLatency
B.ModelLatency
C.5XXError
D.4XXError
AnswerA

OverheadLatency measures time spent outside the model, covering request routing, queueing and response handling. Since Invocations rose while ModelLatency stayed flat, the added delay sits in this overhead, not in model execution, pinpointing the source.

Why this answer

OverheadLatency measures the time taken by the SageMaker infrastructure to handle requests before and after model inference, including request routing, authentication, and response processing. Since ModelLatency is stable but total endpoint latency has increased, the extra time must be in the overhead component, making OverheadLatency the correct metric to investigate.

Exam trap

The trap here is that candidates assume increased Invocations directly cause higher ModelLatency, but the exam tests the distinction between inference time and infrastructure overhead, leading them to incorrectly select ModelLatency instead of OverheadLatency.

How to eliminate wrong answers

Option B is wrong because ModelLatency is explicitly stated as stable, so it cannot be the source of increased latency. Option C is wrong because 5XXError indicates server-side errors, not latency; while errors can correlate with latency, the question asks for the metric directly measuring the latency increase. Option D is wrong because 4XXError indicates client-side errors (e.g., invalid requests), which do not directly cause increased endpoint latency.

237
MCQhard

A data science team wants to host 50 different models for a recommendation engine. Each model is small (under 100 MB) and traffic patterns are unpredictable. They need to minimize cost and operational overhead. Which approach should they take?

A.Deploy each model to its own real-time endpoint
B.Use SageMaker serverless inference for each model
C.Use a single multi-model endpoint (MME)
D.Use a single multi-container endpoint
AnswerC

A multi-model endpoint hosts many models behind one container, so 50 small models share a single endpoint rather than 50 separate ones. This directly minimises cost and operational overhead while handling unpredictable traffic through shared autoscaling.

Why this answer

A single multi-model endpoint (MME) allows hosting multiple models (up to thousands) on the same endpoint, sharing the underlying compute instance. This minimizes cost and operational overhead for small models (under 100 MB) with unpredictable traffic, as the endpoint dynamically loads and unloads models from Amazon S3 into memory based on incoming requests, eliminating the need for separate endpoints or idle compute.

Exam trap

AWS often tests the distinction between multi-model endpoints (for multiple independent models) and multi-container endpoints (for a single model with multiple containers), leading candidates to confuse the two and incorrectly choose option D.

How to eliminate wrong answers

Option A is wrong because deploying each model to its own real-time endpoint would require 50 separate endpoints, each with its own compute instance, leading to high cost and operational overhead due to idle resources during unpredictable traffic patterns. Option B is wrong because SageMaker serverless inference is designed for infrequent or sporadic traffic, but it still requires a separate serverless endpoint per model, incurring per-request costs and cold-start latency for each model, which does not minimize cost or overhead for 50 models. Option D is wrong because a single multi-container endpoint is intended for hosting multiple containers that serve a single model (e.g., pre-processing and inference), not for hosting multiple independent models; it does not support dynamic model loading and would require separate containers for each model, defeating the purpose of consolidation.

238
MCQmedium

A company uses SageMaker Clarify to detect bias in their training data. They find that the model has a high disparate impact for a protected attribute. What should they do to mitigate this bias during training?

A.Use SageMaker Clarify’s built-in bias mitigation algorithm during training
B.Remove the protected attribute from the dataset
C.Increase the model complexity to capture more patterns
D.Preprocess the data using techniques like reweighing or resampling to reduce bias
AnswerD

Reweighing assigns weights to training examples so the protected and unprotected groups contribute proportionally, while resampling adjusts class balance. Applied before training, these preprocessing techniques reduce the disparate impact measured by SageMaker Clarify at its source.

Why this answer

SageMaker Clarify detects bias but does not automatically mitigate it during training. To reduce disparate impact, the correct approach is to preprocess the training data using bias mitigation techniques such as reweighing (assigning weights to instances to balance outcomes across groups) or resampling (oversampling underrepresented groups or undersampling overrepresented ones). These methods adjust the data distribution before training, directly addressing the source of bias and reducing disparate impact.

Exam trap

MLA-C01 often tests the misconception that SageMaker Clarify can automatically mitigate bias during training, when in fact it only detects and explains bias; mitigation requires separate preprocessing, in-processing, or post-processing techniques.

How to eliminate wrong answers

Option A is wrong because SageMaker Clarify is a bias detection and explainability tool, not a bias mitigation algorithm; it does not modify training or apply corrections automatically. Option B is wrong because simply removing the protected attribute does not eliminate bias—other features may act as proxies, and the model can still produce disparate outcomes. Option C is wrong because increasing model complexity can exacerbate bias by fitting to spurious correlations and does not address the underlying data imbalance.

239
MCQhard

A machine learning engineer is deploying a model to a SageMaker endpoint and wants to ensure that the model's predictions can be explained. The engineer needs to understand which features contributed most to each prediction. Which SageMaker feature should be used?

A.SageMaker Model Monitor
B.SageMaker Clarify
C.SageMaker Debugger
D.SageMaker Experiments
AnswerB

SageMaker Clarify provides explainability by computing feature attributions using algorithms like SHAP. It helps understand which features contributed most to individual predictions. This directly meets the requirement to explain predictions. Clarify can be integrated with SageMaker endpoints and provides both global and local explanations.

Why this answer

SageMaker Clarify is specifically designed to provide feature attributions and explain model predictions. It uses SHAP to compute the contribution of each feature to a prediction, which is exactly what the engineer needs. Other services like Debugger, Model Monitor, and Experiments serve different purposes and do not offer prediction-level explainability.

Exam trap

The trap here is confusing monitoring or debugging services with explainability, or assuming that Model Monitor provides feature importance.

240
MCQmedium

A media company runs a real-time recommendation model on a SageMaker endpoint. Traffic triples every evening between 18:00 and 22:00 and drops to near zero overnight, and the team wants to cut costs without underserving evening users. They want the endpoint to scale out automatically as invocations rise and scale back in when traffic falls. Which solution should they implement?

A.Create a second endpoint and use an Application Load Balancer to route evening traffic between the two endpoints.
B.Deploy the model to a SageMaker Serverless Inference endpoint and set the maximum concurrency to a fixed value equal to the evening peak.
C.Enable automatic scaling on the endpoint with a target-tracking scaling policy based on the SageMakerVariantInvocationsPerInstance metric.
D.Increase the instance count of the endpoint to the evening peak permanently and rely on Savings Plans to offset the extra cost.
AnswerC

Target tracking on SageMakerVariantInvocationsPerInstance lets Application Auto Scaling add instances when per-instance invocation load exceeds the target and remove them when it drops, matching the nightly peak-and-trough pattern without manual intervention. This is the documented approach for provisioning real-time endpoint capacity dynamically, and it preserves latency by scaling out before instances saturate.

Why this answer

Target-tracking automatic scaling based on SageMakerVariantInvocationsPerInstance is the native mechanism for elastic real-time endpoints. It registers the variant as a scalable target with Application Auto Scaling, then adds or removes instances to hold the metric near the target. This directly matches a recurring evening peak followed by an overnight trough while keeping latency low, unlike fixed serverless concurrency or static over-provisioning.

Exam trap

The trap here is assuming Serverless Inference is always the cheapest elastic option, when a fixed maximum concurrency reintroduces the same idle cost as a permanently over-provisioned endpoint.

241
MCQmedium

A team uses SageMaker Clarify to monitor bias drift on a deployed model. They have defined a baseline with training data and set up a monitoring schedule. After one month, they receive a violation report indicating that the post-training metrics have deviated from the baseline. What does this violation indicate?

A.The model's predictions relative to sensitive attributes have shifted compared to the training baseline
B.The SHAP values for features have changed
C.The model's predictions are becoming less accurate
D.The distribution of input features has changed
AnswerA

Clarify's post-training bias metrics compare predicted labels across sensitive attribute groups against the training baseline. A violation means those group-conditional prediction rates have drifted, indicating the model now treats protected groups differently than when trained.

Why this answer

SageMaker Clarify bias drift monitoring compares predicted outcomes (post-training) against the baseline to detect changes in fairness metrics like disparate impact. It does not measure prediction accuracy or data quality.

242
MCQmedium

A team is using Amazon SageMaker Processing for data preprocessing. They have a Parquet dataset in Amazon S3. Which configuration will provide the most efficient reading of the dataset during processing?

A.Read the Parquet files as text using SparkContext.textFile
B.Split the dataset into many small Parquet files (e.g., 1 MB each)
C.Convert the Parquet files to CSV before processing
D.Read the Parquet files directly using SparkSession.read.parquet
AnswerD

SparkSession.read.parquet reads Parquet's columnar, compressed format directly, enabling predicate pushdown and column pruning so SageMaker Processing scans only needed columns and row groups from Amazon S3, avoiding full-dataset deserialisation and delivering the most efficient read.

Why this answer

SageMaker Processing natively integrates with Apache Spark, and reading Parquet files directly via `SparkSession.read.parquet` leverages columnar storage, predicate pushdown, and compression (e.g., Snappy) to minimize I/O and deserialization overhead. This approach is far more efficient than text-based or format-conversion methods, as Parquet is optimized for analytical workloads and preserves schema information.

Exam trap

AWS often tests the misconception that many small files improve parallelism, but in distributed systems like Spark on SageMaker, small files increase S3 API call overhead and scheduler latency, making larger Parquet files (e.g., 128 MB–1 GB) far more efficient for reading.

How to eliminate wrong answers

Option A is wrong because `SparkContext.textFile` reads data as plain text lines, which is incompatible with binary Parquet format and would result in corrupted data or require manual parsing, losing all columnar optimization. Option B is wrong because splitting the dataset into many small 1 MB Parquet files increases S3 LIST and GET request overhead, causing task scheduling delays and poor I/O throughput due to excessive file metadata operations. Option C is wrong because converting Parquet to CSV before processing introduces unnecessary serialization/deserialization costs, increases data size (CSV lacks compression and columnar storage), and discards schema and type information, leading to slower read performance.

243
MCQeasy

A machine learning engineer wants to automatically track hyperparameters, metrics, and artifacts for multiple training runs. Which SageMaker feature should they use?

A.SageMaker Debugger
B.SageMaker Model Monitor
C.SageMaker Experiments
D.SageMaker Clarify
AnswerC

SageMaker Experiments automatically captures hyperparameters, metrics, and artifacts across training runs, satisfying the requirement to track multiple runs without manual logging. It records each trial as a run within an experiment, enabling comparison and reproducibility.

Why this answer

SageMaker Experiments is the feature designed to track, organize, and compare machine learning training runs, including hyperparameters, metrics, and artifacts. It automatically logs these elements when integrated with SageMaker training jobs, enabling reproducibility and experiment comparison. This directly matches the requirement to track multiple runs.

Exam trap

MLA-C01 often tests the confusion between SageMaker Experiments (tracking) and SageMaker Debugger (debugging) — candidates pick Debugger because 'tracking training' sounds like debugging.

How to eliminate wrong answers

Option A is wrong because SageMaker Debugger focuses on real-time debugging of training jobs (tensor analysis, profiling), not on tracking hyperparameters and artifacts across runs. Option B is wrong because SageMaker Model Monitor detects drift and quality issues in deployed models, not training run metadata. Option D is wrong because SageMaker Clarify provides bias detection and explainability, not experiment tracking.

244
MCQhard

A financial services company deploys a fraud detection model on a SageMaker real-time endpoint. The inference logic includes a pre-processing step that requires access to a DynamoDB table for user metadata. The model container is a custom Docker image. How should the team grant the endpoint access to DynamoDB?

A.Store IAM credentials in the container image as environment variables
B.Attach an IAM instance profile to the underlying EC2 instance
C.Create an IAM role with DynamoDB read access and assign it to the SageMaker endpoint as the execution role
D.Retrieve temporary credentials from AWS Secrets Manager within the container code
AnswerC

The endpoint's execution role is the identity SageMaker assumes when running the container, so attaching DynamoDB read permissions to that role lets the custom inference code call DynamoDB during pre-processing. This satisfies the requirement without embedding credentials in the image.

Why this answer

SageMaker endpoints require an IAM execution role to be assigned at creation time. This role defines the permissions the endpoint's container has when making AWS API calls, such as reading from DynamoDB. By attaching a policy with DynamoDB read access to this execution role, the endpoint securely obtains temporary credentials via the AWS STS service, eliminating the need to hardcode or manage long-term credentials.

Exam trap

The trap here is that candidates confuse SageMaker endpoints with EC2-based deployments and incorrectly think they need to manage instance profiles or embed credentials, when in fact SageMaker abstracts the underlying compute and uses an execution role for all API access.

How to eliminate wrong answers

Option A is wrong because storing IAM credentials as environment variables in a container image is a security anti-pattern; credentials would be baked into the image, exposed in the container's environment, and not automatically rotated. Option B is wrong because SageMaker endpoints do not run on EC2 instances that you manage; they run on SageMaker-managed infrastructure, so attaching an instance profile to an underlying EC2 instance is not applicable. Option D is wrong because while Secrets Manager can store credentials, the container code would still need permissions to access Secrets Manager itself, and the standard, simpler approach is to use the endpoint's execution role rather than managing temporary credentials manually.

245
MCQmedium

A data scientist is performing text preprocessing for a sentiment analysis model. The dataset contains many stop words and rare words. Which combination of preprocessing steps will reduce dimensionality and improve model performance?

A.Remove all words shorter than 3 characters and apply label encoding
B.Only tokenization without any removal
C.Tokenization, stop-word removal, and TF-IDF
D.Tokenization and one-hot encoding of words
AnswerC

Tokenisation splits text into units, stop-word removal discards high-frequency, low-information terms, and TF-IDF down-weights terms appearing across many documents while rewarding rare, discriminative ones. Together these directly reduce the feature space created by the dataset's many stop words and rare words, satisfying the dimensionality-reduction constraint.

Why this answer

Tokenization splits raw text into individual terms, stop-word removal eliminates high-frequency but low-information words (e.g., 'the', 'is', 'and'), and TF-IDF assigns weights that downscale terms appearing in many documents while upscaling rare, discriminative terms. Together they reduce the feature space and emphasize sentiment-bearing vocabulary, which directly improves model performance on text classification tasks. This is the standard NLP preprocessing pipeline for sentiment analysis.

Exam trap

MLA-C01 often tests the misconception that any encoding scheme (label or one-hot) is a valid substitute for TF-IDF in text preprocessing, when in fact only TF-IDF captures term importance and reduces the impact of stop words.

How to eliminate wrong answers

Option A is wrong because removing short words discards meaningful tokens like 'no', 'not', 'bad', and 'sad' that are critical for sentiment, and label encoding is inappropriate for text features since it imposes a false ordinal relationship on words. Option B is wrong because tokenization alone retains all stop words and rare noise, leaving dimensionality high and adding no discriminative signal. Option D is wrong because one-hot encoding produces an extremely sparse, high-dimensional vector per word with no weighting for term importance, and it ignores stop-word removal entirely.

246
MCQhard

An organization uses SageMaker Studio and needs to restrict Studio's internet access while allowing users to install custom packages from a private PyPI mirror hosted in a VPC. Which networking configuration should they use?

A.Use a NAT gateway to allow outbound traffic to the private PyPI mirror
B.Disable internet access for Studio and rely on SageMaker's default VPC configuration
C.Disable internet access for Studio and configure VPC-only mode, then use a VPC endpoint to the private PyPI mirror
D.Enable internet access for Studio and use a VPC endpoint to the private PyPI mirror
AnswerC

VPC-only mode blocks internet; VPC endpoint to the private mirror allows package installation from the VPC.

Why this answer

To restrict Studio's internet access while allowing access to a private PyPI mirror in a VPC, you must disable internet access for Studio and configure VPC-only mode, then create a VPC endpoint (e.g., an interface endpoint or a private link) to the mirror. This keeps all traffic within the VPC and avoids public internet exposure. The VPC endpoint provides private connectivity to the mirror without a NAT gateway or internet gateway.

Exam trap

MLA-C01 often tests the confusion between NAT gateway (internet access) and VPC endpoint (private access), tricking candidates into choosing NAT when the requirement is to restrict internet while allowing private resources.

How to eliminate wrong answers

Option A is wrong because a NAT gateway provides outbound internet access, which contradicts the requirement to restrict Studio's internet access — it would allow broader internet connectivity, not just the private mirror. Option B is wrong because disabling internet access and relying on SageMaker's default VPC configuration does not provide a path to the private PyPI mirror; the default VPC may not have the necessary endpoints or routing. Option D is wrong because enabling internet access for Studio violates the restriction requirement, even if a VPC endpoint is also used.

247
MCQeasy

An organization needs to ensure that all data used for inference on a SageMaker endpoint is encrypted at rest. The endpoint uses a SageMaker-provided container. Which configuration should be applied?

A.Use a custom container with built-in encryption
B.Specify a KMS key in the endpoint configuration
C.Enable network isolation mode
D.Enable inter-container traffic encryption
AnswerB

A KMS key encrypts the ML storage volume attached to the endpoint, ensuring data at rest is encrypted.

Why this answer

SageMaker endpoints use AWS KMS for encryption at rest. By specifying a KMS key in the endpoint configuration, the data in the attached ML storage volume is encrypted. Inter-container traffic encryption is for encryption in transit.

248
MCQeasy

A data engineer is setting up a Glue ETL job to process a large dataset stored in Amazon S3. The job needs to read data in Parquet format, apply a filter, and write the results back to S3 in Parquet. The engineer wants to minimize the cost and runtime. Which optimization technique is MOST effective?

A.Increase the number of DPUs to the maximum allowed.
B.Use column pruning to read only the columns needed in the transformation.
C.Convert the data to JSON format for faster read performance.
D.Use a single large file instead of partitioning.
AnswerB

Parquet is columnar, so column pruning reads only the columns the filter and transformation reference, skipping irrelevant data on disk. This cuts bytes scanned and I/O, reducing both runtime and cost more effectively than row-based filtering or partition changes alone.

Why this answer

Column pruning is the most effective optimization because AWS Glue with the Spark engine can push column selection down to the Parquet reader, so only the required columns are read from S3 and processed. Parquet is a columnar format, meaning each column is stored separately, so skipping unneeded columns dramatically reduces I/O, memory footprint, and shuffle volume. This directly lowers DPU-hour consumption and runtime, which are the two cost drivers for Glue ETL jobs.

Exam trap

MLA-C01 often tests the misconception that throwing more DPUs at a Glue job is the universal performance fix, when the real lever is reducing the data read via column pruning and predicate pushdown.

How to eliminate wrong answers

Option A is wrong because increasing DPUs to the maximum scales cost linearly and often does not reduce runtime proportionally — it treats the symptom rather than eliminating unnecessary I/O, and Glue charges per DPU-hour. Option C is wrong because converting Parquet to JSON is a downgrade: JSON is row-based, uncompressed by default, and far slower to parse, so it would increase both runtime and cost. Option D is wrong because using a single large file instead of partitioning removes parallelism — Spark splits work by file and partition, so one large file forces a single task to read everything, eliminating the distributed advantage.

249
Multi-Selecthard

A company is running a SageMaker endpoint serving multiple models. They need to monitor for data drift and model quality. Which THREE actions are necessary? (Choose three.)

Select 3 answers
A.Deploy a shadow endpoint for comparison
B.Enable data capture on the endpoint
C.Use SageMaker Debugger for monitoring
D.Create a SageMaker Model Monitor schedule
E.Configure baseline constraints from training data
AnswersB, D, E

Data capture records endpoint request and response payloads to Amazon S3, supplying the actual inference data that Model Monitor analyses. Without captured data there is nothing to compare against baselines, so this is a prerequisite for detecting data drift and measuring model quality.

Why this answer

Option B is correct because SageMaker Model Monitor requires data capture to be enabled on the endpoint (via DataCaptureConfig) so that inference requests and responses are logged to Amazon S3 for drift and quality analysis. Option D is correct because a Model Monitor schedule (created with CreateMonitoringSchedule) is what actually runs the monitoring jobs on a recurring basis against the captured data. Option E is correct because Model Monitor compares live data to a baseline; constraints and statistics generated from the training dataset (via DefaultModelMonitor.suggest_baseline) define the expected schema and thresholds used to detect drift and quality violations.

Option A is not required, since a shadow endpoint is for testing a new model variant, not for monitoring drift or quality of the existing endpoint. Option C is not required, because SageMaker Debugger is used for debugging training jobs (tensors, metrics, profiling), not for production data drift or model quality monitoring.

Exam trap

The trap here is that candidates confuse SageMaker Debugger (for training) with SageMaker Model Monitor (for inference), leading them to select Debugger instead of the correct monitoring schedule and baseline configuration.

250
MCQmedium

A financial services company trains multiple models on SageMaker and needs to track hyperparameters, metrics, and artifacts for each experiment. Which SageMaker feature should they use to organize and compare experiments?

A.SageMaker Model Registry
B.SageMaker Pipelines
C.SageMaker Experiments
D.SageMaker Debugger
AnswerC

SageMaker Experiments groups training runs into experiments and trials, automatically capturing hyperparameters, metrics, and artifacts for each job. This directly satisfies the requirement to organise and compare runs across multiple models, providing the lineage and side-by-side analysis the financial services company needs.

Why this answer

SageMaker Experiments is purpose-built for tracking, organizing, and comparing machine learning experiments — it captures hyperparameters, metrics, input datasets, and output artifacts for each trial and groups them into experiments and runs. This lets data scientists visualize and compare results across many training jobs in a single view.

Exam trap

MLA-C01 often tests the confusion between SageMaker Experiments (tracking/comparing runs) and Model Registry (versioning/approving models) — both involve 'models' but serve different lifecycle stages.

How to eliminate wrong answers

Option A is wrong because Model Registry is for cataloging, versioning, and approving trained models for deployment, not for tracking training-time hyperparameters and metrics. Option B is wrong because Pipelines is an orchestration service for CI/CD-style ML workflows, not an experiment-tracking tool. Option D is wrong because Debugger focuses on detecting training anomalies (vanishing gradients, overfitting) via tensor analysis, not on organizing and comparing experiments.

251
MCQhard

A company operates an e-commerce platform that uses a machine learning model to recommend products to users. The model is deployed on an Amazon SageMaker endpoint with automatic scaling enabled based on average CPU utilization. The model was trained on historical data and is updated weekly. Recently, the platform experienced a flash sale event that caused a sudden spike in traffic. During the event, the endpoint's latency increased dramatically, and many requests timed out. After the event, the team reviews the CloudWatch metrics and notices that the CPU utilization never exceeded 70%, and the scaling policy was triggered but instances took several minutes to become available. The team wants to prevent similar issues in future flash sales. Which course of action would be MOST effective?

A.Use predictive scaling based on historical traffic patterns.
B.Lower the CPU utilization threshold for the scaling policy to 40%.
C.Switch to larger instance types to handle higher CPU loads.
D.Implement scheduled scaling to add capacity ahead of known flash sales.
AnswerD

Scheduled scaling provisions instances before the flash sale begins, eliminating the several-minute cold-start delay that caused timeouts. Since traffic spikes are predictable events, pre-warming capacity bypasses the reactive CPU-based policy's lag entirely, keeping latency low from the first request.

Why this answer

Scheduled scaling allows you to proactively add capacity ahead of known traffic events like flash sales, eliminating the cold-start delay that occurs when reactive scaling policies (like those based on CPU utilization) must launch new instances. During the flash sale, the scaling policy was triggered but instances took minutes to become available, causing timeouts; scheduled scaling pre-warms the endpoint by adjusting the desired instance count before the traffic spike hits.

Exam trap

The trap here is that candidates assume reactive scaling (lowering thresholds or using predictive scaling) can handle sudden spikes, but the exam tests your understanding that provisioning latency is the bottleneck, and only proactive scheduled scaling can eliminate that delay for known events.

How to eliminate wrong answers

Option A is wrong because predictive scaling relies on historical traffic patterns to forecast future demand, but a flash sale is an irregular, planned event that may not follow those patterns, and predictive scaling still involves a delay in provisioning instances. Option B is wrong because lowering the CPU threshold to 40% would cause the scaling policy to trigger earlier, but it does not address the fundamental issue that new instances take several minutes to become available (cold-start latency), so requests would still time out during that provisioning window. Option C is wrong because switching to larger instance types increases the per-instance capacity but does not eliminate the cold-start delay when scaling out; during a sudden spike, even larger instances would eventually be overwhelmed if the scaling action itself is too slow.

252
MCQeasy

An ML team wants to deploy a model that was trained using XGBoost in SageMaker. They want to use the built-in XGBoost algorithm container for inference. Which inference option requires the least custom code?

A.Create a custom Docker container with XGBoost and deploy to an endpoint
B.Deploy to a real-time endpoint using the built-in XGBoost container
C.Attach Elastic Inference to a generic container
D.Use SageMaker Python SDK to download the model and run local inference
AnswerB

SageMaker's built-in XGBoost container already implements the inference handler and serving stack, so deploying it to a real-time endpoint needs only a model artefact and an endpoint configuration. No custom inference script or Dockerfile is required, minimising code.

Why this answer

The built-in XGBoost container in SageMaker is pre-configured with the XGBoost serving stack, including the necessary inference code and dependencies. Deploying a model trained with XGBoost to a real-time endpoint using this container requires no custom inference script or Docker image, only the model artifact and endpoint configuration. This minimizes custom code to just the SageMaker SDK calls for creating the model and endpoint.

Exam trap

AWS often tests the misconception that Elastic Inference can accelerate any ML model, but it is specifically designed for deep learning models and does not apply to tree-based algorithms like XGBoost.

How to eliminate wrong answers

Option A is wrong because creating a custom Docker container with XGBoost introduces unnecessary custom code and maintenance overhead, whereas the built-in container already provides the same functionality. Option C is wrong because Elastic Inference is an acceleration technology for deep learning models (e.g., TensorFlow, PyTorch) and is not compatible with XGBoost, which is a gradient boosting framework; attaching it to a generic container would not reduce custom code and would be architecturally incorrect. Option D is wrong because using the SageMaker Python SDK to download the model and run local inference moves the inference workload outside of SageMaker's managed infrastructure, requiring custom orchestration code and defeating the purpose of a managed deployment.

253
MCQhard

A company wants to serve a large ensemble of models using NVIDIA Triton Inference Server on SageMaker for high throughput GPU inference. Which SageMaker inference option supports this?

A.Asynchronous Inference
B.Multi-model endpoint
C.Serverless Inference
D.Real-time endpoint with a custom container running Triton
AnswerD

A real-time endpoint with a custom container lets you run NVIDIA Triton Inference Server directly, enabling ensemble execution, dynamic batching and concurrent model execution on GPU instances. This satisfies the high-throughput GPU inference requirement, which single-model SageMaker containers cannot provide.

Why this answer

NVIDIA Triton Inference Server is a custom inference server that supports multiple frameworks, model ensembles, and dynamic batching. To use it on SageMaker, you deploy a real-time endpoint with a custom container that runs Triton, which supports GPU inference and high throughput. This is the only option that explicitly supports Triton.

Exam trap

MLA-C01 often tests the assumption that any SageMaker hosting option can run Triton, when only a real-time endpoint with a custom container supports arbitrary inference servers like Triton.

How to eliminate wrong answers

Option A is wrong because Asynchronous Inference is a request-queueing mode for large payloads/long processing, not a mechanism for running Triton. Option B is wrong because Multi-Model Endpoint is a SageMaker hosting feature for many models on one endpoint, but it does not natively run Triton; Triton requires a custom container. Option C is wrong because Serverless Inference does not support custom containers with GPU and has cold-start/memory limits incompatible with high-throughput Triton ensembles.

254
MCQhard

A data engineer is using Amazon SageMaker Data Wrangler to create a data preparation flow for a dataset with 500 columns, many of which are highly correlated. The goal is to reduce dimensionality while preserving interpretability. Which built-in transform in Data Wrangler should be applied?

A.Imputation
B.Principal Component Analysis (PCA)
C.StandardScaler
D.Feature Selection (correlation-based)
AnswerD

Correlation-based feature selection drops one of each pair of highly correlated columns, directly reducing dimensionality while retaining the original, interpretable features. Unlike PCA, which produces opaque linear combinations, it satisfies the stem's interpretability constraint and scales to 500 columns without manual inspection.

Why this answer

The goal is to reduce dimensionality while preserving interpretability. SageMaker Data Wrangler's built-in Feature Selection (correlation-based) transform identifies and removes highly correlated columns, directly reducing the number of features without transforming the original variables into new, uninterpretable components. This preserves the meaning of each selected column, which is essential when interpretability is a priority.

Exam trap

The trap here is that candidates often confuse dimensionality reduction with PCA, assuming it is always the best choice, but the question explicitly requires preserving interpretability, which PCA inherently sacrifices.

How to eliminate wrong answers

Option A is wrong because imputation is used to fill missing values, not to reduce dimensionality or handle correlated columns. Option B is wrong because Principal Component Analysis (PCA) creates new orthogonal components that are linear combinations of original features, which reduces dimensionality but destroys interpretability since the components are not directly tied to the original columns. Option C is wrong because StandardScaler only standardizes features by removing the mean and scaling to unit variance; it does not reduce the number of columns or address correlation.

255
MCQmedium

A retailer runs a nightly batch scoring job that processes 40 GB of transaction data and writes predictions to S3. Occasionally a single partition is corrupt, causing the entire job to fail after several hours. The team wants the job to skip the corrupt partition, log which partition failed, and still complete processing of the remaining data with minimal changes to their existing SageMaker Processing job. Which change should they make?

A.Wrap the per-partition processing logic in try/except inside the Processing container script, record the failing partition key to CloudWatch Logs, and continue to the next partition.
B.Configure the Processing job with a retry policy and a maximum retry count of three so transient failures are retried automatically.
C.Convert the Processing job to a SageMaker Training job with checkpointing enabled so it can resume after the corrupt partition.
D.Increase the Processing job's instance count and volume size so the corrupt partition is retried on a different instance.
AnswerA

Handling exceptions per partition inside the container lets the job skip a corrupt input, emit the partition key to CloudWatch Logs, and proceed with the remaining data. It requires only a code change to the existing Processing job, preserving the current orchestration and S3 output path. This directly satisfies fault isolation, failure logging, and job completion with minimal disruption.

Why this answer

Adding per-partition exception handling inside the Processing container isolates corruption to the affected partition, logs its key for follow-up, and allows the job to finish the rest of the data. This is a localized code change that preserves the existing SageMaker Processing job structure and S3 outputs. It meets fault isolation, failure visibility, and completion requirements with the least disruption.

Exam trap

The trap here is reaching for infrastructure-level retries or resizing, when a deterministic data corruption must be handled in application code to isolate the bad partition.

256
MCQhard

A machine learning engineer is using SageMaker Automatic Model Tuning (AMT) to optimize hyperparameters for a random forest model. The engineer notices that the tuning job is taking too long and many hyperparameter combinations are being evaluated but not improving the objective metric. Which action should the engineer take to make the tuning more efficient?

A.Switch the strategy from Bayesian to random search
B.Use a smaller instance type for each training job
C.Increase the maximum number of training jobs
D.Enable early stopping for the tuning job
AnswerD

Early stopping halts poorly performing training trials once the objective metric stops improving, so AMT stops evaluating unpromising hyperparameter combinations. This directly addresses the slow tuning job and the wasted evaluations that never improve the objective metric.

Why this answer

Enabling early stopping in SageMaker Automatic Model Tuning (AMT) terminates poorly performing training jobs before they complete, which reduces wasted compute time and speeds up the tuning process. This is especially effective when using Bayesian optimization, as it allows the algorithm to focus on promising hyperparameter regions and avoid evaluating combinations that are unlikely to improve the objective metric.

Exam trap

The trap here is that candidates may confuse early stopping with reducing instance size or changing search strategies, not realizing that early stopping directly addresses wasted computation on poor trials without sacrificing search quality.

How to eliminate wrong answers

Option A is wrong because switching from Bayesian to random search would likely make the tuning less efficient, as random search does not use past results to guide future evaluations and often requires more trials to find optimal hyperparameters. Option B is wrong because using a smaller instance type for each training job reduces per-job compute capacity, which can slow down individual training runs and may not address the core issue of evaluating many unproductive combinations. Option C is wrong because increasing the maximum number of training jobs would evaluate even more hyperparameter combinations, prolonging the tuning job and potentially increasing wasted resources without improving efficiency.

257
MCQhard

A team uses SageMaker Pipelines to train and register a model. They want to conditionally run a hyperparameter tuning step only if the data quality check passes. Which pipeline step type should they use to branch the execution?

A.TuningStep
B.TrainingStep
C.ConditionStep
D.TransformStep
AnswerC

ConditionStep evaluates a Boolean condition against a pipeline property, such as the data quality check's output, and branches execution accordingly. It gates the tuning step so it runs only when the check passes, satisfying the conditional-execution requirement without altering the underlying training logic.

Why this answer

The ConditionStep allows comparing values and branching to different steps. If data quality passes, the tuning step runs; otherwise, the pipeline stops or runs an alternative step. Other steps do not provide conditional branching.

258
MCQeasy

Which SageMaker built-in algorithm is specifically designed for time series forecasting?

A.Image Classification
B.BlazingText
C.DeepAR
D.XGBoost
AnswerC

DeepAR is a supervised recurrent neural network algorithm built into Amazon SageMaker specifically for time series forecasting. It learns from many related series, producing probabilistic forecasts with quantiles, which matches the requirement for a purpose-built forecasting algorithm rather than generic regression or classification.

Why this answer

DeepAR is a SageMaker built-in algorithm based on recurrent neural networks (RNNs) specifically designed for time series forecasting. It learns from historical time series data and can predict future values, making it the correct choice for forecasting tasks. It supports both univariate and multivariate time series and can incorporate related time series.

Exam trap

MLA-C01 often tests the confusion between general-purpose algorithms (XGBoost) and purpose-built forecasting algorithms (DeepAR) — candidates pick XGBoost because it can be adapted for time series, but the question asks for one 'specifically designed' for forecasting.

How to eliminate wrong answers

Option A is wrong because Image Classification is a computer vision algorithm for categorizing images, unrelated to time series. Option B is wrong because BlazingText is a natural language processing algorithm for text classification and word embeddings. Option D is wrong because XGBoost is a general-purpose gradient boosting algorithm for classification and regression, not specifically designed for time series forecasting.

259
Multi-Selectmedium

A company has a SageMaker real-time endpoint that serves predictions. They want to set up automated monitoring and remediation for when the number of 5XX errors exceeds a threshold. Which TWO steps should they take? (Choose TWO.)

Select 2 answers
A.Use SageMaker Model Monitor to detect 5XX errors
B.Configure the CloudWatch Alarm to publish to an SNS topic
C.Set up a scheduled EventBridge rule to check 5XXError every minute
D.Write a custom script on EC2 to poll the endpoint and check for errors
E.Create a CloudWatch Alarm on the 5XXError metric
AnswersB, E

Remediation requires an action trigger, not just detection. Publishing the CloudWatch Alarm to an SNS topic satisfies the stem's automated remediation requirement by fanning out to subscribers such as Lambda or email, enabling an automated response once the 5XX threshold is breached.

Why this answer

A CloudWatch Alarm on the 5XXError metric can be configured to publish to an SNS topic, enabling automated notifications or remediation actions (e.g., via Lambda) when the alarm state is triggered. This is the standard AWS approach for alerting on endpoint errors without custom polling.

Exam trap

The trap here is that candidates confuse SageMaker Model Monitor (for data quality) with CloudWatch metrics (for operational health), leading them to select option A instead of recognizing that 5XX errors are operational metrics monitored via CloudWatch Alarms.

260
Multi-Selecthard

A machine learning team is deploying a model to a SageMaker endpoint and needs to implement A/B testing between two model versions. They want to split traffic 80/20 and monitor performance metrics for each variant. Which two actions should they take? (Choose two.)

Select 2 answers
A.Configure an Application Load Balancer to distribute traffic between two separate endpoints.
B.Use SageMaker Model Monitor to automatically compare the accuracy of the two variants.
C.Create a SageMaker endpoint configuration with two production variants, each specifying a different model and initial weight.
D.Create two separate SageMaker endpoints and use AWS Lambda to route requests based on a random number.
E.Enable data capture on the endpoint to log request and response data for analysis.
AnswersC, E

An endpoint configuration with multiple production variants allows you to deploy multiple models to a single endpoint. Each variant can have a different model and an initial weight that determines the traffic split. This is the foundational step for A/B testing, as it enables simultaneous serving of both model versions.

Why this answer

To perform A/B testing on SageMaker, you create an endpoint configuration with multiple production variants, each with a different model and weight to split traffic. Enabling data capture logs the requests and responses for each variant, allowing you to analyze performance metrics. Other options either do not provide native A/B testing or add unnecessary complexity.

Exam trap

The trap here is assuming that Model Monitor automatically compares variant performance, but it only monitors individual models for drift and quality.

261
MCQmedium

A company uses AWS Glue ETL jobs to clean and transform data from S3 before training. The data contains a column with 40% missing values. The column is normally distributed. Which imputation strategy should the data engineer use?

A.Impute with mode
B.Impute with mean
C.Impute with median
D.Drop rows with missing values
AnswerB

With a normally distributed column, the mean equals the central tendency, so mean imputation preserves the distribution's centre and avoids bias. At 40% missing, dropping rows would lose too much data, making mean the appropriate strategy.

Why this answer

For normally distributed data, imputing with the mean is the standard approach because it is efficient and unbiased. Option A (mode) is for categorical data and inappropriate for numeric data. Option C (median) is robust but less efficient for normal data.

Option D (drop rows) would lose 40% of data, which is excessive.

262
MCQeasy

A data scientist is preparing a CSV dataset in Amazon S3 for a SageMaker training job. Several rows contain missing values in numeric feature columns, and the chosen algorithm cannot handle NaNs. The scientist wants a repeatable, code-based transformation that runs inside a SageMaker Processing job before training. Which step is the MOST appropriate?

A.Use an S3 Lifecycle rule to expire objects containing missing values so only clean data remains in the bucket.
B.Open the CSV in a SageMaker notebook, manually edit the missing cells, and save the file back to S3.
C.Configure the SageMaker training job's input channel with a content type that tells the algorithm to ignore NaN values automatically.
D.Write a preprocessing script that uses pandas to impute or drop missing values and run it in a SageMaker Processing job with a scikit-learn container.
AnswerD

SageMaker Processing jobs run a containerized script against data in S3 and write outputs back to S3, which makes the transformation repeatable, versionable, and independent of notebook state. Using pandas for imputation or row removal inside the scikit-learn container is a standard, well-supported approach that produces a clean dataset the training job can consume directly.

Why this answer

Running a pandas-based cleaning script in a SageMaker Processing job makes the imputation or row removal reproducible and code-driven, and the job reads from and writes to S3 so the cleaned dataset is available to training. This is the canonical pattern for repeatable preprocessing that must run before a training job and be re-executed as data changes.

Exam trap

The trap here is treating missing-value handling as an algorithm setting or a storage-lifecycle concern, when it is really a preprocessing step that belongs in a scripted Processing job.

263
MCQmedium

A company needs to deploy a new model version to a SageMaker real-time endpoint. They want to route 5% of traffic to the new version initially to monitor for errors before full rollout. Which deployment strategy should they use?

A.Blue/green deployment
B.Shadow testing
C.Canary deployment with production variants
D.Multi-model endpoint
AnswerC

Canary deployment with production variants lets SageMaker split endpoint traffic, sending 5% to the new model variant while the old version serves the rest. This directly satisfies the requirement to monitor errors before full rollout.

Why this answer

A canary deployment with production variants allows you to route a specific percentage of traffic (e.g., 5%) to the new model version by adjusting the `InitialVariantWeight` parameter in the production variant configuration. This enables gradual traffic shifting while monitoring errors, and you can later increase the weight to 100% for full rollout. SageMaker real-time endpoints support this natively by hosting multiple model variants behind the same endpoint.

Exam trap

The trap here is that candidates confuse canary deployment with shadow testing, mistakenly thinking shadow testing also routes live user traffic, when in fact shadow testing only duplicates traffic for validation without affecting the user experience.

How to eliminate wrong answers

Option A is wrong because blue/green deployment switches all traffic from the old version to the new version at once, not a gradual 5% routing, which defeats the purpose of initial error monitoring. Option B is wrong because shadow testing sends a copy of live traffic to the new version but does not serve responses to users; it is used for validation without impacting production traffic, not for routing a percentage of user-facing traffic. Option D is wrong because a multi-model endpoint hosts multiple models on the same endpoint but does not provide traffic splitting or weighted routing between model versions; it is designed for cost efficiency with many models, not gradual rollout.

264
MCQmedium

A machine learning engineer observes that a SageMaker training job fails with the error shown in the exhibit. What is the most likely cause of the failure?

A.The SageMaker execution role does not have an IAM policy that grants read access to the S3 bucket containing the training data.
B.The training data is stored in an unsupported format like Parquet.
C.The training job is using an incorrect AWS Region for the S3 bucket.
D.The VPC configuration prevents the training job from reaching the S3 bucket.
AnswerA

SageMaker training jobs assume an execution role whose IAM policy must permit s3:GetObject on the training data bucket. Without that read permission, the job cannot download the dataset and fails immediately, matching the access-denied error in the exhibit.

Why this answer

The error shown in the exhibit is a standard SageMaker access-denied error, which occurs when the SageMaker execution role lacks the necessary IAM permissions to read the training data from the S3 bucket. SageMaker uses the execution role's IAM policy to determine access to S3 resources; without a policy granting s3:GetObject (and optionally s3:ListBucket) on the bucket and objects, the training job fails at the data-loading stage.

Exam trap

The trap here is that candidates often confuse an access-denied error with a VPC or region issue, but the specific error message 'AccessDenied' (or similar) directly points to an IAM permissions problem, not a network or configuration mismatch.

How to eliminate wrong answers

Option B is wrong because SageMaker supports Parquet and other columnar formats natively (e.g., via Spark or built-in algorithms like XGBoost with Parquet input), so an unsupported format is not the cause of this specific access-denied error. Option C is wrong because SageMaker training jobs can access S3 buckets in any region as long as the bucket policy and IAM permissions allow cross-region access; the error message does not indicate a region mismatch. Option D is wrong because a VPC configuration issue would typically produce a timeout or network connectivity error, not an explicit access-denied error; the error shown is an IAM permissions failure, not a network reachability problem.

265
MCQeasy

A machine learning engineer needs to run a one-time scoring job over 500 GB of data stored in Amazon S3 using a trained model, and the results must be written back to S3. There is no requirement for a persistent HTTPS endpoint. Which SageMaker feature should the engineer use?

A.Run a SageMaker Batch Transform job that reads input objects from S3 and writes inference output back to S3.
B.Create a real-time endpoint and send each S3 object as an HTTPS request through the InvokeEndpoint API.
C.Use a SageMaker Processing job with a custom script that loads the model and writes predictions to S3.
D.Deploy a serverless inference endpoint and invoke it repeatedly until all 500 GB has been processed.
AnswerA

Batch Transform is purpose-built for offline scoring of large datasets in S3. It provisions the needed compute, distributes the work across the input objects, and writes output back to a specified S3 location, then tears the resources down, matching the one-time job with no persistent endpoint.

Why this answer

Batch Transform exists precisely for offline, high-volume inference where inputs and outputs live in Amazon S3 and no persistent endpoint is required. It manages the compute fleet for the duration of the job, parallelizes across input data, writes results to S3, and then releases resources, which fits a one-time 500 GB scoring task.

Exam trap

The trap here is reaching for a real-time or serverless endpoint for bulk scoring, when those are optimized for interactive request/response rather than large offline datasets.

266
MCQeasy

A company requires that all SageMaker notebook instances be created within a private VPC without internet access. Which configuration step is mandatory?

A.Use a SageMaker Studio notebook instead.
B.Configure VPC settings when creating the notebook instance, choosing a private subnet.
C.Enable SageMaker direct internet access.
D.Assign a public IP to the notebook instance.
AnswerB

Selecting a private subnet during notebook instance creation places the instance's network interface in a VPC with no route to an internet gateway, satisfying the no-internet-access constraint. Direct internet access is then disabled by default, so traffic stays internal unless a NAT gateway or VPC endpoint is explicitly configured.

Why this answer

When creating a SageMaker notebook instance, you must explicitly configure the VPC settings and select a private subnet to ensure the instance is launched within a private VPC without internet access. This is mandatory because SageMaker notebook instances, by default, are created with internet access enabled unless you specify a VPC with no direct internet access. Selecting a private subnet ensures the instance uses only VPC endpoints or NAT gateways for outbound traffic, meeting the no-internet requirement.

Exam trap

The trap here is that candidates often assume SageMaker Studio notebooks are automatically private, but they also require explicit VPC configuration to restrict internet access, and the question specifically asks about notebook instances, not Studio.

How to eliminate wrong answers

Option A is wrong because using a SageMaker Studio notebook does not inherently enforce a private VPC without internet access; Studio notebooks also require VPC configuration and can have internet access if not properly restricted. Option C is wrong because enabling SageMaker direct internet access would explicitly allow the notebook instance to reach the internet, which contradicts the requirement. Option D is wrong because assigning a public IP to the notebook instance would provide direct internet access, violating the no-internet requirement.

267
MCQhard

A company is fine-tuning a large language model using LoRA on SageMaker. They want to reduce GPU memory usage during training. Which configuration change would help?

A.Use QLoRA (quantized LoRA) with 4-bit quantization
B.Enable gradient accumulation
C.Increase the sequence length
D.Increase the batch size
AnswerA

QLoRA quantises the frozen base model weights to 4-bit, so they occupy roughly a quarter of the memory that 16-bit weights require, while LoRA adapters remain trainable in higher precision. This directly satisfies the stem's constraint of reducing GPU memory during fine-tuning, with minimal accuracy loss.

Why this answer

QLoRA extends LoRA by quantizing the frozen base model weights to 4-bit (typically NF4) while keeping LoRA adapters in higher precision, dramatically reducing GPU memory required for fine-tuning. This lets you train larger models on smaller GPUs with minimal accuracy loss, directly addressing the goal of reducing memory usage. It is the standard memory-optimization technique for LoRA fine-tuning on SageMaker.

Exam trap

The trap is choosing gradient accumulation because it is a common memory-related technique — but MLA-C01 tests that only quantization (QLoRA) reduces the base model's memory footprint, while accumulation and batch/sequence changes affect activation memory differently.

How to eliminate wrong answers

Option B is wrong because gradient accumulation reduces effective batch size pressure by summing gradients over steps, but it does not reduce the memory footprint of model weights, activations, or optimizer states — it mainly helps simulate larger batches. Option C is wrong because increasing sequence length increases activation memory quadratically in attention, making memory usage worse, not better. Option D is wrong because increasing batch size increases activation and gradient memory, raising GPU memory consumption rather than lowering it.

268
MCQmedium

A machine learning engineer is building a pipeline to preprocess text data for a sentiment analysis model. The data consists of customer reviews. The engineer wants to convert the text into numerical features while preserving the semantic meaning of words. Which technique should be used?

A.One-hot encoding of each word
B.Bag-of-words with TF-IDF
C.Hashing vectorizer
D.Word embeddings (e.g., Word2Vec or GloVe)
AnswerD

Word embeddings map tokens to dense vectors whose geometric relationships encode semantic similarity, so reviews with comparable meaning produce comparable features. This preserves semantic meaning, unlike bag-of-words or TF-IDF, which treat terms as independent and lose contextual relationships.

Why this answer

Word embeddings (like Word2Vec or GloVe) are dense vector representations that capture semantic relationships between words based on their context in a large corpus. For sentiment analysis, preserving semantic meaning (e.g., 'good' and 'excellent' having similar vectors) is critical, and embeddings directly encode this, unlike sparse or count-based methods.

Exam trap

The trap here is that candidates often choose TF-IDF (Option B) because it is a common text preprocessing technique, but they overlook the explicit requirement to 'preserve semantic meaning,' which only dense embeddings can achieve.

How to eliminate wrong answers

Option A is wrong because one-hot encoding treats each word as an independent binary feature with no semantic similarity—vectors for 'good' and 'excellent' are orthogonal, losing all contextual meaning. Option B is wrong because bag-of-words with TF-IDF produces sparse, high-dimensional vectors based on word frequency and inverse document frequency, which ignore word order and context, failing to capture semantic relationships. Option C is wrong because a hashing vectorizer uses a hash function to map words to fixed-size indices, which can cause collisions and still produces sparse, frequency-based features without any semantic understanding.

269
MCQeasy

Refer to the exhibit. A user launches a SageMaker notebook instance with this lifecycle configuration. What happens?

A.The script runs every time the notebook is started
B.The script runs after each kernel reset
C.The script runs only on the first start
D.The script runs only when creating the instance
AnswerA

A lifecycle configuration script placed in the start-notebook lifecycle hook executes each time the instance transitions to running, including after stop and restart. This satisfies the scenario's launch behaviour, since the script is invoked on every start rather than only at initial creation.

Why this answer

The lifecycle configuration script is set to run 'on-start', which means it executes every time the SageMaker notebook instance transitions to the 'InService' state, including initial creation and subsequent starts. This is distinct from 'on-create', which runs only once during instance provisioning. Therefore, the script runs on every start, making option A correct.

Exam trap

The key trap in this question is the distinction between 'on-start' and 'on-create' lifecycle events in SageMaker. Candidates often assume that a lifecycle script runs only once or on kernel resets, but 'on-start' executes every time the notebook instance is started, including after stopping and starting.

How to eliminate wrong answers

Option B is wrong because lifecycle configuration scripts are tied to the instance lifecycle (start/stop/create), not to kernel events; kernel resets are internal to the Jupyter environment and do not trigger lifecycle scripts. Option C is wrong because the script runs on every start, not only on the first start; the 'on-start' event fires each time the instance is started after being stopped. Option D is wrong because the script runs on every start, not only during instance creation; the 'on-create' event would be used for creation-only execution.

270
MCQhard

A data scientist trains a binary classification model using SageMaker and obtains an AUC of 0.95 on the test set. However, the precision-recall curve shows low precision for high recall thresholds. The business requires a model that performs well on the minority class. Which metric should the team primarily optimize during hyperparameter tuning?

A.Accuracy
B.F1-score on the validation set
C.AUC (Area Under the ROC Curve)
D.Log loss
AnswerB

F1-score balances precision and recall, directly addressing the low-precision-at-high-recall problem and the minority-class requirement. AUC aggregates performance across all thresholds and can look strong despite poor minority-class precision, so tuning against F1 on the validation set targets the constraint the business actually cares about.

Why this answer

The F1-score is the harmonic mean of precision and recall, so it directly penalizes models that achieve high recall at the cost of low precision — exactly the failure mode described. Since the business cares about the minority class, optimizing F1 during hyperparameter tuning forces the model to balance both false positives and false negatives on that class. AUC-ROC can remain deceptively high (0.95) even when precision collapses at high recall because it aggregates performance across all thresholds and is dominated by the majority class.

Exam trap

MLA-C01 often tests the misconception that a high AUC-ROC means the model is good for imbalanced data, when in fact AUC can be high while precision at the required recall is unusable — the fix is to tune on F1 or PR-AUC, not AUC.

How to eliminate wrong answers

Option A is wrong because accuracy is misleading on imbalanced binary classification — a model predicting the majority class exclusively can score >95% accuracy while completely failing the minority class. Option C is wrong because AUC-ROC summarizes ranking quality across all thresholds and is insensitive to class imbalance; a high AUC does not guarantee usable precision at the operating threshold, which is precisely the symptom described. Option D is wrong because log loss measures probabilistic calibration, not thresholded classification performance on the minority class, so minimizing it does not directly optimize precision-recall trade-off.

271
MCQmedium

A data engineer is preparing a dataset for a time series forecasting model. The dataset contains a timestamp column and a target variable. The engineer wants to create additional features such as lag values and rolling averages. Which SageMaker Data Wrangler transform should be used to generate these time series features?

A.Use the 'Handle Outliers' transform to create lag features.
B.Use the 'Balance Data' transform to generate rolling averages.
C.Use the 'Time Series' transform to create lag and rolling window features.
D.Use the 'Featurize Text' transform to extract date parts.
AnswerC

SageMaker Data Wrangler includes a 'Time Series' transform that can generate lag features, rolling statistics (like mean, sum), and other time-based features from a timestamp column. This transform is specifically designed for time series data preparation, allowing the engineer to create the required lag values and rolling averages efficiently.

Why this answer

Time series feature engineering often involves creating lag features (past values) and rolling statistics (e.g., moving averages) to capture temporal patterns. In SageMaker Data Wrangler, the 'Time Series' transform provides a dedicated set of operations for these tasks, including lag, rolling window, and date part extraction. This transform simplifies the process and ensures correct handling of time order.

Other transforms like text featurization or outlier handling serve different purposes and cannot produce these features.

Exam trap

The trap here is assuming any transform that manipulates numerical data can create lag features; only the Time Series transform is designed for temporal feature engineering.

272
MCQmedium

A data science team deployed a model on Amazon SageMaker and enabled Model Monitor to detect data drift. After a week, they receive alerts indicating that the distribution of a key feature has shifted significantly. However, the model's accuracy on the recent production data remains high. Which action should the team take next?

A.Disable the data drift alert since accuracy is not affected.
B.Increase the sample size for monitoring to reduce false positives.
C.Retrain the model immediately because data drift always degrades performance.
D.Investigate the root cause of the drift as it may be benign or may lead to future degradation.
AnswerD

Drift alerts flag distributional shift, not performance loss. Since accuracy stays high, the shift may be benign, but monitoring exists precisely to catch silent degradation before it harms outcomes. Investigating the root cause satisfies the need to determine whether the drift is harmless or a precursor to future accuracy decline.

Why this answer

Data drift does not always immediately impact model accuracy; the drift may be benign (e.g., a shift in a non-predictive feature) or may indicate a precursor to future degradation. Amazon SageMaker Model Monitor detects distribution shifts using statistical tests like Kolmogorov-Smirnov or Chi-squared, but the team must investigate the root cause—such as changes in data collection, seasonal patterns, or upstream pipeline issues—before taking corrective action. Disabling alerts or retraining blindly could mask underlying problems or waste resources.

Exam trap

The trap here is that candidates assume data drift always implies model degradation, but the exam tests the understanding that drift can be benign and requires root-cause analysis before any action.

How to eliminate wrong answers

Option A is wrong because disabling the alert ignores the potential for future performance degradation; drift can be a leading indicator of model failure even if current accuracy is high. Option B is wrong because increasing the sample size may reduce variance but does not address the underlying cause of the drift; false positives are not the issue here since the drift is real. Option C is wrong because data drift does not always degrade performance; retraining immediately without investigation could introduce unnecessary cost and complexity, and may even harm the model if the drift is benign.

273
MCQeasy

A data science team deploys a regression model using Amazon SageMaker. After one week, the model's prediction accuracy drops significantly. The team needs to detect this degradation automatically and trigger retraining. Which AWS service should they use to monitor the model's performance over time and set up alerts?

A.AWS CloudWatch
B.Amazon SageMaker Model Monitor
C.Amazon Inspector
D.AWS Config
AnswerB

SageMaker Model Monitor continuously captures inference data and compares it against a baseline, detecting data drift and model quality degradation. It publishes metrics to CloudWatch, letting the team set alarms that automatically trigger retraining pipelines when accuracy falls below threshold.

Why this answer

Amazon SageMaker Model Monitor is the correct choice because it is purpose-built to continuously monitor machine learning models deployed on SageMaker endpoints for data drift, feature attribution drift, and prediction quality degradation. It automatically compares live inference data against a baseline, triggers alerts when performance drops, and can be configured to initiate retraining pipelines via AWS Lambda or Step Functions, directly addressing the need to detect accuracy degradation and trigger retraining.

Exam trap

The trap here is that candidates often confuse general-purpose monitoring services like CloudWatch with model-specific monitoring tools, overlooking that SageMaker Model Monitor provides built-in drift detection and retraining triggers tailored for ML models, whereas CloudWatch requires extensive custom scripting to achieve the same functionality.

How to eliminate wrong answers

Option A is wrong because AWS CloudWatch is a general-purpose monitoring service for metrics, logs, and alarms, but it lacks native capabilities to detect model-specific degradation like data drift or prediction accuracy drop without custom code and manual baseline setup. Option C is wrong because Amazon Inspector is a vulnerability management service that scans workloads for software vulnerabilities and unintended network exposure, not for monitoring ML model performance or triggering retraining. Option D is wrong because AWS Config is a service for evaluating, auditing, and recording changes to AWS resource configurations, not for monitoring model prediction accuracy or detecting performance degradation over time.

274
MCQmedium

A data scientist notices that a production model's accuracy has degraded over the past week. The training data distribution remains unchanged, but the relationship between features and the target has shifted. Which type of drift is occurring, and which monitoring approach should be used?

A.Bias drift; use SageMaker Clarify post-deployment bias monitoring
B.Data drift; use SageMaker Model Monitor data quality monitoring
C.Feature attribution drift; use SageMaker Clarify
D.Concept drift; use SageMaker Model Monitor model quality monitoring with ground truth labels
AnswerD

Concept drift is a change in the relationship between features and target while input distribution stays stable, so predictions degrade despite unchanged data. Model quality monitoring with ground truth labels detects this by comparing predicted against actual outcomes.

Why this answer

Concept drift occurs when the underlying relationship between features and target changes. Model quality monitoring (comparing predictions against ground truth) detects this. Data drift monitors feature distribution changes, which are not present here.

275
MCQhard

A company needs to detect bias in a pre-trained model before deployment. They want to compute metrics like disparate impact and equal opportunity difference. Which AWS service should they use?

A.SageMaker Clarify
B.Amazon Rekognition
C.SageMaker Model Monitor
D.SageMaker Debugger
AnswerA

SageMaker Clarify computes pre-training and post-training bias metrics, including disparate impact and equal opportunity difference, directly against a pre-trained model. It satisfies the stem's requirement to detect bias before deployment by running bias analysis on model predictions without retraining, unlike SageMaker Model Monitor, which only tracks drift on live endpoints.

Why this answer

SageMaker Clarify is purpose-built to detect bias in datasets and models, computing pre-training and post-training bias metrics such as disparate impact (DI) and equal opportunity difference (EOD). It integrates with SageMaker training and hosting, and can also generate explainability reports using SHAP values. Because the question asks specifically for bias metrics like DI and EOD, Clarify is the correct service.

Exam trap

MLA-C01 often tests the confusion between Clarify (bias/explainability) and Model Monitor (drift/quality), so candidates who see 'model' and pick Model Monitor miss the bias-specific requirement.

How to eliminate wrong answers

Option B is wrong because Amazon Rekognition is a computer-vision service for image and video analysis, not a bias-detection tool. Option C is wrong because SageMaker Model Monitor detects data drift and model quality degradation in production, not pre-deployment bias metrics. Option D is wrong because SageMaker Debugger focuses on training-job debugging (tensor inspection, vanishing gradients), not fairness or bias analysis.

276
MCQmedium

A data engineer needs to prepare a large dataset (10 TB) stored in Amazon S3 for a training job on SageMaker. The data is in CSV format, but the training algorithm expects Parquet for performance. The engineer must transform the data with minimal cost and without writing custom code. Which service should be used?

A.Use AWS Glue to create a crawler and ETL job that converts CSV to Parquet.
B.Use SageMaker Processing with a TensorFlow script to read CSV and write Parquet.
C.Use Amazon S3 Select to convert the data to Parquet during retrieval.
D.Use Amazon EMR with a Spark job to convert the files.
AnswerA

AWS Glue's crawler infers the CSV schema and its serverless Spark ETL writes Parquet directly to S3, satisfying the no-custom-code and minimal-cost constraints for the 10 TB conversion. Unlike Lambda, which caps runtime and memory, Glue scales to terabyte workloads without cluster management.

Why this answer

AWS Glue is the correct choice because it provides a serverless, pay-per-use ETL service that can automatically convert CSV to Parquet without writing custom code. The Glue crawler infers the schema, and the ETL job uses built-in transforms to efficiently handle 10 TB of data with minimal cost, as it only charges for the resources consumed during the job execution.

Exam trap

The trap here is that candidates often confuse Amazon S3 Select's ability to filter data with the ability to transform data formats, but S3 Select only returns filtered results in the original format and cannot perform format conversion like CSV to Parquet.

How to eliminate wrong answers

Option B is wrong because SageMaker Processing with a TensorFlow script requires writing custom code, which violates the 'without writing custom code' requirement. Option C is wrong because Amazon S3 Select only supports filtering data using SQL queries on CSV or JSON objects; it cannot convert data to Parquet format. Option D is wrong because Amazon EMR with a Spark job requires provisioning and managing a cluster, incurring higher costs and operational overhead compared to the serverless Glue approach.

277
Multi-Selectmedium

A data scientist needs to create a feature group in Amazon SageMaker Feature Store for real-time recommendations. Which TWO configurations are required? (Select TWO.)

Select 2 answers
A.Enable the offline store
B.Provide a feature description
C.Specify a record identifier feature
D.Set a time-to-live (TTL) for the records
E.Enable the online store
AnswersC, E

Every SageMaker Feature Store feature group requires a record identifier feature, which uniquely identifies each record and enables point-in-time lookups and upserts. Without it the feature group cannot be created, satisfying the stem's requirement for a mandatory configuration for real-time recommendations.

Why this answer

Online store must be enabled for real-time serving, and a record identifier is required to uniquely identify records. Offline store and feature description are optional; time-to-live is not a standard feature.

278
MCQhard

A data scientist is training a binary classifier on a highly imbalanced dataset (1:100 class ratio). The dataset contains 500,000 rows and 30 features. The data is stored in S3 in Parquet format. The data scientist wants to use SageMaker's built-in XGBoost algorithm. Which data preparation technique should the data scientist apply to best address the class imbalance without causing data leakage?

A.Undersample the majority class to create a balanced dataset, then split.
B.Use the scale_pos_weight parameter in XGBoost to assign higher weight to the minority class.
C.Oversample the minority class using SMOTE on the entire dataset before splitting into train/validation sets.
D.Randomly oversample the minority class by duplicating rows, then perform stratified train/test split.
AnswerB

scale_pos_weight multiplies the minority class's gradient contribution during XGBoost training, countering the 1:100 skew. Because it adjusts only the loss weighting rather than duplicating or synthesising rows, no information crosses between train and validation splits, avoiding the leakage that resampling before splitting would cause.

Why this answer

The scale_pos_weight parameter in XGBoost directly adjusts the loss function to penalize misclassifications of the minority class more heavily, effectively handling class imbalance without modifying the dataset. This avoids data leakage because the weighting is applied during training only, not during preprocessing, and does not involve any synthetic data generation or resampling that could inadvertently expose test information.

Exam trap

AWS often tests the misconception that resampling techniques (like SMOTE or random oversampling) are always safe, when in fact applying them before splitting introduces data leakage, whereas built-in parameters like scale_pos_weight avoid this pitfall.

How to eliminate wrong answers

Option A is wrong because undersampling the majority class reduces the dataset size significantly (from 500,000 rows to ~10,000 rows), discarding valuable information and potentially degrading model performance, and it does not inherently prevent data leakage if done before splitting. Option C is wrong because applying SMOTE on the entire dataset before splitting causes data leakage: synthetic samples generated from the full dataset can incorporate information from the test set, leading to overly optimistic validation metrics. Option D is wrong because randomly oversampling the minority class by duplicating rows before splitting can cause data leakage if duplicates of the same row appear in both training and validation sets, and it does not introduce new variance, leading to overfitting.

279
MCQmedium

A machine learning team deploys a fraud detection model on a SageMaker endpoint. The model's predictions are used in real-time. The team wants to monitor for data drift by comparing incoming data distributions against a baseline created from the training data. Which SageMaker capability should they use?

A.SageMaker Model Monitor - Model Quality Monitor
B.SageMaker Model Monitor - Data Quality Monitor
C.SageMaker Model Monitor - Feature Attribution Drift Monitor
D.SageMaker Model Monitor - Bias Drift Monitor
AnswerB

Data Quality Monitor compares incoming request distributions against a baseline computed from training data, detecting drift in feature values. This directly satisfies the requirement to monitor real-time endpoint traffic for data drift, unlike Model Quality Monitor, which tracks prediction accuracy against ground truth labels.

Why this answer

SageMaker Model Monitor's Data Quality Monitor is specifically designed to detect data drift by comparing the statistical distribution of incoming inference data against a baseline computed from the training dataset. This capability tracks metrics like mean, variance, and quantiles for each feature, alerting when significant deviations occur. For a fraud detection model requiring real-time monitoring of input distributions, this is the correct choice.

Exam trap

The trap here is that candidates often confuse 'data drift' (input distribution changes) with 'model quality drift' (prediction performance changes), leading them to select Model Quality Monitor instead of Data Quality Monitor.

How to eliminate wrong answers

Option A is wrong because Model Quality Monitor focuses on monitoring the model's predictive performance metrics (e.g., accuracy, precision, recall) against a baseline, not the distribution of input features. Option C is wrong because Feature Attribution Drift Monitor uses SHAP-based feature importance to detect shifts in how features contribute to predictions, not the raw data distributions themselves. Option D is wrong because Bias Drift Monitor tracks fairness metrics and bias over time, such as demographic parity or equal opportunity, which is unrelated to general data distribution drift.

280
MCQhard

An ML team uses SageMaker Pipelines to automate model retraining. They want to skip redundant training steps when input data has not changed. Which feature should they enable?

A.Pipeline caching
B.Pipeline variable expressions
C.Model registry approval
D.Step parallelism
AnswerA

Pipeline caching reuses a step's outputs when its inputs, code and parameters are unchanged, so the training step is skipped entirely rather than re-executed. This directly satisfies the stem's requirement to avoid redundant training when input data has not changed, saving compute time and cost.

Why this answer

SageMaker Pipelines caching stores the output of a step keyed by the step's inputs (code, data, hyperparameters, etc.). When a pipeline run executes and the inputs are unchanged, the cached output is reused and the step is skipped, avoiding redundant training. This directly addresses the requirement to skip training when input data hasn't changed.

Exam trap

MLA-C01 often tests the confusion between caching (skip unchanged steps) and parallelism (run steps concurrently) — candidates must match the optimization goal to the correct feature.

How to eliminate wrong answers

Option B is wrong because pipeline variable expressions are for parameterizing pipeline definitions (e.g., passing values between steps or at runtime), not for detecting unchanged inputs and skipping execution. Option C is wrong because model registry approval is a governance workflow for promoting models to production, unrelated to step execution optimization. Option D is wrong because step parallelism runs independent steps concurrently to reduce wall-clock time, but it does not skip steps whose inputs are unchanged.

281
MCQmedium

A company plans to deploy a large foundation model using SageMaker JumpStart. They are concerned about costs because the model will be used intermittently. Which deployment option is MOST cost-effective for intermittent traffic?

A.Purchase SageMaker Savings Plans for the endpoint
B.Deploy as a serverless endpoint
C.Use a batch transform job for each request
D.Deploy as a real-time endpoint with a multi-model endpoint
AnswerB

Serverless endpoints scale to zero when idle, so you pay only for inference requests rather than continuous instance hours. This directly satisfies the intermittent-traffic constraint, where a real-time endpoint would bill for provisioned capacity around the clock. Cold-start latency is the trade-off, but cost efficiency dominates for sporadic workloads.

Why this answer

Serverless endpoints in SageMaker automatically scale to zero when not in use, so you pay only for the compute time consumed during inference requests. This makes them the most cost-effective option for intermittent traffic, as you avoid paying for idle compute capacity.

Exam trap

The trap here is that candidates often confuse 'multi-model endpoints' with 'serverless' and assume they both scale to zero, but multi-model endpoints still run on provisioned instances that incur hourly costs regardless of traffic.

How to eliminate wrong answers

Option A is wrong because Savings Plans provide a discount on consistent usage but still require you to pay for a minimum baseline of compute, which is wasteful for intermittent traffic. Option C is wrong because batch transform jobs are designed for processing large datasets asynchronously, not for handling individual requests in real time, and they incur startup costs per job. Option D is wrong because a multi-model endpoint still runs on persistent instances that incur costs even when idle, and while it improves utilization across models, it does not eliminate idle costs for intermittent traffic.

282
MCQhard

Refer to the exhibit. The training job failed. What is the MOST likely cause?

A.The learning rate is too high
B.The instance type does not have SSD storage
C.The instance type does not have enough memory
D.The training data size exceeds the available EBS volume size
E.The number of epochs is too low
AnswerD

SageMaker training jobs download data to an attached EBS volume; if the dataset exceeds that volume's capacity, the job fails with a disk-space error. The exhibit's failure therefore points to training data size exceeding available EBS volume size.

Why this answer

The error message in the exhibit indicates an 'OSError: [Errno 28] No space left on device' during the training job. This occurs when the training data size exceeds the available EBS volume size attached to the SageMaker training instance. SageMaker uses EBS volumes for storing training data and intermediate outputs; if the dataset is larger than the provisioned EBS storage, the job fails with this specific disk-full error.

Exam trap

The trap here is that candidates confuse disk space errors with memory errors (Option C) or incorrectly attribute the failure to hyperparameters (Option A or E), when the specific 'No space left on device' error directly points to insufficient EBS volume size.

How to eliminate wrong answers

Option A is wrong because a high learning rate would cause divergence or NaN loss values, not a disk space error. Option B is wrong because SSD storage is not a requirement for SageMaker training instances; the error is about disk space, not storage type. Option C is wrong because insufficient memory would manifest as an out-of-memory (OOM) error or process kill, not a 'No space left on device' error.

Option E is wrong because a low number of epochs would result in underfitting or poor convergence, not a disk space error.

283
Multi-Selecthard

A team is deploying a model using SageMaker Pipelines. They have defined a pipeline with steps: preprocessing, training, evaluation, and conditional registration. The evaluation step produces a JSON file with metrics. If accuracy > 0.9, the model is registered; else, the pipeline fails. Which TWO statements about this pipeline are correct? (Choose TWO.)

Select 2 answers
A.The evaluation step must output a JSON file in a specific format to be used by the condition step.
B.The condition step can reference the accuracy value using a pipeline parameter or property file.
C.The conditional step should be implemented as a separate Lambda function called from the pipeline.
D.The pipeline will automatically retry the training step if the condition fails.
E.The model registration step should be placed before the condition step to ensure the model is always registered.
AnswersA, B

The condition step parses the evaluation step's output as JSON, so that file must follow SageMaker's expected property-file schema with metric names and values; otherwise the accuracy comparison cannot be evaluated and the branch fails.

Why this answer

Option A is correct because SageMaker Pipelines' ConditionStep consumes a JSON property file produced by a preceding step (typically via a PropertyFile output), and that file must follow the supported JSON structure so the condition can parse and evaluate the metric value. Option B is correct because the condition can reference the accuracy value either through a pipeline parameter (e.g., a ParameterString/ParameterFloat passed in) or through the step's property file using JsonGet (e.g., JsonGet(step_name=eval_step, property_file='evaluation.json', json_path='metrics.accuracy')), which is exactly how the threshold comparison is expressed. Option C is incorrect because conditional branching in SageMaker Pipelines is natively handled by the ConditionStep, not by invoking a separate Lambda function.

Option D is incorrect because a failing condition causes the pipeline to fail or take the defined branch; it does not trigger an automatic retry of the training step. Option E is incorrect because registering the model before evaluating the condition would defeat the purpose of gating registration on accuracy > 0.9.

Exam trap

The trap here is that candidates may confuse the built-in ConditionStep with a Lambda-based custom step, or assume that pipeline failure triggers automatic retries, when in fact SageMaker Pipelines requires explicit retry policies and does not retry on condition failures.

284
MCQmedium

A machine learning engineer is deploying a real-time inference endpoint on Amazon SageMaker AI for a fraud detection model. The model must serve predictions with consistent latency under 50 ms and the team expects traffic to fluctuate unpredictably, with occasional bursts. The engineer wants to automatically adjust the number of instances based on actual workload while minimizing cost during idle periods. Which SageMaker AI feature should the engineer configure?

A.SageMaker AI multi-model endpoints with a shared inference container
B.SageMaker AI automatic scaling with a target-tracking policy based on the SageMakerVariantInvocationsPerInstance metric
C.SageMaker AI Inference Recommender to select the optimal instance type and count
D.SageMaker AI Asynchronous Inference with an auto-scaling policy on queue depth
AnswerB

Target-tracking automatic scaling uses CloudWatch metrics like SageMakerVariantInvocationsPerInstance to adjust instance count dynamically, matching capacity to demand. This directly addresses unpredictable traffic and minimizes cost during low usage while maintaining latency. It is the native SageMaker AI scaling mechanism for real-time endpoints.

Why this answer

Automatic scaling with target-tracking on the invocations-per-instance metric is the correct approach because it continuously adjusts the number of instances to match actual traffic, ensuring latency targets are met while scaling down during idle periods to save cost. The other options either provide static recommendations or are designed for different inference patterns.

Exam trap

The trap here is confusing a one-time right-sizing recommendation from Inference Recommender with runtime automatic scaling that responds to live traffic.

285
MCQeasy

A data scientist is training a deep learning model on SageMaker and notices that the training loss oscillates and does not converge. They want to debug this issue. Which SageMaker feature can they use to monitor and analyze the training process?

A.SageMaker Profiler
B.SageMaker Gradient Descent optimization
C.SageMaker Debugger
D.SageMaker Automatic Model Tuning
AnswerC

SageMaker Debugger captures training tensors and metrics in real time, letting the data scientist inspect loss curves and detect vanishing gradients, exploding gradients or poor learning rates. This directly addresses the oscillating, non-converging loss by exposing per-step training telemetry rather than only post-training evaluation.

Why this answer

SageMaker Debugger is the correct feature because it provides real-time monitoring and analysis of training metrics, including loss values, gradients, and weights. It can automatically detect issues like oscillating or non-converging loss by setting rules (e.g., loss not decreasing) and emit alerts or capture tensors for later analysis, directly addressing the data scientist's need to debug training instability.

Exam trap

AWS often tests the distinction between monitoring training metrics (Debugger) versus optimizing hyperparameters (Automatic Model Tuning) or profiling system resources (Profiler), leading candidates to confuse Debugger with tuning or profiling features.

How to eliminate wrong answers

Option A is wrong because SageMaker Profiler is designed to analyze system-level performance (e.g., CPU/GPU utilization, I/O bottlenecks) and not training metrics like loss convergence. Option B is wrong because SageMaker does not offer a feature named 'Gradient Descent optimization'; gradient descent is an algorithm, not a SageMaker service, and this option represents a misconception that SageMaker provides a built-in optimizer for debugging. Option D is wrong because SageMaker Automatic Model Tuning (hyperparameter tuning) is used to find optimal hyperparameters, not to monitor or debug the training process in real time.

286
MCQmedium

A team has a large number of models that need to be deployed for batch inference weekly. They want to minimize cost and management overhead. Which approach is MOST efficient?

A.Use SageMaker Pipelines to run inference as part of the pipeline.
B.Use SageMaker Batch Transform with separate jobs for each model.
C.Create a single SageMaker endpoint for all models and update the model periodically.
D.Deploy each model to a separate SageMaker endpoint and delete after use.
AnswerB

SageMaker Batch Transform provisions compute only for the duration of each job, then tears it down, so weekly batch inference incurs no idle endpoint cost. Running separate jobs per model isolates each workload, satisfying the stem's cost-minimisation and low-management-overhead constraints without persistent infrastructure.

Why this answer

SageMaker Batch Transform is the most efficient approach for weekly batch inference because it automatically provisions and terminates compute resources for each job, minimizing cost and management overhead. Running separate jobs for each model allows independent scaling and avoids the complexity of managing persistent endpoints or multi-model hosting for batch workloads.

Exam trap

AWS often tests the distinction between batch and real-time inference, where candidates mistakenly choose persistent endpoints (Option C or D) for batch workloads, overlooking that Batch Transform is purpose-built for cost-efficient, ephemeral batch processing.

How to eliminate wrong answers

Option A is wrong because SageMaker Pipelines is an orchestration service for building and managing ML workflows, not optimized for running batch inference; using it for inference would add unnecessary complexity and cost without the automatic resource teardown of Batch Transform. Option C is wrong because a single endpoint for all models would require frequent model updates and cannot efficiently handle batch inference at scale, leading to idle costs and management overhead. Option D is wrong because deploying each model to a separate endpoint and deleting after use incurs significant provisioning delays and cost for endpoint creation/teardown, whereas Batch Transform handles this automatically with managed instances.

287
MCQmedium

A company wants to deploy a foundation model from SageMaker JumpStart with the lowest possible inference cost, given that latency requirements are flexible. They have a mix of traffic volumes. Which approach should they take?

A.Use SageMaker Savings Plans to get a discount on on-demand instances
B.Deploy the model on the largest GPU instance to handle peak load
C.Deploy the model on a serverless inference endpoint
D.Select the smallest instance type that meets throughput requirements and enable automatic scaling
AnswerD

Flexible latency permits the smallest viable instance, and automatic scaling matches capacity to fluctuating traffic volumes. Together these minimise inference cost while meeting throughput, since you pay only for the instances actually needed at each moment.

Why this answer

SageMaker JumpStart provides pre-built models; for cost optimization, choosing the smallest suitable instance type and enabling auto-scaling based on demand reduces cost while handling varying traffic.

288
MCQmedium

A company wants to deploy a machine learning model using infrastructure as code to ensure reproducibility. They need to define the SageMaker Studio domain, user profiles, and the endpoint configuration. Which tool should they use?

A.AWS CloudFormation or AWS CDK
B.SageMaker Pipelines
C.AWS Step Functions
D.SageMaker Studio
AnswerA

AWS CloudFormation and AWS CDK both declare SageMaker resources — Studio domains, user profiles, endpoint configurations — as versioned templates, so the same stack redeploys identically across environments. This satisfies the reproducibility constraint, unlike console-based provisioning or imperative scripts, which drift and cannot be diffed or rolled back.

Why this answer

AWS CloudFormation and AWS CDK are infrastructure-as-code (IaC) tools that allow you to define, provision, and manage AWS resources declaratively. For this use case, they can model the entire SageMaker Studio domain, user profiles, and endpoint configuration in templates or code, ensuring reproducibility and version control. This aligns directly with the requirement to deploy ML infrastructure as code.

Exam trap

The trap here is that candidates confuse SageMaker Pipelines (a CI/CD service for ML steps) with infrastructure-as-code tools, forgetting that Pipelines does not manage underlying infrastructure resources like Studio domains or endpoint configurations.

How to eliminate wrong answers

Option B (SageMaker Pipelines) is wrong because it is a purpose-built CI/CD service for ML workflows (training, tuning, batch transforms), not for defining and provisioning infrastructure resources like Studio domains or endpoints. Option C (AWS Step Functions) is wrong because it is a serverless workflow orchestration service for coordinating distributed applications and microservices, not for defining infrastructure resources declaratively. Option D (SageMaker Studio) is wrong because it is the web-based IDE for ML development, not a tool for defining or deploying infrastructure as code.

289
MCQmedium

An ML engineer creates a SageMaker inference pipeline with two containers: a preprocessor and a predictor. The preprocessor is a lightweight Python script that transforms input data. How should the engineer structure the endpoints to ensure both containers run sequentially?

A.Use batch transform with two transform jobs chained together.
B.Use an AWS Lambda function as a proxy to invoke the preprocessor and then the predictor separately.
C.Combine the preprocessor and predictor into a single Docker container.
D.Create a PipelineModel in SageMaker with both containers listed in order: first preprocessor, then predictor.
AnswerD

A PipelineModel chains containers so inference requests flow sequentially through each stage, satisfying the requirement that the preprocessor runs before the predictor. Listing them in order ensures the preprocessor's transformed output feeds the predictor within a single endpoint invocation.

Why this answer

SageMaker's PipelineModel allows you to define an ordered sequence of containers that are executed sequentially within a single HTTPS endpoint. When an inference request is made, the preprocessor container transforms the input, and the output is passed directly to the predictor container, all within the same endpoint invocation. This ensures low latency and tight coupling without needing external orchestration.

Exam trap

The trap here is that candidates often assume chaining containers requires external orchestration (like Lambda or separate jobs), but SageMaker's PipelineModel natively supports sequential container execution within a single endpoint, which is the simplest and most efficient approach.

How to eliminate wrong answers

Option A is wrong because Batch Transform runs separate transform jobs asynchronously, not as a real-time sequential pipeline; chaining two jobs introduces intermediate storage and latency, and is not designed for low-latency inference endpoints. Option B is wrong because using an AWS Lambda function as a proxy adds unnecessary network hops, cold-start latency, and complexity; SageMaker provides native PipelineModel support for sequential container execution without external orchestration. Option C is wrong because combining the preprocessor and predictor into a single container violates the separation of concerns and defeats the purpose of a modular inference pipeline; it also prevents independent scaling and updating of the preprocessing logic.

290
MCQeasy

Which SageMaker feature compiles a trained model into an optimized binary for a specific hardware target (e.g., Intel, ARM, NVIDIA, or edge devices) to improve inference performance?

A.SageMaker Model Monitor
B.SageMaker Neo
C.Amazon Elastic Inference
D.SageMaker Clarify
AnswerB

SageMaker Neo compiles trained models into optimised executables for specific hardware targets, including Intel, ARM, NVIDIA and edge devices. This directly satisfies the stem's requirement for a hardware-specific binary that improves inference performance, unlike training or hosting features that leave the model format unchanged.

Why this answer

SageMaker Neo is a model compilation service that optimizes models for specific hardware targets. Amazon Elastic Inference attaches GPU acceleration to endpoints, but does not compile models. Model Monitor monitors quality.

SageMaker Clarify explains predictions.

291
MCQmedium

A team is training a large language model using PyTorch on SageMaker. They need to reduce training time. The model has 10 billion parameters. Which distributed training strategy should they use?

A.Data parallelism with Horovod
B.Single GPU training
C.Use a larger instance type without parallelism
D.Model parallelism with SageMaker distributed
AnswerD

SageMaker distributed model parallelism partitions the 10-billion-parameter model across GPUs, holding each shard separately so the full weight set need not fit on one device. This suits the stem's large language model, where data parallelism would replicate all parameters per GPU and exhaust memory.

Why this answer

Model parallelism with SageMaker distributed is correct because a 10-billion-parameter model cannot fit into the memory of a single GPU, so the model itself must be partitioned across multiple GPUs/devices. SageMaker's distributed model parallelism library shards model layers and parameters across instances, enabling training of models that exceed single-device memory. This directly addresses the memory bottleneck that prevents scaling with data parallelism alone.

Exam trap

MLA-C01 often tests the misconception that adding more GPUs via data parallelism solves all scaling problems, when in fact model size exceeding single-device memory requires model parallelism.

How to eliminate wrong answers

Option A is wrong because Horovod data parallelism replicates the full model on every worker, so a 10B-parameter model will not fit in a single GPU's memory regardless of how many workers are added. Option B is wrong because single-GPU training cannot hold a 10B-parameter model in memory and offers no scaling path. Option C is wrong because simply choosing a larger instance type does not solve the fundamental memory constraint for very large models and does not provide the multi-device sharding needed for efficient training.

292
MCQhard

A machine learning engineer is configuring a SageMaker Processing job that runs a custom container to compute bias metrics on a dataset containing personally identifiable information. The job reads input data from one S3 bucket and writes reports to another, and the security team requires that the container has no outbound internet access and that the input and output buckets are reached without traversing the public internet. Which configuration satisfies these requirements?

A.Attach a VpcConfig with private subnets and a NAT gateway in the route table, and leave network isolation disabled so the container can reach S3.
B.Set NetworkConfig.EnableNetworkIsolation to true, attach a VpcConfig with private subnets, and create S3 gateway VPC endpoints that the subnet route tables reference.
C.Attach a VpcConfig with public subnets and an internet gateway, and set EnableNetworkIsolation to true so the container cannot use the gateway.
D.Set NetworkConfig.EnableNetworkIsolation to true and rely on the default SageMaker service-linked network path to reach S3.
AnswerB

Network isolation blocks the container from making outbound network calls, while the VPC configuration places the processing instances in private subnets. S3 gateway endpoints attached to those subnets' route tables let the job download inputs and upload reports over the AWS private network, so both the no-internet and no-public-traversal requirements are met without opening a NAT path.

Why this answer

Network isolation removes the container's ability to make outbound calls, and the VPC configuration places processing instances in subnets you control. S3 gateway endpoints on those subnets' route tables provide a private path for reading inputs and writing reports, satisfying both the no-internet and no-public-traversal constraints without a NAT gateway.

Exam trap

The trap here is assuming that enabling network isolation by itself secures S3 access, when private bucket reachability requires VPC endpoints on the processing subnets.

293
MCQmedium

A company deploys a model for fraud detection. They want to monitor if the model's predictions become less accurate over time due to changes in the underlying data distribution, but they do not have immediate access to ground truth labels. Which type of drift should they monitor as a proxy?

A.Feature attribution drift
B.Model quality drift
C.Data drift
D.Concept drift
AnswerC

Without ground truth labels, accuracy cannot be measured directly, so data drift is monitored as a proxy: it compares the live input feature distribution against the training baseline. A significant divergence signals that the model is operating on data unlike what it learned from, indicating likely degradation.

Why this answer

Data drift (option C) is the correct proxy to monitor when ground truth labels are unavailable because it detects changes in the input feature distribution over time. If the underlying data distribution shifts, the model's predictions are likely to become less accurate even if the relationship between features and labels remains stable. This allows teams to trigger retraining or investigation before model quality degrades.

Exam trap

AWS often tests the distinction between data drift and concept drift, and the trap here is that candidates confuse 'changes in data distribution' (data drift) with 'changes in the relationship between features and labels' (concept drift), assuming both require labels when only concept drift does.

How to eliminate wrong answers

Option A is wrong because feature attribution drift measures changes in the importance of features to the model's predictions, not shifts in the input data distribution itself, and it still requires some form of baseline comparison that may not directly indicate accuracy loss without labels. Option B is wrong because model quality drift requires access to ground truth labels to compute metrics like accuracy or F1-score, which the scenario explicitly states are unavailable. Option D is wrong because concept drift refers to changes in the underlying relationship between features and the target variable (the function mapping inputs to outputs), which cannot be detected without labels to compare predicted vs. actual outcomes.

294
MCQhard

A company is training a deep learning model with SageMaker and wants to reduce training time by using pipe mode instead of file mode for a large dataset stored as TFRecord files in Amazon S3. After switching the estimator's input mode to Pipe, the training job fails immediately with a dataset format error. The data scientist confirms the files are valid TFRecords and that the same script works with File mode. What is the most likely cause?

A.The S3 bucket and the training job are in different AWS Regions, so Pipe mode cannot stream the objects across Regions.
B.Pipe mode only supports RecordIO-encoded data for built-in algorithms, so TFRecord files must be converted to RecordIO before use.
C.The estimator is missing the enable_network_isolation parameter, which is required for Pipe mode to establish the streaming connection to S3.
D.The training script is still trying to read files from the local filesystem path instead of consuming the named pipe provided by SageMaker.
AnswerD

In Pipe mode, SageMaker streams channel data through a named pipe (FIFO) whose path is given by SM_CHANNEL_TRAIN, not as ordinary files. A script that calls standard file listing or opens a directory will fail because the pipe looks like a single stream. The script must read from the pipe sequentially, which explains why File mode worked and Pipe mode does not.

Why this answer

Pipe mode exposes training data as a named pipe rather than as files on disk. A script written for File mode typically lists files and opens them by path, which fails when SM_CHANNEL_TRAIN points to a FIFO. To use Pipe mode, the script must consume the stream sequentially, for example with a framework reader designed for pipes.

Format conversion and Region settings are unrelated to this error.

Exam trap

The trap here is assuming Pipe mode is a drop-in replacement for File mode, when scripts must be adapted to read from a stream instead of a directory.

295
MCQmedium

A company needs to serve real-time predictions from a large ensemble of three deep learning models, each requiring different inference environments (PyTorch, TensorFlow, MXNet). Which SageMaker endpoint type supports running multiple inference containers together?

A.Multi-model endpoint
B.Real-time endpoint with a single container
C.Multi-container endpoint
D.Asynchronous endpoint
AnswerC

Multi-container endpoints run up to fifteen containers together on one instance, letting each model use its own PyTorch, TensorFlow or MXNet environment. This satisfies the stem's need for multiple inference containers co-hosted, which single-model and serverless options cannot provide.

Why this answer

Amazon SageMaker multi-container endpoints allow you to run multiple inference containers (e.g., PyTorch, TensorFlow, MXNet) within a single endpoint, each handling different models or inference environments. This is achieved by deploying multiple containers behind a single endpoint with a serial or direct invocation pattern, enabling real-time predictions from the ensemble without managing separate endpoints.

Exam trap

The trap here is that candidates often confuse 'multi-model endpoint' (multiple models in one container) with 'multi-container endpoint' (multiple containers with different environments), leading them to incorrectly select Option A.

How to eliminate wrong answers

Option A is wrong because a multi-model endpoint hosts multiple models within a single container, not multiple containers with different inference environments; it uses a shared serving container and loads models dynamically from Amazon S3. Option B is wrong because a real-time endpoint with a single container can only run one inference environment, making it impossible to serve the three different deep learning frameworks required by the ensemble. Option D is wrong because an asynchronous endpoint is designed for large payloads and long processing times, not for real-time predictions, and it still uses a single container per endpoint.

296
MCQmedium

A company wants to build a customer service chatbot that answers questions about their internal policy documents. The documents are updated monthly, and the team cannot afford to retrain a model each time. Which approach is MOST appropriate?

A.Use a larger foundation model with a longer context window and paste all documents into each prompt
B.Use Retrieval-Augmented Generation (RAG) with the policy documents indexed in a vector store
C.Fine-tune a base LLM on the policy documents monthly
D.Train a custom model from scratch on the policy documents each month
AnswerB

RAG retrieves relevant passages from the indexed policy documents at query time and supplies them as context to the model, so monthly updates only require re-indexing the vector store rather than retraining. This satisfies the constraint that retraining is unaffordable.

Why this answer

RAG allows the LLM to retrieve relevant document sections at inference time, so knowledge stays current without retraining.

297
MCQeasy

An organization stores raw data in Amazon S3 as CSV files. They need to perform serverless data transformation and convert the data to Parquet format for efficient ML training. Which AWS service is most appropriate?

A.AWS Glue
B.Amazon EMR
C.Amazon Athena
D.Amazon Redshift
AnswerA

AWS Glue provides serverless Spark-based ETL that reads CSV from Amazon S3 and writes Parquet, satisfying the stem's transformation and format-conversion requirement. Its crawler and job model needs no cluster management, and Parquet's columnar layout accelerates downstream ML training.

Why this answer

AWS Glue is the most appropriate service because it is a fully managed, serverless ETL service designed specifically for data transformation tasks like converting CSV to Parquet. It automatically handles schema inference, data partitioning, and optimization for ML training workloads without requiring infrastructure management.

Exam trap

The trap here is that candidates often confuse Amazon Athena's ability to query Parquet data with the ability to transform data into Parquet, but Athena is a query engine, not an ETL transformation service.

How to eliminate wrong answers

Option B (Amazon EMR) is wrong because it requires provisioning and managing clusters, which contradicts the 'serverless' requirement; it is better suited for large-scale big data processing with frameworks like Spark or Hadoop, not simple serverless transformations. Option C (Amazon Athena) is wrong because it is an interactive query service for analyzing data directly in S3 using SQL, not a transformation engine; it cannot convert file formats like CSV to Parquet. Option D (Amazon Redshift) is wrong because it is a data warehouse for analytics and SQL-based querying, not a serverless transformation service; it requires loading data into a cluster and does not natively convert CSV to Parquet in S3.

298
MCQmedium

A team is building a fraud detection model using SageMaker and wants to detect anomalies in user login events. Which SageMaker built-in algorithm is specifically designed for anomaly detection in event-based data?

A.Factorisation Machines
B.IP Insights
C.Random Cut Forest
D.K-Means
AnswerB

IP Insights learns normal patterns of entity-to-IP address associations and flags unusual login events, making it the built-in algorithm designed for anomaly detection in event-based data. It differs from unsupervised outlier algorithms that operate on tabular feature vectors rather than entity-IP interaction history.

Why this answer

IP Insights is a SageMaker built-in algorithm specifically designed to learn patterns of IP address usage and detect anomalous behaviour in event-based data such as login events. It embeds IP addresses and entities (e.g., user IDs) into a vector space and flags deviations, making it ideal for fraud detection on login events. The other algorithms serve different purposes.

Exam trap

MLA-C01 often tests the distinction between IP Insights (IP/entity anomaly) and Random Cut Forest (general numeric anomaly), so candidates who see 'anomaly' and pick RCF miss the IP-event specificity.

How to eliminate wrong answers

Option A is wrong because Factorization Machines are used for recommendation and click-through-rate prediction, not anomaly detection in event data. Option C is wrong because Random Cut Forest is a general-purpose anomaly detection algorithm for numeric/tabular data, not specifically for IP-entity event patterns. Option D is wrong because K-Means is a clustering algorithm for grouping data, not for detecting anomalies in login events.

299
MCQmedium

A retail company has a SageMaker model that predicts customer churn. The model was trained on data that included a 'customer_zipcode' feature. After deployment, the data science team notices that the model's predictions for certain zip codes have become less accurate over time. They suspect that the relationship between zip code and churn has changed due to a recent relocation of a major employer. Which SageMaker monitoring capability should they use to detect this type of drift?

A.SageMaker Model Monitor bias drift monitoring
B.SageMaker Model Monitor model quality monitoring
C.SageMaker Model Monitor feature attribution drift monitoring
D.SageMaker Model Monitor data quality monitoring
AnswerB

Model quality monitoring evaluates the model's predictive performance against ground truth labels over time. It can detect concept drift by measuring metrics like accuracy or AUC and alerting when they degrade. Since the relationship between zip code and churn has changed, the model's predictions become less accurate, which model quality monitoring will catch.

Why this answer

Model quality monitoring is designed to monitor the performance of a model by comparing predictions to actual labels. When the relationship between a feature and the target changes, the model's accuracy drops, and model quality monitoring will detect this drift. Data quality monitoring only looks at input distributions, bias drift focuses on fairness, and feature attribution drift looks at feature importance, none of which directly measure predictive performance.

Exam trap

The trap here is assuming that any change in feature distribution or importance is equivalent to concept drift, when actually concept drift is a change in the underlying relationship that degrades model performance.

300
MCQmedium

A data scientist is preparing text data for sentiment analysis. They need to convert the text into numerical features while reducing the impact of common words. Which feature extraction method should they use?

A.Word2Vec embeddings
B.TF-IDF vectorization
C.Label encoding of each word
D.CountVectorizer with n-grams
AnswerB

TF-IDF vectorization weights each term by its frequency within a document, scaled inversely by how many documents contain it. This downweights common words such as "the" while emphasising distinctive terms, directly satisfying the requirement to reduce the impact of frequent words during numerical feature extraction.

Why this answer

TF-IDF (Term Frequency–Inverse Document Frequency) converts text to numerical features while down-weighting common words (like 'the', 'is') that appear across many documents, exactly matching the requirement to reduce the impact of common words. It is the standard feature extraction method for sentiment analysis when common-word influence must be minimized.

Exam trap

MLA-C01 often tests the confusion between CountVectorizer (raw counts) and TF-IDF (weighted counts) — candidates pick D, but only TF-IDF reduces the impact of common words via inverse document frequency.

How to eliminate wrong answers

Option A is wrong because Word2Vec produces dense embeddings that capture semantic similarity but do not inherently down-weight common words — frequency-based weighting is not part of Word2Vec. Option C is wrong because label encoding assigns arbitrary integers to words, implying false ordinal relationships and providing no frequency weighting. Option D is wrong because CountVectorizer with n-grams counts raw frequencies without inverse-document-frequency weighting, so common words still dominate.

Page 3

Page 4 of 9

Page 5

All pages