Courseiva

AWS Certified Machine Learning Engineer Associate MLA-C01 (MLA-C01) — Questions 76–150

665 questions total · 9pages · All types, answers revealed

Page 1

Page 2 of 9

Page 3
76
Multi-Selectmedium

A data scientist is using SageMaker to train a custom PyTorch model for image classification. They want to use SageMaker Debugger to detect training issues. Which TWO built-in rules are most relevant for detecting common training problems? (Select TWO.)

Select 2 answers
A.DataDistribution
B.Overfit
C.ExplodingGradients
D.ImageQuality
E.ConfusionMatrix
AnswersB, C

The Overfit rule in SageMaker Debugger monitors the validation loss relative to the training loss; if validation loss begins to increase while training loss continues to decrease, the rule emits a warning. This directly addresses the image classification scenario, where a custom PyTorch model can easily memorise training data rather than generalising, satisfying the stem’s requirement to detect common training problems.

Why this answer

Option B (Overfit) is correct because SageMaker Debugger's built-in Overfit rule monitors the gap between training and validation loss across steps and raises an issue when validation loss stops improving while training loss keeps decreasing, which is the classic signature of overfitting in a PyTorch image-classification job. Option C (ExplodingGradients) is correct because the ExplodingGradients rule inspects the gradients tensor emitted by the framework and flags abnormally large gradient values or spikes, which cause unstable or diverging training and are a common problem in deep networks. Option A (DataDistribution) is not a built-in Debugger rule for detecting training problems; it relates to SageMaker Clarify/Model Monitor data and bias analysis rather than Debugger's training-issue rule set.

Option D (ImageQuality) is not a SageMaker Debugger built-in rule; image-quality checks would be a custom preprocessing concern, not a Debugger rule. Option E (ConfusionMatrix) is not a Debugger training-issue rule; confusion matrices are evaluation artifacts computed after training (for example with SageMaker Clarify or custom code), not a built-in Debugger rule for detecting training problems.

Exam trap

The trap is selecting plausible-sounding but non-existent Debugger rules — candidates must know the actual built-in rule names (Overfit, ExplodingGradients, VanishingGradient, etc.) and not confuse them with evaluation metrics or data quality tools.

77
MCQeasy

A company has a SageMaker endpoint that serves a recommendation model. The security team wants to ensure that the model artifacts stored in Amazon S3 are encrypted at rest and that access to the S3 bucket is limited to the SageMaker execution role only. The team also wants to receive alerts if the bucket policy is changed. Which combination of actions should the machine learning engineer take?

A.Enable S3 Transfer Acceleration and restrict access using an IAM policy attached to the SageMaker execution role.
B.Enable S3 server access logging and use Amazon GuardDuty to monitor for unauthorized access to the bucket.
C.Use S3 Object Lock in compliance mode and attach a bucket policy that denies all principals except the SageMaker execution role.
D.Enable default encryption on the S3 bucket using AWS KMS, restrict the bucket policy to the SageMaker execution role, and enable AWS CloudTrail logging for S3 bucket policy changes with a CloudWatch alarm.
AnswerD

Enabling default encryption with AWS KMS ensures artifacts are encrypted at rest. Restricting the bucket policy to the SageMaker execution role enforces least privilege. CloudTrail logs S3 bucket policy changes, and a CloudWatch alarm on the relevant event pattern provides alerting when the policy is modified.

Why this answer

Encryption at rest is achieved with S3 default encryption using AWS KMS. Least privilege access is enforced by a bucket policy that allows only the SageMaker execution role. Alerting on bucket policy changes requires CloudTrail to log the event and a CloudWatch alarm to notify when the policy is modified.

Exam trap

The trap here is assuming that server access logging or GuardDuty provides encryption and policy-change alerting, when they only log access or detect threats.

78
MCQhard

A team is using SageMaker to run a large-scale distributed training job for a language model. They are using SageMaker's Pipe mode to stream data from S3 to reduce IO. They observe that the training throughput is lower than expected, and the CPU utilization is high while GPU utilization is low. The training script uses PyTorch's DataLoader with num_workers=0. The data preprocessing is minimal. Which change is most likely to improve GPU utilization?

A.Increase the number of data loading workers (num_workers).
B.Use a larger instance with more vCPUs.
C.Increase the number of GPUs per instance.
D.Switch from Pipe mode to File mode.
AnswerA

With num_workers=0, PyTorch loads and preprocesses batches synchronously on the main process, starving the GPU while CPUs work. Raising num_workers enables parallel background data loading, keeping the GPU fed and lifting utilisation during distributed training.

Why this answer

With num_workers=0, PyTorch's DataLoader loads data in the main training process, creating a CPU bottleneck that keeps GPUs idle. Increasing num_workers parallelizes data loading across multiple subprocesses, which reduces CPU strain and feeds data faster to GPUs, improving throughput. Adding more GPUs (Option C) or vCPUs (Option B) does not address the root cause, and switching to File mode (Option D) would increase I/O overhead, worsening performance.

79
MCQhard

A machine learning engineer is using Amazon SageMaker Debugger to monitor a training job for a deep neural network. They receive a rule alert indicating 'exploding gradients'. Which action should they take to address this issue?

A.Use a smaller batch size
B.Reduce the learning rate
C.Increase the number of layers to absorb gradients
D.Increase the learning rate
AnswerB

Exploding gradients arise when large updates compound across layers, so lowering the learning rate shrinks each step and restores stability. This directly addresses the alert SageMaker Debugger raised, though gradient clipping is an alternative remedy.

Why this answer

Exploding gradients occur when large error gradients accumulate during backpropagation, causing unstable updates and divergence. Reducing the learning rate directly scales down the parameter update step (Δθ = -η∇J), preventing the weights from overshooting and stabilizing training. This is the standard first-line remedy for exploding gradients in deep networks.

Exam trap

MLA-C01 often tests the confusion between exploding and vanishing gradients, where candidates might think increasing learning rate or adding layers helps, but the correct action is to reduce the learning rate or apply gradient clipping.

How to eliminate wrong answers

Option A is wrong because a smaller batch size increases gradient noise and can actually worsen instability, not fix exploding gradients. Option C is wrong because adding more layers deepens the network, which can exacerbate vanishing/exploding gradients due to repeated multiplicative Jacobians. Option D is wrong because increasing the learning rate amplifies the update magnitude, making exploding gradients worse and likely causing divergence.

80
Multi-Selectmedium

A data scientist is training a deep learning model using SageMaker and wants to use distributed training across multiple GPUs to reduce training time. Which TWO actions should the scientist take to configure distributed training? (Select TWO.)

Select 2 answers
A.Reduce the number of epochs to match the number of GPUs
B.Use the SageMaker distributed data parallelism library
C.Manually split the training data into shards and upload to S3
D.Configure the SageMaker estimator with a distribution parameter
E.Set the instance count to 1 with a multi-GPU instance
AnswersB, D

The SageMaker distributed data parallelism library shards each batch across GPUs and synchronises gradients, which is the mechanism that actually distributes training across multiple GPUs. It satisfies the requirement to configure distributed training rather than merely selecting an instance type.

Why this answer

Option B is correct because the SageMaker distributed data parallelism library (smdistributed.dataparallel) is purpose-built to shard data and synchronize gradients across multiple GPUs, reducing training time for deep learning jobs. Option D is correct because distributed training is enabled by passing a distribution parameter (e.g., {'smdistributed': {'dataparallel': {'enabled': True}}}) to the SageMaker estimator, which tells SageMaker how to launch the distributed framework. Option A is wrong because reducing epochs is unrelated to distributed training and would only degrade model convergence.

Option C is wrong because SageMaker's data parallelism library handles data sharding automatically, so manual sharding and S3 uploads are unnecessary. Option E is wrong because setting instance count to 1 with a multi-GPU instance does not configure distributed training across multiple instances and is not the required configuration action.

Exam trap

The trap here is that candidates confuse single-instance multi-GPU training (option E) with true distributed training across multiple instances, or assume manual data sharding (option C) is required when SageMaker automates it.

81
MCQmedium

A machine learning engineer is configuring auto-scaling for a SageMaker real-time endpoint. The endpoint is expected to have steady traffic during business hours and low traffic at night. The engineer wants to minimize costs by scaling in during low traffic, but the model container has a long start-up time (about 5 minutes). Which scaling policy should the engineer use to prevent request drops during sudden traffic spikes?

A.Use a step scaling policy based on invocations per minute with a step that adds two instances at a time.
B.Use a target tracking scaling policy based on average invocations per minute with a warm-up of 300 seconds.
C.Use a scheduled scaling action to add instances before business hours and remove them after.
D.Use a simple scaling policy based on average CPU utilization with a cooldown period of 5 minutes.
AnswerB

Target tracking on average invocations per minute reacts to demand, while the 300-second warm-up keeps newly added instances out of the metric until the slow-starting container is ready, preventing premature scale-in and request drops during sudden spikes.

Why this answer

Target tracking scaling policies in SageMaker automatically adjust capacity to maintain a target metric value, and the warm-up time of 300 seconds accounts for the 5-minute container start-up latency. This prevents request drops during sudden traffic spikes by ensuring new instances are fully initialized before they receive traffic, while still allowing the endpoint to scale in during low traffic to minimize costs.

Exam trap

The trap here is that candidates often choose a step scaling policy (Option A) because they think adding multiple instances at once handles spikes faster, but they overlook the critical need for a warm-up period to account for container start-up latency, which target tracking with warm-up explicitly addresses.

How to eliminate wrong answers

Option A is wrong because step scaling policies add instances in fixed increments (e.g., two at a time) without considering the long start-up time; this can lead to over-provisioning or under-provisioning during sudden spikes, and the lack of a warm-up period means new instances may not be ready to handle incoming requests, causing drops. Option C is wrong because scheduled scaling actions only handle predictable traffic patterns (e.g., business hours) and cannot react to sudden, unplanned traffic spikes, leaving the endpoint vulnerable to request drops. Option D is wrong because simple scaling policies based on average CPU utilization with a cooldown period of 5 minutes do not account for the model container's start-up latency; the cooldown prevents further scaling actions during the start-up period, but the policy itself cannot pre-warm instances, so traffic spikes during the cooldown can still cause request drops.

82
MCQeasy

A machine learning engineer needs to optimize a trained TensorFlow model for deployment on edge devices with limited compute. Which SageMaker feature should they use to compile the model for target hardware?

A.SageMaker Model Monitor
B.SageMaker Neo
C.SageMaker Debugger
D.SageMaker Elastic Inference
AnswerB

SageMaker Neo compiles trained models into optimised executables for specific target hardware, reducing compute and memory footprint on constrained edge devices. It satisfies the stem's requirement to compile a TensorFlow model for the target edge hardware without manual re-engineering.

Why this answer

SageMaker Neo is the correct choice because it is specifically designed to compile trained machine learning models into an optimized format for target hardware architectures, such as ARM, Intel, or NVIDIA, enabling efficient inference on edge devices with limited compute resources. It uses a compiler to apply hardware-specific optimizations like operator fusion and memory layout tuning, reducing latency and memory footprint without requiring manual code changes.

Exam trap

The trap here is that candidates confuse SageMaker Neo with SageMaker Elastic Inference, mistakenly thinking Elastic Inference compiles models for edge devices, when in fact Elastic Inference only accelerates cloud inference by attaching a fractional GPU and does not perform compilation or target edge hardware.

How to eliminate wrong answers

Option A is wrong because SageMaker Model Monitor is used for detecting data drift and model quality degradation in production, not for compiling or optimizing models for hardware. Option C is wrong because SageMaker Debugger is a tool for monitoring training jobs, capturing tensors and metrics to debug issues like vanishing gradients, not for post-training compilation or hardware-specific optimization. Option D is wrong because SageMaker Elastic Inference attaches a separate accelerator to an endpoint for low-cost GPU acceleration, but it does not compile or optimize the model for edge hardware; it is a runtime acceleration service for cloud inference, not for edge deployment.

83
Multi-Selecthard

A team is preparing text data for a natural language processing (NLP) model. They have a corpus of customer reviews. Which THREE preprocessing steps are essential to reduce noise and improve model performance?

Select 3 answers
A.Apply one-hot encoding to each word
B.Remove punctuation and special characters
C.Compute TF-IDF vectors
D.Perform stemming or lemmatization
E.Convert all text to lowercase
AnswersB, D, E

Removing punctuation and special characters strips non-linguistic tokens that inflate vocabulary size and dilute token frequency statistics, directly reducing noise in the customer review corpus. This normalisation step ensures the NLP model learns from meaningful word patterns rather than artefacts like commas, hashtags or emojis, satisfying the stem's noise-reduction requirement.

Why this answer

Option B is correct because removing punctuation and special characters eliminates non-linguistic symbols that add noise and are typically not useful features for NLP models, helping normalize the token stream. Option D is correct because stemming or lemmatization reduces inflected words to their base or dictionary form (e.g., 'running' to 'run'), decreasing vocabulary size and helping the model generalize across morphological variants. Option E is correct because converting all text to lowercase ensures that words like 'Review' and 'review' are treated as the same token, preventing spurious vocabulary duplication and improving consistency.

Option A is not a noise-reduction preprocessing step; one-hot encoding is a feature representation technique that actually increases dimensionality and does not clean the text. Option C is also not a preprocessing cleaning step; TF-IDF is a numerical vectorization/weighting method applied after text has already been normalized and tokenized.

Exam trap

AWS often tests the distinction between preprocessing steps (cleaning) and feature engineering steps (vectorization), so the trap here is that candidates mistake TF-IDF or one-hot encoding as essential preprocessing for noise reduction when they are actually downstream representation techniques.

84
MCQmedium

An ML engineer wants to use MLflow on SageMaker to track experiments and log metrics. They have set up MLflow on an EC2 instance. How can they best integrate MLflow tracking with SageMaker training jobs?

A.Install MLflow on the SageMaker notebook instance only
B.Use the SageMaker Experiments integration with MLflow
C.Set the MLFLOW_TRACKING_URI environment variable in the training job and use the mlflow library in the training script
D.Use SageMaker Processing to run MLflow after training
AnswerC

Setting MLFLOW_TRACKING_URI in the training job's environment points the mlflow client inside the training container at the EC2-hosted tracking server, so metrics and artefacts log remotely during training. This satisfies the requirement to integrate the existing MLflow setup without altering SageMaker's managed training infrastructure.

Why this answer

The ML engineer can set the `MLFLOW_TRACKING_URI` environment variable in the SageMaker training job definition and use the `mlflow` library inside the training script to log parameters, metrics, and artifacts directly to the MLflow tracking server running on the EC2 instance. This approach allows the training job to communicate with the external MLflow server over HTTP/HTTPS without requiring any additional SageMaker integrations.

Exam trap

A common misconception is that SageMaker Experiments is required for tracking with MLflow, but the correct approach is to directly configure the MLflow tracking URI and use the mlflow library in the training script, as SageMaker does not natively block outbound HTTP connections to an external MLflow server.

How to eliminate wrong answers

Option A is wrong because installing MLflow only on the SageMaker notebook instance does not enable tracking from SageMaker training jobs, which run on separate, ephemeral compute instances that do not have access to the notebook instance's local MLflow server. Option B is wrong because SageMaker Experiments is a separate tracking service that does not natively integrate with an external MLflow server; using it would require additional custom code to bridge the two systems, and it does not replace the need to set the tracking URI. Option D is wrong because SageMaker Processing is designed for data preprocessing and postprocessing, not for real-time metric logging during training; running MLflow after training would miss the ability to log metrics incrementally during the training run.

85
MCQmedium

A machine learning engineer is using Amazon SageMaker Data Wrangler to create a data preparation pipeline. The pipeline includes multiple transforms such as handling missing values, scaling, and encoding. The engineer wants to export the prepared data directly to a feature group in Amazon SageMaker Feature Store for reuse in training and inference. Which export option should the engineer choose?

A.Export to Amazon S3 as a CSV file.
B.Export to a Jupyter notebook for further processing.
C.Export to Amazon Redshift for analysis.
D.Export to a feature group in SageMaker Feature Store.
AnswerD

Exporting to a feature group writes the transformed dataset straight into SageMaker Feature Store, satisfying the requirement to reuse prepared features across both training and inference. Data Wrangler's native feature group export handles schema and ingestion automatically, avoiding intermediate Amazon S3 staging and manual ingestion code that other export targets would demand.

Why this answer

Amazon SageMaker Data Wrangler provides a built-in export destination for SageMaker Feature Store, allowing you to directly write the transformed data to a feature group without additional code. This enables seamless reuse of the prepared features for both training and real-time inference, leveraging the Feature Store's low-latency retrieval and versioning capabilities.

Exam trap

The trap here is that candidates may assume any export to Amazon S3 (Option A) is sufficient for reuse, but the question specifically requires export to a feature group, which is a distinct SageMaker Feature Store construct with its own schema, online/offline stores, and ingestion API—not just a file in S3.

How to eliminate wrong answers

Option A is wrong because exporting to Amazon S3 as a CSV file only stores the data as a flat file, not as a feature group, so it cannot be directly used with SageMaker Feature Store for online or offline inference without additional ingestion steps. Option B is wrong because exporting to a Jupyter notebook generates code for further processing but does not automatically persist the data to a feature group; it requires manual execution and additional configuration. Option C is wrong because exporting to Amazon Redshift is for analytical workloads and does not integrate with SageMaker Feature Store's feature group schema, online store, or low-latency serving for ML inference.

86
MCQmedium

A fraud detection team runs a SageMaker real-time endpoint that logs every request and response to an Amazon S3 bucket. Compliance requires that the model's prediction inputs and outputs be encrypted at rest with a customer-managed AWS KMS key, and that the endpoint be able to read the capture bucket only when necessary. The team enables data capture and specifies a KMS key on the endpoint configuration. Which additional configuration is required for the captured data written to S3 to be encrypted with that customer-managed key?

A.Enable default encryption on the target S3 bucket using SSE-KMS with the customer-managed key, and grant the SageMaker execution role kms:GenerateDataKey and kms:Decrypt permissions on that key.
B.Attach the AWS managed policy AmazonSageMakerFullAccess to the execution role and set the endpoint's KmsKeyId to the capture bucket's bucket key.
C.Create a VPC endpoint for S3 and configure the endpoint policy to require server-side encryption with the customer-managed key for all capture objects.
D.Configure an S3 bucket policy that denies PutObject unless the request includes the aws:kms header, and grant the SageMaker execution role kms:GenerateDataKey and kms:Decrypt.
AnswerA

SageMaker Data Capture writes objects to S3 using the endpoint's execution role. To have those objects encrypted with a customer-managed KMS key, the destination bucket must apply SSE-KMS default encryption with that key, and the role must hold kms:GenerateDataKey and kms:Decrypt on the key. Without both, captures are written with SSE-S3 or fail, so this combination satisfies the compliance requirement.

Why this answer

Captured data is written to S3 by the endpoint's execution role, so encryption with a customer-managed KMS key depends on the destination bucket applying SSE-KMS default encryption with that key and the role holding kms:GenerateDataKey and kms:Decrypt. A bucket policy or VPC endpoint can restrict access but does not encrypt capture objects. The endpoint's KmsKeyId protects attached volumes, not the S3 capture objects.

Exam trap

The trap here is assuming the endpoint's KmsKeyId setting encrypts captured S3 objects, when it only encrypts attached storage volumes and the capture destination's own SSE-KMS configuration governs object encryption.

87
Multi-Selectmedium

A data engineer needs to assess the quality of a dataset containing customer information. The dataset has missing values, outliers, and duplicate records. Which TWO AWS services can be used to perform data quality assessment? (Select TWO.)

Select 2 answers
A.Amazon Athena
B.AWS Glue DataBrew
C.AWS Glue ETL
D.Amazon QuickSight
E.Amazon SageMaker Data Wrangler
AnswersB, E

AWS Glue DataBrew profiles datasets and surfaces missing values, outliers and duplicates through built-in statistics and anomaly detection, without writing code. Its data quality rules then validate these conditions at scale, directly satisfying the stem's requirement to assess customer data quality across all three defect types.

Why this answer

AWS Glue DataBrew (B) is correct because it is a visual data preparation service that provides over 250 built-in transformations and profiling features, including data quality checks such as detecting missing values, outliers, and duplicate records via its profile view and anomaly detection. Amazon SageMaker Data Wrangler (E) is correct because it offers data quality and insights reports that automatically surface missing values, outliers, duplicate rows, and target leakage, allowing a data engineer to assess dataset quality before ML workflows. Amazon Athena (A) is a serverless query service for analyzing data in S3, not a dedicated data quality assessment tool.

AWS Glue ETL (C) is used for building and running extract-transform-load jobs rather than profiling data quality. Amazon QuickSight (D) is a business intelligence visualization service and does not provide built-in data quality assessment capabilities.

Exam trap

The trap here is that candidates often confuse AWS Glue ETL (option C) with AWS Glue DataBrew (option B), assuming the ETL service includes visual data quality assessment, when in fact DataBrew is the dedicated no-code data preparation and quality tool.

88
MCQeasy

A machine learning team needs to deploy a model that was built using scikit-learn. They want to use SageMaker for hosting. Which approach should they take?

A.Create a Jupyter notebook that loads the model and runs predictions on the SageMaker notebook instance
B.Create a custom Docker container with scikit-learn and deploy it on SageMaker
C.Launch a SageMaker training job with the model and use the training instance as an endpoint
D.Package the model artifacts and use the SageMaker built-in scikit-learn container for inference
AnswerD

SageMaker's built-in scikit-learn container already includes the framework and inference toolkit, so packaging model artifacts in the expected tar.gz format lets SageMaker host them directly. This avoids writing custom inference code or managing your own container, matching the requirement to use SageMaker hosting.

Why this answer

SageMaker provides a pre-built, optimized Docker container for scikit-learn that supports inference. By packaging the model artifacts (e.g., a joblib or pickle file) and deploying them using the built-in container, the team avoids the overhead of custom container creation while ensuring compatibility with SageMaker's hosting infrastructure, including automatic scaling and load balancing.

Exam trap

The trap here is that candidates often overcomplicate the solution by assuming a custom Docker container is always required for scikit-learn, overlooking the fact that SageMaker provides a fully managed, built-in container specifically for this framework.

How to eliminate wrong answers

Option A is wrong because a Jupyter notebook on a notebook instance is designed for interactive development and testing, not for production hosting; it lacks the necessary endpoint management, scaling, and availability features of SageMaker hosting. Option B is wrong because while a custom Docker container is a valid approach, it is unnecessary when SageMaker provides a built-in scikit-learn container that already includes the required dependencies and is optimized for inference, making this option over-engineered and more complex than needed. Option C is wrong because a SageMaker training job is ephemeral and intended for model training, not for serving inference requests; using a training instance as an endpoint is not supported, as training instances lack the persistent endpoint infrastructure (e.g., HTTPS endpoints, auto-scaling groups) required for production hosting.

89
MCQmedium

A startup wants to deploy a model that has variable traffic patterns, with some periods of no traffic and occasional spikes. They want to pay only for what they use and do not want to manage instances. Which SageMaker inference option should they choose?

A.Batch transform
B.Real-time endpoint with auto-scaling
C.Serverless inference
D.Multi-model endpoint
AnswerC

Serverless inference provisions compute automatically and scales to zero during idle periods, so the startup pays only per invocation and never manages instances. This matches the variable, spiky traffic pattern and the stated requirement to avoid instance management entirely.

Why this answer

Serverless inference is the correct choice because it automatically scales to zero during periods of no traffic and scales up to handle spikes, charging only for the compute time used. This eliminates the need to manage underlying instances, making it ideal for variable and intermittent traffic patterns.

Exam trap

The trap here is that candidates often confuse auto-scaling with the ability to scale to zero, but real-time endpoints with auto-scaling still maintain a minimum number of instances, incurring costs during idle periods, whereas serverless inference truly scales to zero.

How to eliminate wrong answers

Option A is wrong because batch transform is designed for offline, asynchronous predictions on large datasets, not for real-time or variable traffic patterns with occasional spikes. Option B is wrong because real-time endpoints with auto-scaling still require provisioning and managing underlying instances, and they cannot scale to zero, meaning you incur costs even during no traffic. Option D is wrong because multi-model endpoints reduce hosting costs by sharing instances across models but still require managing instances and cannot scale to zero, so you pay for idle capacity.

90
MCQeasy

A data scientist is preparing a dataset stored in Amazon S3 for a SageMaker training job. The dataset contains missing values in several columns. The scientist wants to impute missing values with the mean of each column. Which SageMaker built-in algorithm or processing method should be used to perform this imputation efficiently?

A.Use the SageMaker 'PCA' algorithm to fill missing values.
B.Use the SageMaker 'Linear Learner' algorithm to predict missing values.
C.Use the SageMaker 'Impute' transform in Data Wrangler.
D.Use the SageMaker 'XGBoost' algorithm to predict missing values.
AnswerC

SageMaker Data Wrangler offers an 'Impute' transform that allows replacing missing values with various strategies, including mean, median, or mode. This is a straightforward and efficient way to handle missing numerical data within the Data Wrangler interface, which can then be exported to a processing job or pipeline. It directly fulfills the requirement of mean imputation.

Why this answer

Mean imputation is a common preprocessing step for numerical features. In SageMaker Data Wrangler, the 'Impute' transform provides a user-friendly way to replace missing values with the column mean, median, or mode. This transform can be applied to multiple columns and is part of the data preparation flow.

It ensures that the dataset is complete before feeding it into a training algorithm, which is essential for algorithms that cannot handle missing values natively.

Exam trap

The trap here is assuming that built-in algorithms like XGBoost handle missing values, but the question specifically asks for mean imputation, which requires a dedicated preprocessing transform.

91
MCQmedium

A data scientist is using SageMaker Experiments to track multiple training runs. They want to compare different hyperparameter configurations and visualize the impact on model accuracy. What should they use to track hyperparameters?

A.SageMaker Debugger
B.SageMaker Autopilot
C.SageMaker Experiments
D.SageMaker Model Monitor
AnswerC

SageMaker Experiments records each training run as a trial, logging hyperparameters, metrics and artefacts so runs can be compared and charted. It directly satisfies the requirement to track hyperparameter configurations and visualise their effect on accuracy across multiple jobs.

Why this answer

SageMaker Experiments allows you to log hyperparameters as parameters. They can be viewed and compared across runs in the SageMaker Studio UI.

92
MCQeasy

A machine learning engineer at a retail company is monitoring a production model that predicts inventory demand. The model's prediction accuracy has dropped significantly over the past week. The engineer checks the model's input data and notices a new product category was introduced with a different distribution. Which concept is most likely causing the performance degradation?

A.Concept drift
B.Covariate shift
C.Data leakage
D.Model decay
AnswerB

Covariate shift occurs when input feature distributions change between training and production while the underlying relationship remains. The new product category alters the input distribution, so the model encounters patterns unlike its training data, degrading prediction accuracy.

Why this answer

B is correct because covariate shift occurs when the distribution of the input features changes while the relationship between features and the target remains the same. In this scenario, the introduction of a new product category with a different distribution alters the input data distribution, causing the model to encounter unseen patterns and degrade in prediction accuracy.

Exam trap

AWS often tests the distinction between covariate shift and concept drift, and the trap here is that candidates confuse a change in input distribution (covariate shift) with a change in the relationship between inputs and outputs (concept drift), leading them to incorrectly select concept drift.

How to eliminate wrong answers

Option A is wrong because concept drift refers to a change in the underlying relationship between input features and the target variable over time, not a change in the input distribution itself. Option C is wrong because data leakage involves the accidental inclusion of future information or target data in the training features, which is not indicated by a new product category with a different distribution. Option D is wrong because model decay is a general term for performance degradation over time, but it does not specifically describe the cause as a shift in input distribution; covariate shift is the precise technical concept here.

93
MCQmedium

A data engineer needs to prepare a dataset for a fraud detection model. The dataset contains a highly skewed numerical feature with extreme outliers. The engineer decides to apply a logarithmic transformation to this feature before training. Which SageMaker Data Wrangler transform should be used to apply the logarithmic transformation?

A.Use the 'Standard Scaler' transform in Data Wrangler.
B.Use the 'One-Hot Encoding' transform in Data Wrangler.
C.Use the 'Log Transform' transform in Data Wrangler.
D.Use the 'Min-Max Scaler' transform in Data Wrangler.
AnswerC

The 'Log Transform' transform in SageMaker Data Wrangler applies a natural logarithm (base e) to the selected numeric column. This is specifically designed to reduce right skewness and mitigate the impact of extreme outliers, making the feature more suitable for models that assume normality or are sensitive to scale. It directly addresses the scenario's need for a logarithmic transformation.

Why this answer

The logarithmic transformation is a common technique to reduce right skewness and stabilize variance in numerical data. In SageMaker Data Wrangler, the 'Log Transform' transform applies a natural logarithm to the selected column, effectively compressing the scale of large values and making the distribution more symmetric. This helps models that are sensitive to feature distributions, such as linear models, to perform better.

Other transforms like scaling or encoding do not change the distribution shape.

Exam trap

The trap here is confusing scaling transforms with distribution-changing transforms; scaling does not alter skewness or outliers.

94
MCQmedium

A team is using Amazon SageMaker for feature engineering. They have a dataset with a column 'TransactionDate' in string format (e.g., '2023-01-15 10:30:00'). They need to create features: year, month, day, hour, and day_of_week. What is the most efficient way to do this in a SageMaker processing job?

A.Use pandas datetime functions and then split
B.Use SageMaker built-in first party algorithms
C.Use AWS Glue for transformation
D.Use SQL query in Athena on S3 data
AnswerA

Pandas' `pd.to_datetime` parses the string column into datetime64 in one vectorised pass, then `.dt.year`, `.dt.month`, `.dt.day`, `.dt.hour` and `.dt.dayofweek` extract all five features directly. This satisfies the efficiency constraint by avoiding per-row Python loops or repeated parsing inside the SageMaker processing job.

Why this answer

Using pandas datetime functions within a SageMaker processing job is the most efficient approach for this task. SageMaker processing jobs run custom Python scripts, and pandas provides vectorized operations (e.g., `pd.to_datetime()`, `.dt.year`, `.dt.month`, `.dt.day`, `.dt.hour`, `.dt.dayofweek`) that parse the string column and extract all required features in a single pass without external dependencies or data movement.

Exam trap

AWS often tests the misconception that SageMaker built-in algorithms can handle feature engineering, but they are strictly for training and inference, not data preprocessing — the trap here is assuming 'first-party algorithms' include data transformation capabilities.

How to eliminate wrong answers

Option B is wrong because SageMaker built-in first-party algorithms (e.g., XGBoost, Linear Learner) are designed for model training, not for feature engineering or data transformation tasks like datetime parsing. Option C is wrong because AWS Glue is an ETL service that introduces additional overhead (e.g., Spark cluster startup, schema inference) and is less efficient for a simple in-memory pandas operation within a SageMaker processing job. Option D is wrong because using SQL in Athena on S3 data requires querying the raw data from S3, which incurs scan costs and latency, and Athena's SQL functions for datetime extraction (e.g., `EXTRACT`) are less flexible and slower than pandas for this specific transformation.

95
MCQmedium

A team is preparing text data for sentiment analysis. They have a large corpus of customer reviews. They want to convert the text into numerical features using a technique that captures word importance relative to the whole corpus. Which feature extraction method should they use?

A.Word2Vec embeddings
B.One-hot encoding
C.TF-IDF
D.Bag-of-words (CountVectorizer)
AnswerC

TF-IDF weights each term by its frequency within a review against its rarity across the whole corpus, so words that distinguish individual reviews score higher than common ones. That corpus-relative importance weighting is exactly what the scenario requires.

Why this answer

TF-IDF (Term Frequency–Inverse Document Frequency) weights each word by how often it appears in a document relative to how often it appears across the entire corpus, so common words like 'the' are down-weighted while distinctive words are up-weighted. This directly captures word importance relative to the whole corpus, which is exactly what the team needs for sentiment analysis feature extraction.

Exam trap

MLA-C01 often tests the distinction between frequency-based methods (CountVectorizer, TF-IDF) and embedding-based methods (Word2Vec) — candidates pick Word2Vec for 'importance' when the question specifically asks for corpus-relative term weighting.

How to eliminate wrong answers

Option A is wrong because Word2Vec produces dense embeddings that capture semantic similarity between words but do not weight terms by corpus-wide importance — it is a prediction-based embedding, not a corpus-relative weighting scheme. Option B is wrong because one-hot encoding represents each word as a sparse binary vector with no notion of importance or frequency, and it scales poorly with vocabulary size. Option D is wrong because bag-of-words (CountVectorizer) only counts term frequency within documents and ignores how common a term is across the corpus, so it cannot distinguish discriminative words from ubiquitous ones.

96
Multi-Selecthard

A company is building a real-time fraud detection system. They need to store features with historical context for model training and also support low-latency lookups for inference. Which THREE configurations should they set up in Amazon SageMaker Feature Store? (Select THREE.)

Select 3 answers
A.Enable point-in-time queries to retrieve historical feature values
B.Use the GetRecord API for real-time inference
C.Create a feature group with both online and offline store enabled
D.Use the BatchGetRecord API for all feature retrieval
E.Disable the offline store to reduce costs
AnswersA, B, C

Point-in-time queries are needed to reconstruct feature values at training time.

Why this answer

Point-in-time queries retrieve historical feature values at a specific time. Online store provides low-latency reads for inference. A feature group organizes features; creating one is necessary.

Offline store is for batch but not required for low-latency inference.

97
MCQhard

A machine learning team needs to ensure that all model training and inference jobs within SageMaker Studio run in a private network without internet access. The team also requires that inter-container traffic within the same training job be encrypted. Which configurations should they combine?

A.Configure SageMaker Studio in VPC-only mode and use KMS encryption
B.Use a VPC with a NAT gateway and enable network isolation
C.Enable inter-container traffic encryption and use a VPC with VPC endpoints
D.Enable network isolation mode and inter-container traffic encryption
AnswerD

Network isolation mode blocks all outbound internet and VPC traffic from the training container, satisfying the private-network requirement. Inter-container traffic encryption secures communication between containers within the same job, meeting the encryption constraint for distributed training.

Why this answer

To run SageMaker jobs in a private network without internet access, you enable network isolation mode, which prevents containers from accessing the internet. To encrypt inter-container traffic within the same training job, you enable inter-container traffic encryption. These two configurations together meet both requirements.

Exam trap

MLA-C01 often tests the confusion between VPC-only mode (which controls Studio access) and network isolation (which controls job containers), and candidates may overlook the need for inter-container encryption as a separate setting.

How to eliminate wrong answers

Option A is wrong because VPC-only mode for SageMaker Studio does not necessarily prevent internet access for training jobs, and KMS encryption is for data at rest, not inter-container traffic. Option B is wrong because a NAT gateway provides internet access, which contradicts the requirement for no internet access. Option C is wrong because using a VPC with VPC endpoints provides private access to AWS services but does not enforce network isolation or encrypt inter-container traffic; inter-container encryption is a separate setting.

98
MCQhard

A company uses Amazon SageMaker Ground Truth to label a dataset for object detection. To reduce labeling costs, they want to use active learning. Which configuration should they set up in Ground Truth?

A.Use a private workforce to label all data manually
B.Set the labeling job to random sampling of data
C.Configure the labeling job to use only bounding box annotations
D.Enable automated data labeling with a pre-trained model to select uncertain samples
AnswerD

Automated data labeling trains a model on already-labelled data, then selects the most uncertain samples for human review, cutting the number of manual annotations. This directly satisfies the goal of reducing labelling costs through active learning in Ground Truth.

Why this answer

Active learning in Amazon SageMaker Ground Truth reduces labeling costs by automatically selecting the most uncertain or informative data samples for human review, rather than labeling all data. Option D correctly configures automated data labeling with a pre-trained model to select uncertain samples, which is the core mechanism of active learning in Ground Truth.

Exam trap

The trap here is that candidates may confuse active learning with simply using a pre-trained model for inference (like option C's annotation type) or with random sampling (option B), missing that active learning specifically requires a feedback loop to select uncertain samples for human review.

How to eliminate wrong answers

Option A is wrong because using a private workforce to label all data manually does not implement active learning; it increases costs by labeling every sample without any automated selection. Option B is wrong because random sampling of data does not prioritize uncertain or informative samples; it treats all data equally, which defeats the purpose of active learning's cost-saving strategy. Option C is wrong because configuring the labeling job to use only bounding box annotations is a choice of annotation type, not an active learning configuration; it does not involve any automated selection or uncertainty sampling.

99
MCQmedium

A company is using SageMaker Autopilot to automatically build a regression model on a dataset. They want to understand which features are most important for the model's predictions. Which feature of Autopilot can provide this insight?

A.Autopilot candidate definition notebook
B.Autopilot model leaderboard
C.Autopilot data exploration report
D.Autopilot explainability report
AnswerD

The Autopilot explainability report uses SHAP values to quantify each feature's contribution to individual predictions, directly satisfying the requirement to identify which features most influence the regression model. It is generated automatically after candidate training, providing the feature-importance insight without manual analysis.

Why this answer

SageMaker Autopilot can generate explainability reports that include feature importance, either through SHAP or other methods, depending on the model type.

100
MCQhard

A company uses Amazon SageMaker Ground Truth to create a labeled dataset. They want to monitor the accuracy of human labelers during the labeling process. Which metric should they track?

A.Labeling job cost
B.Number of tasks completed
C.Accuracy against blinded ground truth
D.Task acceptance rate
AnswerC

Blinded ground truth compares each labeler's output against hidden known-correct labels, yielding a per-worker accuracy score. This directly satisfies the requirement to monitor labeler accuracy during labelling, unlike throughput or time-per-task metrics, which measure speed rather than correctness.

Why this answer

Ground Truth supports a built-in quality control mechanism where a percentage of tasks are sent to multiple workers, and one worker's answer is treated as the 'ground truth' (blinded to the others). By comparing each labeler's output against this blinded ground truth, the labeling job computes per-worker accuracy metrics, which is the direct measure of labeler quality. This is the metric designed specifically to monitor human labeler accuracy during the job.

Exam trap

MLA-C01 often tests the distinction between operational metrics (cost, throughput, acceptance rate) and quality metrics (accuracy vs. blinded ground truth) — candidates pick 'task acceptance rate' because it sounds like a quality signal, but it measures workflow behavior, not correctness.

How to eliminate wrong answers

Option A is wrong because labeling job cost is a financial/operational metric, not a measure of labeler correctness. Option B is wrong because the number of tasks completed measures throughput or productivity, not accuracy — a worker can complete many tasks with poor quality. Option D is wrong because task acceptance rate reflects how often a worker accepts or skips tasks, which is a workflow metric, not a correctness metric.

101
MCQeasy

A data science team deploys a PyTorch model on Amazon SageMaker for real-time inference. The model requires GPU for low latency. Which instance type is MOST cost-effective while meeting the GPU requirement?

A.ml.m5.2xlarge
B.ml.p4d.24xlarge
C.ml.p3.2xlarge
D.ml.c5.2xlarge
AnswerC

ml.p3.2xlarge pairs a single NVIDIA V100 GPU with the lowest cost among GPU-backed instances, satisfying the GPU constraint while avoiding the expense of multi-GPU types such as ml.p3.8xlarge. It delivers the required low-latency inference economically.

Why this answer

(ml.p3.2xlarge) is correct because it provides a GPU (NVIDIA V100) necessary for low-latency PyTorch inference on SageMaker, while being the most cost-effective among GPU options. The ml.p3.2xlarge offers a single GPU with sufficient compute for many real-time inference workloads, avoiding the higher cost of larger instances like ml.p4d.24xlarge.

Exam trap

The trap here is that candidates may assume any GPU instance is equally cost-effective, overlooking that ml.p4d.24xlarge is overprovisioned for typical inference, while CPU-only instances like ml.m5 and ml.c5 are tempting but fail the explicit GPU requirement.

How to eliminate wrong answers

Option A (ml.m5.2xlarge) is wrong because it is a general-purpose CPU instance with no GPU, failing to meet the GPU requirement for low-latency PyTorch inference. Option B (ml.p4d.24xlarge) is wrong because, while it provides powerful GPUs (NVIDIA A100), it is significantly more expensive than necessary for typical real-time inference, making it not the most cost-effective choice. Option D (ml.c5.2xlarge) is wrong because it is a compute-optimized CPU instance with no GPU, which cannot satisfy the GPU requirement for low-latency inference.

102
Multi-Selectmedium

A data engineer needs to perform feature selection on a dataset with 500 numeric features to train a regression model. The engineer wants to remove features that are redundant or have low predictive power. Which TWO techniques should the engineer consider? (Select TWO.)

Select 2 answers
A.Oversampling
B.Standardization
C.One-hot encoding
D.Correlation analysis
E.Lasso regularization (L1)
AnswersD, E

Correlation analysis identifies pairs of numeric features whose values move together, exposing redundancy so one of each highly correlated pair can be dropped. It directly satisfies the stem's requirement to remove redundant features, and with 500 numeric columns it scales cheaply, unlike wrapper methods that retrain the model per subset.

Why this answer

Correlation analysis (D) is correct because it identifies pairs of numeric features whose values move together, letting the engineer drop redundant, highly correlated predictors before training the regression model. Lasso regularization (E) is correct because its L1 penalty shrinks the coefficients of weak or irrelevant features exactly to zero, effectively performing embedded feature selection on the 500 numeric inputs. Oversampling (A) is a class-imbalance technique that duplicates or synthesizes minority-class samples and does nothing to remove redundant or low-predictive-power features.

Standardization (B) merely rescales features to comparable ranges and does not eliminate any feature. One-hot encoding (C) is a categorical-variable transformation that expands categories into binary columns and is irrelevant to a dataset of numeric features.

Exam trap

MLA-C01 often tests the difference between preprocessing (standardization, one-hot, oversampling) and feature selection (correlation, Lasso) — candidates pick standardization or one-hot thinking they reduce features, when they actually transform or expand them.

103
MCQeasy

A data scientist is working on a binary classification problem and wants to use AWS Glue for data preparation. The dataset has missing values in several numeric columns. Which imputation strategy is MOST appropriate for the scientist to apply in AWS Glue ETL?

A.Use the FillMissingValues transform to replace missing values with the mean of each column
B.Use a machine learning model to predict missing values
C.Drop all rows with missing values using the Drop transform
D.Set missing values to zero
AnswerA

FillMissingValues (or Imputer) with mean is a standard imputation strategy for numeric features.

Why this answer

AWS Glue ETL (PySpark) supports the `Imputer` transformer which can impute missing numeric values using the mean or median of the column. This is a built-in, straightforward approach.

104
Multi-Selectmedium

A data science team uses SageMaker to train and deploy models. They need to track model lineage, including datasets, training jobs, and model versions, to ensure reproducibility. Which THREE actions should they take? (Select THREE)

Select 3 answers
A.Enable SageMaker ML Lineage Tracking
B.Register all models in the SageMaker Model Registry
C.Store trained models in a public S3 bucket
D.Use SageMaker Experiments to organize training runs
E.Tag all resources with metadata such as project ID and training run ID
AnswersA, B, E

Lineage tracking automatically records artifacts, actions, and contexts.

Why this answer

A is correct because SageMaker ML Lineage Tracking automatically captures the relationships between datasets, training jobs, and model versions, creating a directed acyclic graph (DAG) of the ML workflow. This enables full reproducibility by allowing you to trace which data and code produced a specific model, without manual intervention.

Exam trap

The trap here is that candidates confuse SageMaker Experiments (which tracks metrics and parameters) with ML Lineage Tracking (which tracks the full provenance graph), leading them to select D instead of A, even though Experiments alone does not capture the inter-resource relationships needed for reproducibility.

105
MCQhard

A data engineer is using Amazon SageMaker Processing to run a data preprocessing script on a dataset with 500 million rows. The script runs out of memory on a single ml.r5.24xlarge instance. The engineer needs to modify the processing job to handle the dataset size. Which approach is most cost-effective and scalable?

A.Configure the Processing job with multiple instances and use ShardedByS3Key for data splitting.
B.Write the script to process data in chunks and write intermediate results to local ephemeral storage.
C.Increase the instance type to a larger one like ml.p3dn.24xlarge with more memory.
D.Reduce the number of instances to one and increase the volume size for swap space.
AnswerA

ShardedByS3Key distributes objects across multiple instances, so each processes a subset in parallel and memory per instance stays bounded. Scaling horizontally on smaller instances costs less than one oversized ml.r5.24xlarge and handles 500 million rows.

Why this answer

SageMaker Processing with ShardedByS3Key splits the input dataset by S3 object boundaries across multiple instances, allowing distributed processing of the 500 million rows without exceeding memory on any single instance. This approach is cost-effective as it uses multiple smaller instances (e.g., ml.r5.xlarge) rather than a single oversized instance, and scales linearly with data size.

Exam trap

AWS often tests the misconception that increasing instance size or using swap space is the primary solution for memory issues, whereas the correct approach is to distribute the workload horizontally using SageMaker's built-in data sharding feature.

How to eliminate wrong answers

Option B is wrong because writing intermediate results to local ephemeral storage does not solve the out-of-memory issue; the script still loads the entire dataset into memory before chunking, and local storage is limited and not designed for large-scale intermediate data. Option C is wrong because increasing to a larger instance like ml.p3dn.24xlarge (which has 192 GB memory vs. ml.r5.24xlarge's 768 GB) actually reduces memory, and GPU instances are not optimized for memory-intensive preprocessing; this approach is neither cost-effective nor scalable. Option D is wrong because reducing to a single instance and increasing volume size for swap space relies on disk-based swapping, which is orders of magnitude slower than RAM and will cause severe performance degradation or job failure due to I/O bottlenecks.

106
MCQmedium

A machine learning team needs to deploy a new model version for A/B testing, gradually shifting traffic from the old version to the new version over 24 hours. Which deployment strategy should they use?

A.Blue/green deployment
B.Shadow testing
C.Direct deployment with immediate full traffic
D.Canary deployment
AnswerD

Canary deployment routes a small percentage of live traffic to the new model version first, then incrementally increases that share as metrics stay healthy. This directly satisfies the stem's requirement to shift traffic gradually over 24 hours while limiting blast radius if the new version underperforms.

Why this answer

Canary deployment is the correct strategy because it allows gradual traffic shifting from the old model version to the new one over a specified time period (e.g., 24 hours) while monitoring for errors or performance degradation. This approach minimizes risk by exposing only a small percentage of users to the new version initially, then incrementally increasing traffic as confidence grows, which aligns perfectly with the A/B testing requirement.

Exam trap

AWS often tests the distinction between canary and blue/green deployment, where candidates mistakenly choose blue/green because both involve two versions, but blue/green is an instant switch, not a gradual traffic shift.

How to eliminate wrong answers

Option A is wrong because blue/green deployment involves switching all traffic from the old environment (blue) to the new environment (green) in a single cutover, not gradual traffic shifting over 24 hours. Option B is wrong because shadow testing runs the new model version in parallel with the old one but sends traffic only to the old version, comparing outputs offline without affecting live users, so it does not gradually shift traffic. Option C is wrong because direct deployment with immediate full traffic replaces the old version instantly, providing no gradual rollout or A/B testing capability.

107
MCQmedium

A data science team wants to track the lineage of models, including datasets, training jobs, and endpoints, for reproducibility and audit. They need a solution that captures relationships between artifacts automatically during training and deployment. Which service should they use?

A.Amazon S3 object versioning
B.SageMaker Experiments
C.SageMaker Model Registry
D.SageMaker ML Lineage Tracking
AnswerD

SageMaker ML Lineage Tracking automatically records relationships between datasets, training jobs, model artefacts, and endpoints as they are created, forming a queryable lineage graph. This satisfies the requirement for automatic capture during training and deployment for reproducibility and audit.

Why this answer

SageMaker ML Lineage Tracking automatically captures relationships among datasets, training jobs, model artifacts, and endpoints as the pipeline runs, producing a queryable lineage graph for reproducibility and audit. It records entities and associations without custom instrumentation, which is exactly what the team needs. SageMaker Experiments tracks runs and metrics but does not build the full artifact relationship graph.

Exam trap

MLA-C01 often tests the distinction between lineage tracking and model registry; candidates pick Model Registry because it sounds like governance, but it only catalogs models, not their data ancestry.

How to eliminate wrong answers

Option A is wrong because S3 object versioning only preserves object history; it has no concept of training jobs, endpoints, or relationships between artifacts. Option B is wrong because SageMaker Experiments organizes trials and metrics for comparison but does not automatically capture dataset-to-model-to-endpoint lineage. Option C is wrong because SageMaker Model Registry catalogs model versions and approval status, not the upstream dataset and training-job relationships required for full lineage.

108
MCQmedium

A data scientist wants to track hyperparameters, metrics, and artifacts for multiple training runs in SageMaker. They need to compare runs and identify the best performing model. Which SageMaker feature should they use?

A.SageMaker Model Monitor
B.SageMaker Debugger
C.SageMaker Autopilot
D.SageMaker Experiments
AnswerD

SageMaker Experiments groups training runs into experiments and trials, logging hyperparameters, metrics, and artefacts for each run. This enables side-by-side comparison of runs and identification of the best performing model, satisfying the tracking and comparison requirement.

Why this answer

SageMaker Experiments is the feature designed to track, organize, and compare machine learning training runs. It automatically captures hyperparameters, metrics, and artifacts for each run, allowing data scientists to analyze and identify the best performing model. It provides a centralized view of experiments and runs, facilitating reproducibility and collaboration.

Exam trap

The trap is confusing SageMaker Experiments with Debugger or Model Monitor; candidates often think Debugger tracks experiments, but it actually focuses on debugging training jobs, while Experiments is for tracking and comparing runs.

How to eliminate wrong answers

Option A is wrong because SageMaker Model Monitor detects drift and anomalies in deployed models, not tracking training runs. Option B is wrong because SageMaker Debugger provides real-time debugging of training jobs, such as tensor analysis, but does not manage experiment tracking. Option C is wrong because SageMaker Autopilot automates model building and tuning, but it does not provide a dedicated experiment tracking interface for comparing multiple runs.

109
MCQhard

A machine learning engineer is using Amazon SageMaker Processing to preprocess a large dataset. The processing job runs a custom Python script that uses the pandas library to read multiple CSV files from an S3 input prefix. The script must write the processed output to a different S3 prefix. Which configuration of the ProcessingInput and ProcessingOutput parameters is correct for this scenario?

A.ProcessingInput with source set to the S3 input prefix and destination set to the S3 output prefix; ProcessingOutput with source set to the S3 input prefix and destination set to '/opt/ml/processing/output'.
B.ProcessingInput with source set to '/opt/ml/processing/input' and destination set to '/opt/ml/processing/output'; ProcessingOutput with source set to the S3 input prefix and destination set to the S3 output prefix.
C.ProcessingInput with source set to the S3 input prefix and destination set to '/opt/ml/processing/input'; ProcessingOutput with source set to '/opt/ml/processing/output' and destination set to the S3 output prefix.
D.ProcessingInput with source set to '/opt/ml/processing/input' and destination set to the S3 input prefix; ProcessingOutput with source set to the S3 output prefix and destination set to '/opt/ml/processing/output'.
AnswerC

This configuration correctly maps the S3 input prefix to a local directory inside the processing container and the local output directory back to S3. The source for ProcessingInput is the S3 URI, and destination is the local path where the data will be available. For ProcessingOutput, source is the local path where the script writes outputs, and destination is the S3 URI. This is the standard pattern for SageMaker Processing jobs.

Why this answer

In SageMaker Processing, ProcessingInput specifies the S3 source and the local destination path where the data will be mounted in the container. ProcessingOutput specifies the local source path where the script writes output and the S3 destination where it will be uploaded. This bidirectional mapping ensures that the container can access input data and persist output data to S3.

The correct configuration uses S3 URIs for external locations and local paths for container locations.

Exam trap

The trap here is mixing up the source and destination fields for ProcessingInput and ProcessingOutput; source always refers to the origin (S3 for input, local for output) and destination to the target (local for input, S3 for output).

110
MCQmedium

During data preparation for a regression model, a data scientist notices that two features have a Pearson correlation coefficient of 0.95. The scientist is concerned about multicollinearity. Which action should be taken to address this issue?

A.Apply PCA to reduce dimensionality to a single component
B.Remove one of the two features
C.Keep both features as they are because linear models are robust to multicollinearity
D.Standardize both features using StandardScaler
AnswerB

A Pearson coefficient of 0.95 between two features signals strong linear redundancy, so their coefficients become unstable and interpretation unreliable. Removing one feature eliminates the collinearity directly, satisfying the stem's multicollinearity concern while retaining the remaining feature's predictive information.

Why this answer

Removing one of the highly correlated features reduces multicollinearity without losing much information, as they are nearly linearly dependent. Standardization does not fix multicollinearity. PCA would reduce dimensionality but may harm interpretability.

Keeping both can destabilize coefficient estimates.

111
MCQeasy

A company needs to perform time-series forecasting on historical sales data. Which SageMaker built-in algorithm is BEST suited for this task?

A.BlazingText
B.Linear Learner
C.XGBoost
D.DeepAR
AnswerD

DeepAR is a supervised recurrent neural network algorithm purpose-built for time-series forecasting, handling multiple related series and probabilistic predictions. It satisfies the stem's historical sales forecasting requirement, unlike classification, regression, or clustering algorithms that ignore temporal ordering.

Why this answer

DeepAR is a SageMaker built-in algorithm specifically designed for time-series forecasting using recurrent neural networks (RNNs). It is well-suited for predicting future values in a sequence, such as sales data, and can handle multiple related time series. The other algorithms are for classification, regression, or text, not time-series forecasting.

Exam trap

MLA-C01 often tests the distinction between general ML algorithms (Linear Learner, XGBoost) and purpose-built time-series algorithms (DeepAR) — candidates pick XGBoost because it can be used for regression, missing the specialized forecasting requirement.

How to eliminate wrong answers

Option A is wrong because BlazingText is a natural language processing algorithm for text classification and word embeddings, not time-series forecasting. Option B is wrong because Linear Learner is a general-purpose supervised learning algorithm for classification and regression on tabular data, not specialized for sequential time-series forecasting. Option C is wrong because XGBoost is a gradient-boosted decision tree algorithm for classification and regression, which can be used for time-series with feature engineering but is not purpose-built for forecasting like DeepAR.

112
MCQeasy

A team uses SageMaker Pipelines to automate model retraining. After a successful pipeline run, they want to register the new model version in the SageMaker Model Registry so that it can be reviewed for approval. Which step type should they add to the pipeline?

A.RegisterModelStep
B.ConditionStep
C.TransformStep
D.TrainingStep
AnswerA

RegisterModelStep packages the trained model artefacts and creates a new model version in the SageMaker Model Registry, enabling the approval workflow the team requires. It runs as a pipeline step, so registration happens automatically after each successful retraining run without manual intervention.

Why this answer

The correct step type is `RegisterModelStep`, which is specifically designed to create a new model version in the SageMaker Model Registry after a training or processing step completes. This step captures the model artifacts, training metrics, and metadata, and registers them under a specified model package group for approval workflows. Other step types serve different pipeline functions and do not interact with the Model Registry.

Exam trap

AWS often tests the distinction between executing a training job and registering the resulting model, leading candidates to mistakenly select `TrainingStep` when the question specifically asks about adding a model to the registry.

How to eliminate wrong answers

Option B is wrong because `ConditionStep` is used for branching logic within a pipeline (e.g., evaluating a metric to decide whether to proceed), not for registering a model. Option C is wrong because `TransformStep` performs batch inference on a dataset using a deployed model, not registration. Option D is wrong because `TrainingStep` runs a training job but does not automatically register the resulting model; a separate `RegisterModelStep` is required for registration.

113
MCQmedium

A gaming company uses a SageMaker endpoint for real-time player churn prediction. The model is updated weekly. After a recent retraining, the team notices that the endpoint's predicted probabilities for churn have shifted dramatically: the average predicted probability dropped from 0.3 to 0.05. The team suspects concept drift (the relationship between features and target changed) rather than data drift. They have SageMaker Model Monitor set up for data drift and quality metrics, but not for bias or explainability. The team needs to confirm concept drift and take corrective action. Which approach should the team take FIRST?

A.Configure SageMaker Model Monitor's model quality monitoring to compare predictions against actual outcomes collected from a week of production traffic
B.Immediately retrain the model using the most recent month of data and redeploy to the endpoint
C.Use Amazon SageMaker Clarify to compute SHAP values and understand which features are driving the new predictions
D.Investigate data drift by reviewing the Model Monitor feature distribution constraints and comparing recent input data to the baseline
AnswerA

Model quality monitoring compares live predictions with ground-truth outcomes, directly measuring whether the feature-target relationship has shifted. This confirms concept drift rather than merely detecting input distribution changes, which data drift monitoring already covers. It is the prerequisite step before retraining or recalibration.

Why this answer

To detect concept drift, the team needs to compare the model's predictions against actual observed outcomes (ground truth). SageMaker Model Monitor's quality monitoring can track prediction accuracy over time if ground truth is provided. Option A (Configure SageMaker Model Monitor's model quality monitoring) is the correct first step.

Option B (retrain with more recent data) might help but does not confirm drift. Option D (investigate data drift by reviewing feature distribution) checks feature distribution, not concept drift. Option C (use Clarify for SHAP values) is for feature importance, not drift detection.

114
Multi-Selectmedium

An ML team has deployed a model to a SageMaker real-time endpoint and wants to set up automated monitoring for model quality. Which TWO elements are required to configure SageMaker Model Monitor for model quality? (Select TWO.)

Select 2 answers
A.SHAP values for feature attribution
B.A constraints file with allowed deviation thresholds
C.A ground truth labels dataset for comparison
D.The endpoint's prediction output captured in real-time
E.A baseline statistics file derived from the training data
AnswersC, D

Ground truth labels are essential to compare against predictions and compute model quality metrics.

Why this answer

SageMaker Model Monitor for model quality requires a ground truth labels dataset to compare the model's predictions against actual outcomes. This comparison is essential for calculating quality metrics like accuracy, precision, recall, or F1 score, which indicate how well the model is performing over time.

Exam trap

The trap here is that candidates confuse the requirements for model quality monitoring (which needs ground truth labels and captured predictions) with those for data quality monitoring (which needs a baseline statistics file and constraints), leading them to select options B or E incorrectly.

115
MCQmedium

A data engineer must prepare a 4 TB Parquet dataset stored in Amazon S3 for a SageMaker training job that runs on 8 ml.p4d.24xlarge instances. The engineer wants the fastest possible data throughput during training while minimizing per-epoch I/O overhead. The dataset is immutable for the duration of the training run. Which approach BEST meets these requirements?

A.Copy the dataset once into an Amazon FSx for Lustre file system linked to the S3 bucket, and point the training job's channel at the FSx mount.
B.Use SageMaker FastFile mode, which streams objects from S3 on demand with a local cache and exposes a POSIX-like interface.
C.Use SageMaker File mode, which copies the dataset to the attached Amazon EBS volume before training begins.
D.Use SageMaker Pipe mode with the dataset converted to protobuf RecordIO and streamed directly to the training algorithm.
AnswerA

FSx for Lustre provides a high-throughput, low-latency POSIX file system that can lazily load data from the linked S3 bucket on first access. For a large immutable dataset read repeatedly across epochs, this eliminates repeated S3 GET overhead, gives shared high-bandwidth access to all training instances, and is the recommended pattern for accelerating SageMaker training on large datasets.

Why this answer

FSx for Lustre linked to S3 offers a shared, high-throughput file system that materializes data on first access and serves subsequent epochs at much lower latency than repeated S3 reads. Because the dataset is immutable during training, the lazy-load-then-cache behavior is safe and efficient, and all eight instances read from the same fast mount rather than each pulling from S3.

Exam trap

The trap here is assuming Pipe mode or FastFile mode always beats a file system for very large datasets, when in fact a linked FSx for Lustre mount removes repeated object-store latency for multi-epoch immutable data.

116
MCQmedium

A team uses SageMaker Model Monitor to track data quality. They notice that the monitor's constraint violations are increasing but the model performance remains good. What should they do?

A.Disable the monitor because it is not affecting performance.
B.Relax the constraint thresholds to reduce alerts.
C.Retrain the model using the latest data.
D.Investigate the specific features that are violating constraints to see if they are still relevant.
AnswerD

Investigating the specific violating features is right because Model Monitor's data quality constraints compare live traffic against the training baseline's statistical properties, not model accuracy. A drifting feature may be irrelevant to predictions, so examining which features breach constraints distinguishes harmless drift from genuine data issues, satisfying the need to act despite stable performance.

Why this answer

Increasing constraint violations in SageMaker Model Monitor do not necessarily indicate model degradation; they may reflect benign data drift where feature distributions shift but the model's predictive performance remains intact. Investigating specific violating features allows the team to determine whether the drift is meaningful (e.g., due to a real-world change that the model should adapt to) or irrelevant (e.g., a feature that is no longer used in the inference pipeline). This aligns with the monitoring best practice of separating data quality alerts from model performance metrics to avoid unnecessary retraining or threshold tuning.

Exam trap

The trap here is that candidates assume increasing constraint violations always mean the model is failing, leading them to choose retraining (Option C) or threshold relaxation (Option B), when the correct first step is to investigate the specific features to distinguish benign drift from harmful drift.

How to eliminate wrong answers

Option A is wrong because disabling the monitor eliminates visibility into data quality trends, which could mask future issues that do impact performance; Model Monitor is designed for proactive detection, not to be turned off when alerts are inconvenient. Option B is wrong because relaxing constraint thresholds without investigation may hide genuine data quality problems that could later degrade model performance, and it does not address the root cause of why violations are increasing. Option C is wrong because retraining the model on the latest data is premature and resource-intensive without first confirming that the drift is harmful; if the violating features are irrelevant, retraining wastes compute and may introduce unnecessary model churn.

117
MCQhard

An ML team uses AWS Step Functions to orchestrate a retraining pipeline triggered by EventBridge when new training data arrives. The pipeline includes a SageMaker training job and a model evaluation. If evaluation fails, the team wants to send an alert. How should they implement this?

A.Use SQS dead-letter queue for failed training jobs
B.Add a Catch rule in the Step Functions state machine to invoke a Lambda alert function
C.Configure SageMaker training job to publish to SNS on failure
D.Use EventBridge to monitor the training job status
AnswerB

A Catch rule on the evaluation state captures the failure and transitions to a Lambda task that sends the alert. This satisfies the stem's requirement to notify the team when evaluation fails, since Step Functions otherwise terminates the execution without invoking downstream alerting.

Why this answer

Step Functions supports error handling via Catch rules; a Catch on the training or evaluation task can transition to a Lambda function that sends an alert.

118
MCQmedium

An ML engineer monitors a SageMaker endpoint for data drift. They set up SageMaker Model Monitor to compare inference data against a baseline created from the training dataset. The monitoring schedule runs daily and reports violations. Which monitoring type should be configured to detect if the distribution of a numerical feature in real-time inference data differs significantly from the training distribution?

A.Data quality monitoring
B.Feature attribution drift monitoring
C.Bias drift monitoring
D.Model quality monitoring
AnswerA

Data quality monitoring compares the statistical distribution of features in captured inference data against a baseline built from the training dataset, detecting drift in numerical features. This matches the stem's requirement to flag significant distribution differences.

Why this answer

SageMaker Model Monitor's data quality monitoring detects feature distribution drift (statistical drift) between baseline and live data. Model quality monitoring requires ground truth labels, bias drift monitors fairness metrics, and feature attribution drift monitors SHAP values.

119
MCQhard

A company needs to deploy a large language model (LLM) on SageMaker with the Triton Inference Server to maximize GPU utilization and reduce latency. They have an NVIDIA A100 GPU. Which SageMaker inference option supports Triton?

A.SageMaker Batch Transform with Triton
B.SageMaker real-time endpoint using a Triton Inference Server container
C.SageMaker Serverless Inference with a custom container
D.SageMaker Neo compiled model on a CPU endpoint
AnswerB

Why this answer

SageMaker real-time endpoints support the Triton Inference Server through a pre-built container that integrates with NVIDIA A100 GPUs, enabling dynamic batching and concurrent model execution to maximize GPU utilization and reduce latency. Triton is designed for high-throughput inference on GPU hardware, making it the correct choice for this scenario.

Exam trap

The trap here is that candidates may confuse SageMaker Batch Transform with real-time endpoints, assuming Triton can be used for batch processing, but Triton is specifically designed for real-time, low-latency inference and is not supported in Batch Transform jobs.

How to eliminate wrong answers

Option A is wrong because SageMaker Batch Transform does not support the Triton Inference Server; it is designed for offline, asynchronous inference on large datasets without real-time GPU optimization features. Option C is wrong because SageMaker Serverless Inference does not support GPU instances or custom containers with Triton; it is limited to CPU-based inference and automatically managed scaling. Option D is wrong because SageMaker Neo compiles models for CPU or edge devices, not for GPU inference with Triton, and using a CPU endpoint would not leverage the A100 GPU's capabilities.

120
MCQeasy

A machine learning team at a retail company has deployed a product recommendation model using Amazon SageMaker. The model is updated weekly with new data. Recently, the team noticed that the model's accuracy on a holdout evaluation set has been declining over the past month. The data pipeline that feeds the training job has not changed. The team suspects data drift. They have SageMaker Model Monitor enabled on the inference endpoint and have set up Amazon CloudWatch metrics for feature distribution distances. Upon reviewing the CloudWatch dashboards, they see that the feature distribution distance metric for the most important feature 'product_category' has increased significantly. However, the team is unsure if this is the root cause. Which remediation step should the team take FIRST?

A.Retrain the model using the most recent week of data and redeploy to the endpoint
B.Investigate the data pipeline that feeds the training job to ensure consistent data collection and encoding of the 'product_category' feature
C.Rebuild the SageMaker endpoint with a different instance type to improve performance
D.Reduce the number of features in the model by removing 'product_category'
AnswerB

The rising distribution distance on product_category points to an input-side change. Before retraining or altering the model, verify the pipeline still collects and encodes that feature consistently, since encoding drift alone can inflate the metric without genuine concept drift.

Why this answer

The first step when data drift is suspected is to investigate the data pipeline to ensure consistent data collection and encoding. Since the model's accuracy is declining and the feature distribution distance for 'product_category' has increased, the root cause may be a change in how the feature is collected or encoded upstream, not necessarily a change in the underlying data distribution. SageMaker Model Monitor detects drift in feature distributions, but it cannot diagnose the cause; the team must verify the pipeline before retraining or modifying the model.

Exam trap

The trap here is that candidates assume data drift always requires retraining, but the first remediation step should always be to investigate the data pipeline to rule out upstream errors before taking corrective action on the model.

How to eliminate wrong answers

Option A is wrong because retraining with the most recent week of data assumes the drift is due to a natural shift in the data distribution, but if the drift is caused by a pipeline error (e.g., encoding change), retraining on corrupted data will not fix the issue and may degrade the model further. Option C is wrong because changing the instance type addresses compute performance, not data quality or model accuracy; it has no impact on feature distribution drift. Option D is wrong because removing the most important feature 'product_category' would likely reduce model accuracy further, and it does not address the underlying cause of the drift.

121
MCQmedium

A data science team uses SageMaker Pipelines for automated training. They need to conditionally register a model only if evaluation metrics exceed a threshold. Which pipeline step type should they use after the evaluation step?

A.Condition step
B.Processing step
C.Transform step
D.RegisterModel step
AnswerA

A condition step evaluates a JSON condition against the evaluation step's output and branches execution accordingly, so registration only proceeds when metrics exceed the threshold. This satisfies the requirement for conditional model registration within SageMaker Pipelines, unlike a processing or callback step.

Why this answer

The Condition step evaluates a condition and branches the pipeline; if the condition is met, the pipeline proceeds to register the model.

122
Multi-Selectmedium

A machine learning engineer is preparing a dataset for a binary classification model. The dataset has 10,000 rows and 200 features, with 5% positive class. The engineer suspects class imbalance may affect model performance. Which TWO actions should the engineer take to mitigate imbalance? (Choose 2.)

Select 2 answers
A.Perform PCA to reduce dimensions
B.Remove features with low variance
C.Use k-fold cross-validation
D.Apply SMOTE only to training data
E.Use class weights in the algorithm
AnswersD, E

SMOTE synthesises new minority-class observations by interpolating between nearest minority neighbours. Restricting it to the training split keeps those synthetic points out of validation and test data, preventing the leakage and inflated metrics that applying it before splitting would produce.

Why this answer

Option D is correct because SMOTE (Synthetic Minority Over-sampling Technique) generates synthetic samples of the minority class and must be applied only to the training data to avoid data leakage into validation/test folds, directly addressing the 5% positive-class imbalance. Option E is correct because setting class weights in the algorithm (e.g., class_weight='balanced' in scikit-learn) penalizes misclassification of the minority class more heavily, which mitigates imbalance without altering the dataset. Option A is incorrect because PCA is a dimensionality-reduction technique for the 200 features and does nothing to change the 5% class ratio.

Option B is incorrect because removing low-variance features is a feature-selection step unrelated to class distribution. Option C is incorrect because k-fold cross-validation is an evaluation/resampling strategy that provides more reliable performance estimates but does not itself rebalance classes.

Exam trap

The trap here is that candidates may confuse techniques for handling class imbalance with general data preprocessing or evaluation methods, leading them to select PCA or cross-validation as solutions, when in fact only resampling (SMOTE) and cost-sensitive learning (class weights) directly address the imbalance problem.

123
MCQmedium

An ML engineer is using Amazon SageMaker Automatic Model Tuning (AMT) to optimize hyperparameters for a gradient boosting model. The tuning job is taking a long time and has completed many training jobs. The engineer wants to stop training jobs that are unlikely to improve the objective metric. What should they configure?

A.Reduce the number of hyperparameter ranges
B.Use a random search strategy instead of Bayesian
C.Increase the maximum number of training jobs
D.Enable early stopping in the hyperparameter tuning job
AnswerD

Early stopping in automatic model tuning halts underperforming trials once their objective metric cannot beat the best completed trial, freeing capacity for promising configurations. This directly addresses the long-running tuning job by terminating trials unlikely to improve the objective.

Why this answer

Enabling early stopping in the Amazon SageMaker Automatic Model Tuning (AMT) job allows the tuning job to automatically stop training jobs that are unlikely to improve the objective metric based on intermediate results. This reduces the total time and compute cost by terminating poorly performing trials early, which directly addresses the engineer's goal of stopping unpromising training jobs.

Exam trap

The trap here is that candidates may confuse early stopping (which stops individual training jobs) with reducing the search space or changing the search strategy, which only affect the overall tuning job configuration without addressing the need to terminate underperforming trials mid-execution.

How to eliminate wrong answers

Option A is wrong because reducing the number of hyperparameter ranges limits the search space but does not actively stop ongoing training jobs that are underperforming; it only reduces the total number of possible trials. Option B is wrong because using a random search strategy instead of Bayesian does not provide any mechanism to stop individual training jobs early; random search simply samples hyperparameters randomly and runs each job to completion. Option C is wrong because increasing the maximum number of training jobs would allow more trials to run, which would increase the total time and cost, contrary to the goal of stopping unpromising jobs.

124
MCQmedium

A company trains a model daily using Amazon SageMaker and uses the model for real-time inference. They want to detect data drift between the training data and the inference data to decide when to retrain. Which AWS service should they use for this purpose?

A.Amazon Athena
B.Amazon SageMaker Model Monitor
C.AWS Glue
D.AWS Lambda
AnswerB

SageMaker Model Monitor compares live inference data against the training baseline and raises CloudWatch alerts when drift exceeds thresholds. That directly satisfies the requirement to detect data drift between training and inference data and decide when to retrain.

Why this answer

Amazon SageMaker Model Monitor is the correct service because it is specifically designed to continuously monitor machine learning models in production for data drift, feature attribution drift, and quality issues. It compares the distribution of live inference data against the baseline training data statistics and alerts when drift exceeds defined thresholds, enabling timely retraining decisions.

Exam trap

The trap here is that candidates may confuse AWS Glue's data cataloging and ETL capabilities with drift detection, or assume Athena's querying ability can be used for monitoring, but neither service is designed for continuous statistical comparison of ML inference data against training baselines.

How to eliminate wrong answers

Option A is wrong because Amazon Athena is an interactive query service for analyzing data in S3 using SQL, not a monitoring tool for ML model drift. Option C is wrong because AWS Glue is a serverless data integration and ETL service used for preparing and transforming data, not for detecting drift in production ML inference data. Option D is wrong because AWS Lambda is a serverless compute service for running code in response to events; while it could be used to trigger retraining, it does not natively perform drift detection or baseline comparison on inference data.

125
MCQeasy

A machine learning engineer wants to store, share, and manage features for multiple ML models across an organization. The features need to be accessible for both real-time inference (low-latency) and batch training. Which AWS service should the engineer use?

A.Amazon S3 with AWS Glue Data Catalog
B.Amazon Redshift
C.Amazon SageMaker Feature Store
D.Amazon DynamoDB
AnswerC

SageMaker Feature Store provides a centralised repository with an online store for low-latency real-time inference and an offline store for batch training, satisfying both access patterns. It also handles feature sharing and management across models and teams, matching the organisation-wide requirement.

Why this answer

Amazon SageMaker Feature Store is purpose-built for storing, sharing, and managing ML features across teams and models. It provides a unified feature store with both an online store (backed by Amazon DynamoDB or Redis) for low-latency real-time inference and an offline store (backed by Amazon S3) for batch training, directly addressing the requirement for dual access patterns.

Exam trap

The trap here is that candidates confuse a general-purpose database or data lake (like DynamoDB or S3) with a purpose-built ML feature store, overlooking the need for both low-latency online access and offline batch storage with feature-specific governance.

How to eliminate wrong answers

Option A is wrong because Amazon S3 with AWS Glue Data Catalog provides a data lake cataloging solution for batch analytics but lacks a low-latency online store for real-time inference, and it is not designed for feature-specific management like point-in-time consistency or feature sharing across ML models. Option B is wrong because Amazon Redshift is a data warehouse optimized for complex analytical queries on structured data, not for sub-millisecond real-time inference or feature store capabilities such as feature versioning and serving. Option D is wrong because Amazon DynamoDB is a NoSQL key-value database that can serve low-latency reads but does not natively support offline batch storage, feature sharing, or the unified online/offline store abstraction required for ML feature management.

126
MCQeasy

A data engineer needs to split a time-series dataset into training and validation sets for a forecasting model. Which split method should be used to avoid data leakage?

A.Use k-fold cross-validation with random shuffling.
B.Use feature importance scores to weight the splitting process.
C.Random split with 80% training and 20% validation.
D.Temporal split where training uses data up to a cutoff date and validation uses later data.
AnswerD

A temporal split preserves chronological order, training on earlier observations and validating on later ones, which mirrors real forecasting conditions. Random or stratified splits leak future information into training because adjacent time points correlate strongly. This satisfies the stem's requirement to avoid data leakage when validating a time-series forecasting model.

Why this answer

Time-series data has temporal dependencies, and random splits or k-fold cross-validation with shuffling would cause data leakage by allowing future information to influence training. A temporal split ensures that the model is trained only on past data and validated on future data, preserving the chronological order and preventing leakage.

Exam trap

For time-series data in AWS, using random splitting or cross-validation (e.g., with SageMaker) ignores temporal order and leads to data leakage. Always use a temporal split to preserve chronological order.

How to eliminate wrong answers

Option A is wrong because k-fold cross-validation with random shuffling randomly assigns data points to folds, which can place future data in the training set and past data in the validation set, leading to data leakage in time-series forecasting. Option B is wrong because feature importance scores are used for feature selection or model interpretation, not for splitting data, and they do not address the temporal ordering required to avoid leakage. Option C is wrong because a random split disregards the time order of the dataset, allowing the model to learn from future patterns during training and artificially inflate validation performance.

127
MCQeasy

A data scientist has a 40 GB CSV dataset in Amazon S3 that will be used to train a SageMaker model. The training script reads the data with pandas, and the scientist wants to reduce both storage cost and training-time I/O without changing the logical schema. Which data preparation action should be taken?

A.Move the dataset to Amazon EFS and mount it to the training container.
B.Convert the dataset to JSON Lines and enable S3 Transfer Acceleration.
C.Split the CSV into many smaller CSV files and keep the same column layout.
D.Convert the dataset to Parquet with Snappy compression and keep the same columns.
AnswerD

Parquet is a columnar format, so the training script reads only the columns it needs, and Snappy compression shrinks the on-disk footprint substantially compared with text CSV. The logical schema is preserved because column names and types remain the same. This directly reduces S3 storage cost and the bytes transferred during training, satisfying both requirements without changing the data model.

Why this answer

Parquet stores data by column and compresses each column efficiently, so a training script that needs a subset of columns reads far fewer bytes than it would from CSV. Snappy compression further reduces the S3 object size, lowering storage cost. Because the column names and types stay the same, the logical schema is unchanged and the pandas-based script can read the Parquet files with minimal modification.

Exam trap

The trap here is assuming that splitting or relocating files reduces cost, when only changing the storage format and compression actually shrinks the bytes stored and read.

128
MCQhard

A team is deploying a model that requires GPU acceleration for inference. They are using an Amazon SageMaker real-time endpoint. The model is a large language model (LLM) that does not fit on a single GPU. Which configuration should they use to minimize latency while fitting the model?

A.Use data parallelism with Horovod to distribute inference across GPUs.
B.Use SageMaker's model parallelism library to shard the model across multiple GPUs in a single instance.
C.Optimize the model with SageMaker Neo to reduce its size.
D.Deploy the model across multiple endpoints and use a load balancer.
AnswerB

Model parallelism shards a single model's layers across multiple GPUs within one instance, so the LLM fits in aggregate memory. Keeping the shards on one instance avoids cross-node network hops, satisfying the low-latency constraint that multi-instance distribution would violate.

Why this answer

SageMaker's model parallelism library allows you to shard a large language model across multiple GPUs within a single instance, enabling inference for models that exceed a single GPU's memory. This approach minimizes latency by keeping all GPUs in a single instance with high-speed interconnects (e.g., NVLink), avoiding the network overhead of distributing across separate instances.

Exam trap

The trap here is that candidates confuse data parallelism (which replicates the model) with model parallelism (which shards the model), assuming any distributed approach works for large models, but only model parallelism solves the 'does not fit on a single GPU' constraint.

How to eliminate wrong answers

Option A is wrong because data parallelism (e.g., Horovod) replicates the entire model on each GPU, which does not solve the problem of a model that does not fit on a single GPU; it requires the model to fit entirely on each GPU. Option C is wrong because SageMaker Neo optimizes models for target hardware through quantization and compiler optimizations, but it does not reduce the model's memory footprint enough to fit an LLM that exceeds a single GPU's capacity; Neo is for inference acceleration, not model sharding. Option D is wrong because deploying across multiple endpoints with a load balancer distributes requests but does not address the fundamental issue of a model that cannot fit on a single GPU; each endpoint would still need to host the full model, which is impossible without sharding.

129
MCQeasy

Which SageMaker feature allows you to automatically tune hyperparameters using Bayesian optimization?

A.SageMaker Autopilot
B.SageMaker Experiments
C.SageMaker Debugger
D.SageMaker Automatic Model Tuning
AnswerD

SageMaker Automatic Model Tuning runs hyperparameter tuning jobs that search parameter ranges using Bayesian optimisation as the default strategy, selecting configurations that improve the objective metric, which is precisely the automated tuning capability the question asks for.

Why this answer

SageMaker Automatic Model Tuning (AMT) is the feature that automatically searches for the best hyperparameters using strategies like Bayesian optimization. It runs multiple training jobs with different hyperparameter combinations and evaluates them against a chosen objective metric to find the optimal set. This reduces the manual effort of tuning and improves model performance.

Exam trap

The trap is confusing Automatic Model Tuning with Autopilot; candidates often think Autopilot is the tuning tool, but Autopilot is a broader AutoML feature, while AMT specifically focuses on hyperparameter optimization.

How to eliminate wrong answers

Option A is wrong because SageMaker Autopilot automates the entire model building process, including feature engineering and algorithm selection, but it uses AMT internally and is not solely focused on hyperparameter tuning. Option B is wrong because SageMaker Experiments tracks and compares runs but does not perform tuning. Option C is wrong because SageMaker Debugger monitors and debugs training jobs, but it does not tune hyperparameters.

130
MCQeasy

A data engineer wants to transform a categorical feature with 1,000 possible values into numerical features for a linear model. Which feature engineering technique is most appropriate for this high-cardinality feature?

A.One-hot encoding
B.Ordinal encoding
C.Target encoding
D.Label encoding
AnswerC

Target encoding replaces each of the 1,000 categories with a statistic derived from the target, producing a single numerical column rather than 1,000 sparse dummy columns. This directly addresses the high-cardinality constraint while remaining suitable for a linear model.

Why this answer

Target encoding replaces each category with the mean of the target variable for that category, which handles high cardinality without exploding dimensionality. One-hot encoding creates 1,000 columns, which is problematic for linear models.

131
MCQeasy

A company wants to ensure that only authorized users and services can invoke a SageMaker real-time endpoint. Which AWS service can be used to manage access control?

A.Amazon CloudWatch
B.AWS Identity and Access Management (IAM)
C.AWS CloudTrail
D.AWS Config
AnswerB

IAM controls authentication and authorisation for SageMaker endpoint invocation through identity-based policies and resource policies. It satisfies the requirement to restrict endpoint access to authorised users and services, unlike network-level controls that do not manage identity.

Why this answer

AWS Identity and Access Management (IAM) is the correct service because it allows you to create fine-grained permissions policies that control which users, roles, or services can invoke a SageMaker real-time endpoint via the InvokeEndpoint API. By attaching IAM policies to principals (e.g., IAM users, roles, or federated identities), you can restrict invocation based on conditions such as source IP, VPC endpoint, or MFA, ensuring only authorized entities can send inference requests.

Exam trap

The trap here is that candidates confuse monitoring or auditing services (CloudWatch, CloudTrail, Config) with access control, mistakenly thinking they can restrict API calls when they only observe or log them.

How to eliminate wrong answers

Option A is wrong because Amazon CloudWatch is a monitoring and observability service that collects metrics, logs, and alarms; it does not manage access control or authentication for API calls. Option C is wrong because AWS CloudTrail is an audit service that records API activity for governance and compliance; it logs who invoked an endpoint but cannot enforce or deny access. Option D is wrong because AWS Config is a resource inventory and compliance service that evaluates configuration rules; it can detect non-compliant endpoint policies but cannot directly control invocation permissions.

132
MCQmedium

A data team is using Amazon SageMaker Data Wrangler to prepare a dataset. They need to detect potential bias in the data before training a model. Which feature of Data Wrangler should they use?

A.Visual data profiling
B.Built-in transforms for missing values
C.Export to Feature Store
D.Bias detection with Amazon SageMaker Clarify
AnswerD

SageMaker Clarify bias detection integrates into Data Wrangler, computing metrics such as class imbalance and disparate impact across facets before training. This satisfies the requirement to detect potential bias in the prepared dataset prior to model training.

Why this answer

Data Wrangler integrates with Amazon SageMaker Clarify for bias detection. Therefore, the correct answer is D: Bias detection with Amazon SageMaker Clarify. Options A and B are features of Data Wrangler but not specifically for bias detection.

Option C (Export to Feature Store) is unrelated to bias detection.

133
MCQmedium

A team deploys a PyTorch model on Amazon SageMaker for real-time inference. They notice that inference latency is higher than expected. They suspect the serialization format used for input data is inefficient. Which approach would MOST likely reduce latency?

A.Use Amazon SageMaker Batch Transform instead of real-time inference.
B.Change the input serialization format to Protocol Buffers.
C.Enable automatic scaling on the endpoint.
D.Increase the instance type to a compute-optimized instance.
AnswerB

Protocol Buffers encode payloads as compact binary, cutting serialisation and parsing overhead compared with JSON or CSV text. Smaller request bodies also reduce network transfer time. This directly targets the inefficient input serialisation the team suspects is inflating real-time inference latency on SageMaker.

Why this answer

Protocol Buffers (protobuf) are a binary serialization format that is significantly more compact and faster to parse than text-based formats like JSON or CSV. By reducing the size of the input data and the CPU overhead of deserialization, switching to protobuf directly addresses the root cause of high inference latency on SageMaker real-time endpoints.

Exam trap

The trap here is that candidates often confuse throughput improvements (scaling, larger instances) with latency reduction, or mistakenly think Batch Transform can substitute for real-time inference, when the question specifically targets the serialization format as the suspected bottleneck.

How to eliminate wrong answers

Option A is wrong because Batch Transform is designed for offline, asynchronous processing of large datasets and does not reduce latency for real-time inference; it actually increases end-to-end time by batching. Option C is wrong because automatic scaling adjusts the number of instances to handle traffic volume, not the per-request latency caused by serialization inefficiency. Option D is wrong while a compute-optimized instance might improve raw processing speed, it does not fix the underlying serialization bottleneck and is a more expensive, indirect solution compared to changing the serialization format.

134
MCQhard

A machine learning engineer is using SageMaker Automatic Model Tuning to optimize hyperparameters for a regression model. The objective metric is RMSE. The training job is costly, and the engineer wants to find a good configuration quickly. Which tuning strategy should they use?

A.Bayesian optimization
B.Hyperband
C.Random search
D.Grid search
AnswerA

Bayesian optimization builds a probabilistic model of the objective and selects hyperparameter combinations likely to improve RMSE, converging in fewer training jobs than grid or random search. This satisfies the constraint of finding a good configuration quickly while minimising costly training runs.

Why this answer

Bayesian optimization is the SageMaker Automatic Model Tuning strategy that builds a probabilistic model of the objective function and uses it to choose the next hyperparameter combination, which typically finds good configurations in fewer training jobs than random or grid search. This makes it well suited when each training job is expensive and the engineer wants to minimize cost while still optimizing RMSE.

Exam trap

The trap is confusing Hyperband's early-stopping efficiency with Bayesian optimization's sample efficiency; the question emphasizes costly training jobs, which favors Bayesian optimization's ability to learn from each trial.

How to eliminate wrong answers

Option B is wrong because Hyperband is a multi-fidelity strategy that aggressively stops poorly performing trials early, which is efficient but not the default best choice when the goal is to find a good configuration quickly with a costly training job and no mention of early-stopping infrastructure. Option C is wrong because random search samples hyperparameters independently and does not learn from previous trials, so it usually requires more jobs to reach a good configuration. Option D is wrong because grid search exhaustively evaluates a fixed set of combinations, which is the most expensive approach and scales poorly with the number of hyperparameters.

135
MCQmedium

A machine learning team is using Amazon SageMaker to train a model. They notice that the training job is taking longer than expected and the logs show repeated warnings about 'loss not decreasing'. Which SageMaker feature should they use to diagnose and visualize the training process?

A.Amazon SageMaker Clarify
B.Amazon SageMaker Experiments
C.Amazon SageMaker Debugger
D.Amazon SageMaker Model Monitor
AnswerC

SageMaker Debugger captures tensors during training and provides built-in rules that detect issues such as vanishing gradients or loss not decreasing, plus visualisations of the training process. This directly diagnoses the repeated 'loss not decreasing' warnings reported in the job logs.

Why this answer

Amazon SageMaker Debugger is the correct choice because it provides real-time monitoring and visualization of training metrics, including loss values, gradients, and weights. The repeated 'loss not decreasing' warnings indicate a training issue (e.g., vanishing gradients or learning rate problems), and Debugger can capture these tensors and emit alerts or trigger actions (like stopping the job) via built-in or custom rules. It also integrates with SageMaker Studio for interactive visualization of the training progress.

Exam trap

The trap here is that candidates often confuse SageMaker Debugger with SageMaker Experiments, thinking both are for monitoring training metrics, but Experiments only logs high-level metrics (like final loss or accuracy) while Debugger provides deep, step-by-step tensor-level diagnostics for issues like loss stagnation.

How to eliminate wrong answers

Option A is wrong because Amazon SageMaker Clarify is designed for bias detection and explainability of model predictions, not for monitoring training metrics like loss. Option B is wrong because Amazon SageMaker Experiments is used for tracking and comparing different training runs (e.g., hyperparameters, metrics), but it does not provide real-time, in-depth debugging of internal tensors or loss plateaus during a single training job. Option D is wrong because Amazon SageMaker Model Monitor focuses on detecting data drift and quality issues in deployed models (inference endpoints), not on diagnosing training-time problems like loss stagnation.

136
MCQeasy

Refer to the exhibit. A team has configured data capture for a SageMaker endpoint. The endpoint is returning predictions but no captured data appears in the S3 bucket. What is the most likely cause?

A.The InitialSamplingPercentage is too low.
B.The IAM role for the endpoint does not have s3:PutObject permission.
C.The capture status is 'Configured' but not 'Running'.
D.The endpoint is not receiving any traffic.
AnswerB

Without s3:PutObject on the endpoint's execution role, SageMaker cannot write captured inference records to the target bucket, so requests succeed while capture silently fails. The stem's constraint is predictions returning normally yet no objects landing in S3, which isolates the fault to the write permission rather than the capture configuration itself.

Why this answer

The most likely cause is that the IAM role associated with the SageMaker endpoint lacks the `s3:PutObject` permission. Without this permission, the endpoint can generate capture data internally but cannot write it to the specified S3 bucket, resulting in no captured data appearing even though predictions are returned successfully.

Exam trap

The trap here is that candidates often focus on sampling percentages or traffic volume, but the core issue is almost always an IAM permissions misconfiguration when predictions succeed but data capture fails silently.

How to eliminate wrong answers

Option A is wrong because a low `InitialSamplingPercentage` would reduce the amount of data captured, not eliminate it entirely; some data would still appear in S3. Option C is wrong because the capture status 'Configured' is the expected state for a properly enabled data capture configuration; there is no 'Running' status for capture itself—capture runs automatically when the endpoint is active. Option D is wrong because the endpoint is returning predictions, which directly indicates it is receiving traffic; if there were no traffic, no predictions would be returned.

137
MCQhard

An ML engineer is preparing a time-series dataset for a forecasting model that predicts daily sales for the next 30 days. The dataset contains 3 years of daily sales data. Which data splitting strategy should the engineer use to evaluate the model's performance on future data?

A.Leave-one-out cross-validation
B.Random 80/20 train-test split
C.Stratified k-fold cross-validation
D.Walk-forward validation (time-series split)
AnswerD

Walk-forward validation trains on all data up to a cutoff and tests on the immediately following window, then rolls the cutoff forward. This preserves chronological order, so the model is always evaluated on genuinely future observations rather than leaking later sales into training.

Why this answer

Walk-forward validation (time-series split) preserves the temporal order of observations by training on past data and testing on subsequent future periods, which mirrors how a forecasting model is actually used in production. Because daily sales data has trend, seasonality, and autocorrelation, random or stratified splits leak future information into training and produce optimistically biased metrics. Walk-forward validation with a 30-day horizon directly simulates predicting the next 30 days from prior history.

Exam trap

MLA-C01 often tests whether candidates recognize that standard cross-validation techniques (k-fold, stratified, leave-one-out) leak future information in time-series problems, tempting them to pick a familiar but temporally invalid split.

How to eliminate wrong answers

Option A is wrong because leave-one-out cross-validation randomly holds out individual observations, breaking temporal order and allowing the model to learn from future data points adjacent to the held-out sample, which inflates accuracy. Option B is wrong because a random 80/20 split shuffles time, so the training set contains future dates relative to test dates, causing data leakage and unrealistic performance estimates. Option C is wrong because stratified k-fold preserves class proportions but still shuffles observations across folds, which is meaningless for continuous time-series targets and again leaks future information into training.

138
Multi-Selectmedium

A company wants to secure access to a SageMaker real-time endpoint. Which TWO actions should be taken? (Select two.)

Select 2 answers
A.Use an IAM role with sts:AssumeRole for invocation.
B.Attach a resource-based policy to the endpoint.
C.Enable AWS WAF on the endpoint.
D.Use AWS CloudTrail to log all invocations.
E.Configure the endpoint to be private within a VPC and use VPC endpoints.
AnswersB, E

A resource-based policy on the endpoint defines which principals may invoke it, restricting access at the endpoint itself. This satisfies the requirement to secure access by controlling cross-account or explicit principal permissions, complementing identity-based IAM policies.

Why this answer

Option B is correct because SageMaker real-time endpoints support resource-based policies (endpoint policies) that let you grant or restrict InvokeEndpoint access to specific principals, AWS accounts, organizations, or source VPC/VPC endpoint conditions, which directly secures who can invoke the endpoint. Option E is correct because configuring the endpoint as private within a VPC and accessing it through an interface VPC endpoint (AWS PrivateLink) keeps invocation traffic off the public internet and enforces network-level isolation. Option A is not correct because sts:AssumeRole is an IAM permission for obtaining temporary credentials, not a mechanism for securing endpoint invocation itself.

Option C is not correct because AWS WAF protects HTTP(S) resources like CloudFront, ALB, and API Gateway, and cannot be attached to a SageMaker endpoint. Option D is not correct because CloudTrail provides auditing and logging of API activity, not access control or security enforcement.

Exam trap

The trap here is that candidates often confuse sts:AssumeRole with direct invocation permissions, or think that AWS WAF can be applied to any AWS service endpoint, when in fact SageMaker endpoints are not supported by WAF.

139
MCQmedium

A company runs a batch inference job on 10 TB of image data stored in S3. Each image needs to be processed by a GPU-accelerated model. The job is not time-sensitive and cost is the primary concern. Which SageMaker option is MOST appropriate?

A.SageMaker Serverless Inference
B.SageMaker Batch Transform with GPU instance and spot instances
C.SageMaker Async Inference with GPU
D.SageMaker real-time endpoint on GPU instances
AnswerB

Batch Transform runs inference over S3 data without a persistent endpoint, and pairing GPU instances with spot capacity cuts cost substantially for a non-time-sensitive job. This satisfies the cost-primary constraint, since managed spot training applies to training jobs, not batch inference.

Why this answer

Batch Transform with GPU spot instances is the most cost-effective choice for a non-time-sensitive, large-scale batch inference job on 10 TB of data. Spot instances offer up to 90% cost savings over on-demand, and Batch Transform natively handles splitting the dataset, distributing work across instances, and writing results to S3 without requiring a persistent endpoint.

Exam trap

The trap here is that candidates confuse 'batch inference' with 'async inference' and choose Option C, not realizing that Async Inference still requires a running endpoint and is designed for near-real-time processing, not cost-optimized offline batch jobs.

How to eliminate wrong answers

Option A is wrong because SageMaker Serverless Inference is designed for intermittent, low-latency workloads with a maximum payload size of 6 MB and a maximum concurrency of 200, making it unsuitable for processing 10 TB of image data. Option C is wrong because SageMaker Async Inference is optimized for near-real-time requests with large payloads (up to 1 GB) and requires a persistent endpoint, incurring higher costs than a batch job that can use spot instances. Option D is wrong because SageMaker real-time endpoints are provisioned 24/7 and designed for low-latency, high-throughput serving, which is wasteful and expensive for a non-time-sensitive batch job that can tolerate startup delays and interruptions.

140
MCQeasy

A company wants to use SageMaker to serve real-time predictions with a model that has a large memory footprint. They need to ensure the endpoint can handle traffic spikes. Which scaling policy should they use?

A.Simple scaling policy
B.Scheduled scaling policy
C.Target tracking policy
D.Step scaling policy
AnswerC

Target tracking adjusts instance count to hold a chosen metric, such as InvocationsPerInstance, at a target value. For a large-memory model, this scales out on demand during traffic spikes and back in afterwards, matching capacity to load without manual intervention.

Why this answer

Target tracking scaling policy is the correct choice because it automatically adjusts the number of instances in the SageMaker endpoint based on a target metric, such as InvocationsPerInstance or ModelLatency, to handle traffic spikes without manual intervention. This policy is ideal for real-time inference with large memory models because it dynamically scales resources up or down to maintain the target metric, ensuring consistent performance during unpredictable traffic bursts.

Exam trap

The trap here is that candidates often confuse step scaling with target tracking, assuming step scaling is more responsive for spikes, but target tracking is actually the recommended and simpler approach for handling unpredictable traffic in SageMaker real-time endpoints.

How to eliminate wrong answers

Option A is wrong because simple scaling policy only triggers a single adjustment based on a CloudWatch alarm breach and then waits for a cooldown period, which cannot handle rapid traffic spikes effectively and may lead to under- or over-provisioning. Option B is wrong because scheduled scaling policy adjusts capacity at predetermined times, which is unsuitable for unpredictable traffic spikes that do not follow a fixed schedule. Option D is wrong because step scaling policy requires defining multiple step adjustments with thresholds, which is more complex to configure and may not react as smoothly to sudden spikes compared to target tracking, which continuously adjusts to maintain a target metric.

141
MCQmedium

A team is training a PyTorch model using SageMaker and wants to use their own custom training container with a specific PyTorch version. Which approach should they use?

A.Use the SageMaker built-in PyTorch estimator and set the framework_version
B.Use SageMaker Bring Your Own Container (BYOC) with a custom Docker image
C.Use SageMaker Script Mode with a PyTorch script
D.Use SageMaker Autopilot to automatically select the container
AnswerB

BYOC lets the team supply a custom Docker image containing their exact PyTorch version, so the training environment matches their dependency requirements precisely. SageMaker's prebuilt PyTorch containers fix the framework version, which cannot satisfy the stated need for a specific PyTorch build.

Why this answer

When a team needs a custom training container with a specific PyTorch version or custom dependencies that are not available in the built-in SageMaker images, the correct approach is Bring Your Own Container (BYOC), where they build a Docker image and push it to Amazon ECR, then reference it in the SageMaker estimator. This gives full control over the framework version and environment.

Exam trap

The trap is assuming Script Mode allows arbitrary framework versions; in reality Script Mode still relies on a SageMaker-managed container, so only BYOC provides full control over the framework version.

How to eliminate wrong answers

Option A is wrong because the built-in PyTorch estimator only supports the framework versions that SageMaker provides, so it cannot satisfy a requirement for a specific custom PyTorch version. Option C is wrong because Script Mode still uses a SageMaker-provided framework container and only allows the training script to be customized, not the underlying framework version or dependencies. Option D is wrong because SageMaker Autopilot is an automated machine learning feature that selects algorithms and containers automatically, which is the opposite of specifying a custom container.

142
MCQeasy

A data engineer is preparing a dataset for a k-means clustering algorithm. The features have different scales: age (18-100), income ($20k-$200k), and number of purchases (0-50). Without scaling, which feature will dominate the distance calculations?

A.All features will contribute equally
B.Income
C.Number of purchases
D.Age
AnswerB

Euclidean distance sums squared differences, so the feature with the widest numeric range contributes most. Income spans roughly $180k against age's 82 and purchases' 50, making its squared deviations dominate the distance metric and skew cluster assignment.

Why this answer

Income has the largest range (180,000 compared to 82 and 50), so it will dominate Euclidean distance calculations. Standardization or normalization is needed before clustering.

143
MCQeasy

A data scientist is using SageMaker to train a linear regression model on a dataset with a large number of features. They notice that the model's training time is long and want to speed it up by using a more efficient algorithm. They decide to use the SageMaker built-in Linear Learner algorithm. Which of the following is a key advantage of using the Linear Learner algorithm in SageMaker for this scenario?

A.It uses a built-in automatic model tuning feature that always finds the optimal hyperparameters without user intervention.
B.It automatically performs feature engineering and selection, reducing the need for manual preprocessing.
C.It is specifically designed for deep learning models and can leverage GPUs for faster training.
D.It supports both regression and classification and can handle large-scale datasets efficiently using distributed training.
AnswerD

SageMaker Linear Learner is designed for large-scale linear models and supports both regression and classification. It can be trained in distributed mode across multiple instances, which speeds up training on large datasets. This makes it a suitable choice for the data scientist's scenario of a large number of features and long training time.

Why this answer

SageMaker Linear Learner is optimized for large-scale linear models and supports both regression and classification. It can be trained in distributed mode across multiple instances, which significantly reduces training time for datasets with many features. This makes it a suitable choice for the data scientist's scenario.

Exam trap

The trap here is assuming that Linear Learner automatically performs feature engineering or hyperparameter tuning, which it does not; these are separate steps.

144
MCQmedium

Refer to the exhibit. A data scientist configured an automatic model tuning job for a classification model. The tuning job completed after 20 training jobs, but the best validation accuracy was only 0.65. What is the most effective way to potentially improve the result?

A.Increase MaxNumberOfTrainingJobs to 100
B.Change the strategy to Random
C.Change the objective metric to training:accuracy
D.Increase MaxParallelTrainingJobs to 10
AnswerA

The tuning job explored only 20 configurations, likely too few to locate a strong region of the hyperparameter space. Raising MaxNumberOfTrainingJobs to 100 lets the tuner evaluate more candidates, improving the chance of higher validation accuracy.

Why this answer

Increasing MaxNumberOfTrainingJobs to 100 allows the automatic model tuning job to explore a larger hyperparameter space, giving the Bayesian optimization strategy (the default) more trials to converge on a better configuration. With only 20 training jobs, the tuner may not have had enough iterations to balance exploration and exploitation, especially for a complex classification model. More jobs increase the likelihood of finding a hyperparameter combination that yields higher validation accuracy.

Exam trap

AWS often tests the misconception that increasing parallelism (MaxParallelTrainingJobs) improves model quality, when in fact it only speeds up execution without increasing the total number of trials, which is the key lever for better hyperparameter optimization.

How to eliminate wrong answers

Option B is wrong because changing the strategy to Random would ignore the information gained from completed trials, making the search less efficient than Bayesian optimization, which uses past results to guide future hyperparameter choices. Option C is wrong because changing the objective metric to training:accuracy would optimize for training performance rather than generalization, likely leading to overfitting and not improving validation accuracy. Option D is wrong because increasing MaxParallelTrainingJobs to 10 does not increase the total number of trials; it only runs more jobs concurrently, which can reduce wall-clock time but does not expand the search space or improve the final model's accuracy.

145
MCQhard

A social media company is processing a real-time stream of user activity data from Amazon Kinesis Data Streams to train a machine learning model for content recommendation. The raw data includes user ID, timestamp, content ID, interaction type (like, share, comment), and device type. The data scientists need to aggregate features per user over a sliding window of 7 days, including counts of interaction types, unique content IDs engaged, and a moving average of interaction timestamps. The aggregated data will be used to update a user embedding model. The streaming data volume is approximately 500 records per second, and the company uses an AWS Glue streaming ETL job for transformation. However, the Glue job is failing frequently with high latency and checkpoint errors. The team needs a more robust solution to prepare the streaming data features. Which approach should the team take?

A.Increase the DPU count on the Glue streaming ETL job and reduce the checkpoint interval to improve performance.
B.Use Amazon Kinesis Data Analytics for Apache Flink to perform the sliding window aggregations with built-in state management and exactly-once processing, then write the features to S3 and DynamoDB.
C.Use AWS Lambda functions to process records from Kinesis, store intermediate aggregation results in Amazon DynamoDB, and read them back to compute windowed features.
D.Use Amazon SageMaker Processing jobs that run periodically every hour to read data from S3 (landing from Kinesis Firehose) and perform the aggregations batch-wise.
AnswerB

Apache Flink on Kinesis Data Analytics provides native event-time sliding windows with managed keyed state and exactly-once checkpointing, eliminating the Glue job's checkpoint failures and latency at 500 records per second. It satisfies the 7-day per-user aggregation constraint, then sinks features to S3 and DynamoDB.

Why this answer

Amazon Kinesis Data Analytics for Apache Flink provides native support for sliding window aggregations with managed state and exactly-once processing semantics, which directly addresses the high latency and checkpoint errors seen in the Glue streaming ETL job. Flink's checkpointing mechanism ensures fault-tolerant state management for the 7-day sliding window, while Glue's Spark Streaming engine struggles with long-running stateful operations at 500 records/sec due to its micro-batch architecture and checkpoint overhead.

Exam trap

The trap here is that candidates assume increasing resources (DPU) on Glue streaming ETL will fix performance issues, but the root cause is Spark's micro-batch architecture's inability to efficiently manage long-running stateful sliding windows, which Flink's native streaming engine is designed for.

How to eliminate wrong answers

Option A is wrong because increasing DPU count and reducing checkpoint interval on a Glue streaming ETL job exacerbates checkpoint errors and latency due to Spark's micro-batch overhead and lack of native long-lived state management for sliding windows. Option C is wrong because AWS Lambda functions have a maximum execution timeout of 15 minutes and no built-in state management, making them unsuitable for maintaining 7-day sliding window aggregations across 500 records/sec without external state stores that introduce eventual consistency and latency. Option D is wrong because using hourly SageMaker Processing jobs on S3 data from Kinesis Firehose introduces a minimum 1-hour delay, which violates the real-time requirement for updating a user embedding model with sliding window features.

146
Multi-Selectmedium

A machine learning engineer is deploying a model using SageMaker and needs to ensure that the endpoint can automatically scale based on traffic patterns. Which TWO actions should the engineer take? (Choose two.)

Select 2 answers
A.Define a scaling policy using Application Auto Scaling for the SageMaker endpoint variant.
B.Set up an Amazon CloudWatch alarm to trigger scaling based on the InvocationsPerInstance metric.
C.Enable SageMaker Model Monitor to detect data drift.
D.Configure a multi-model endpoint to serve multiple models.
E.Use SageMaker batch transform to handle variable traffic.
AnswersA, B

Application Auto Scaling registers the SageMaker endpoint variant as a scalable target and applies a target-tracking or step policy, enabling automatic replica adjustment. This satisfies the requirement for traffic-based scaling by letting the endpoint add or remove instances as load changes.

Why this answer

Option A is correct because SageMaker endpoint variants are scaled through Application Auto Scaling, which is the AWS service that registers the SageMaker variant as a scalable target and applies a scaling policy (target-tracking or step scaling) to adjust the desired instance count. Option B is correct because Application Auto Scaling policies are driven by Amazon CloudWatch alarms, and the InvocationsPerInstance metric is the standard SageMaker metric used to scale on traffic per instance. Option C is incorrect because SageMaker Model Monitor detects data drift and quality issues, not traffic-based scaling.

Option D is incorrect because multi-model endpoints consolidate multiple models on shared infrastructure to reduce hosting cost, not to autoscale on traffic patterns. Option E is incorrect because batch transform is an offline, batch inference mechanism and does not serve a real-time, autoscaling endpoint.

Exam trap

The trap here is confusing monitoring and scaling: candidates often pick Model Monitor (Option C) because it sounds like it monitors traffic, but it is for data drift, not scaling; similarly, batch transform (Option E) is mistaken for a scaling solution when it is a separate inference mode.

147
MCQeasy

An ML engineer needs to create a feature store that supports both low-latency online inference and large-scale offline training. The features are updated hourly from a streaming source. Which Amazon SageMaker Feature Store configuration should the engineer use?

A.Create a feature group with both online and offline stores enabled.
B.Create a feature group with only an online store enabled.
C.Create two separate feature groups: one for online and one for offline.
D.Create a feature group with only an offline store enabled.
AnswerA

Enabling both online and offline stores in one feature group lets SageMaker Feature Store serve low-latency reads for real-time inference while retaining the full history in Amazon S3 for large-scale offline training, satisfying both latency and volume constraints.

Why this answer

A SageMaker Feature Store feature group can be configured with both an online store (for low-latency real-time inference) and an offline store (for large-scale training). This single feature group design supports both use cases and is the recommended configuration when features are updated hourly from a streaming source.

Exam trap

MLA-C01 often tests the need for both online and offline stores in a single feature group, tricking candidates into thinking separate feature groups are required for online vs. offline use cases.

How to eliminate wrong answers

Option B is wrong because an online-only feature group lacks the offline store needed for large-scale training and historical data access. Option C is wrong because creating two separate feature groups duplicates data and complicates consistency; a single feature group with both stores is the intended design. Option D is wrong because an offline-only feature group cannot serve low-latency online inference.

148
MCQhard

A company wants to forecast monthly sales that show clear seasonality. Which algorithm is most suitable?

A.ARIMA (Seasonal ARIMA)
B.Random forest
C.K-means clustering
D.Linear regression
AnswerA

Seasonal ARIMA explicitly models both trend and repeating seasonal cycles through its seasonal (P, D, Q, s) terms, capturing the monthly sales pattern. Plain ARIMA lacks this seasonal component, so SARIMA satisfies the clear seasonality constraint.

Why this answer

Seasonal ARIMA (SARIMA) extends ARIMA by explicitly modeling seasonal components through seasonal differencing and seasonal autoregressive/moving average terms, making it the most suitable algorithm for forecasting monthly sales with clear seasonality. It captures both trend and seasonal patterns by incorporating parameters for the seasonal period (e.g., 12 for monthly data) and can handle non-stationary time series.

Exam trap

AWS often tests the distinction between time-series-specific algorithms (like ARIMA) and general-purpose machine learning models (like random forest or linear regression), trapping candidates who overlook that seasonal patterns require explicit temporal modeling rather than treating data as independent observations.

How to eliminate wrong answers

Option B (Random forest) is wrong because it is an ensemble tree-based method that does not inherently model temporal dependencies or seasonality; it treats each time point as independent, leading to poor extrapolation and inability to capture periodic patterns. Option C (K-means clustering) is wrong because it is an unsupervised clustering algorithm used for grouping data points, not for forecasting time series; it has no mechanism to predict future values or model seasonal cycles. Option D (Linear regression) is wrong because it assumes a linear relationship between independent variables and the target, but it cannot capture complex seasonal patterns or autocorrelation without extensive feature engineering (e.g., manually adding lagged or seasonal dummy variables), and it fails to handle non-stationary time series effectively.

149
MCQhard

A company uses SageMaker Model Monitor's feature attribution drift monitoring with SHAP. They receive an alert that the average SHAP value for a particular feature has increased significantly compared to the baseline. The feature's input distribution has not changed. What does this likely indicate?

A.The feature is no longer relevant to predictions
B.A bug in the SHAP computation
C.Data drift in that feature
D.Concept drift in the model
AnswerD

Feature attribution drift with unchanged input distribution isolates the model's learned relationship: SHAP values rising while inputs stay stable means the model now weights that feature differently, which is concept drift rather than data drift.

Why this answer

Feature attribution drift monitoring compares SHAP value distributions between baseline and current data. If the input distribution of a feature is unchanged but its average SHAP value shifts significantly, the model's learned relationship between that feature and the target has changed — this is the definition of concept drift. The model itself is unchanged, but the underlying mapping from inputs to outputs in the real world has shifted, so the same input values now contribute differently to predictions.

Exam trap

The trap is conflating data drift (change in input distribution P(X)) with concept drift (change in the relationship P(Y|X)) — the question deliberately states inputs are unchanged to force you to recognize concept drift.

How to eliminate wrong answers

Option A is wrong because an increased SHAP magnitude means the feature is contributing more to predictions, not less — irrelevance would show as SHAP values shrinking toward zero. Option B is wrong because a SHAP computation bug would typically produce erratic or uniform anomalies across many features, not a targeted, consistent increase for one feature while inputs remain stable. Option C is wrong because data drift refers to changes in the input distribution (P(X)), and the question explicitly states the input distribution has not changed — so data drift is ruled out by the premise.

150
MCQhard

A machine learning team is building a feature store using Amazon SageMaker Feature Store. They need to store features that support both real-time inference (low latency) and historical training. Which configuration should they choose?

A.Create two separate feature groups: one online and one offline
B.Create a feature group with both online and offline stores enabled
C.Create a feature group with only an online store enabled
D.Create a feature group with only an offline store enabled
AnswerB

Enabling both online and offline stores on the feature group writes records to a low-latency online store for real-time inference while also landing them in Amazon S3 for historical training queries, satisfying the dual latency and training requirements in one configuration.

Why this answer

Amazon SageMaker Feature Store allows a single feature group to have both an online store and an offline store enabled. The online store provides low-latency access for real-time inference, while the offline store stores historical data in S3 for training and batch scoring. This configuration satisfies both requirements without duplicating feature groups.

Exam trap

MLA-C01 often tests the misconception that you need separate feature groups for online and offline, when a single feature group can enable both stores.

How to eliminate wrong answers

Option A is wrong because creating two separate feature groups would duplicate data and management overhead, and is not the recommended approach when a single feature group can serve both purposes. Option C is wrong because an online-only feature group lacks the historical data needed for training. Option D is wrong because an offline-only feature group cannot serve low-latency real-time inference.

Page 1

Page 2 of 9

Page 3

All pages