Courseiva

CCNA Mla Model Development Questions

33 of 108 questions · Page 2/2 · Mla Model Development topic · Answers revealed

76
Multi-Selecthard

An ML engineer is fine-tuning a foundation model using RLHF on SageMaker. Which THREE components are essential for this workflow? (Select THREE.)

Select 3 answers
A.A reward model trained on the preference data
B.A large validation dataset for final evaluation
C.The PPO (Proximal Policy Optimization) algorithm for model updates
D.A preference dataset with human rankings
E.A PEFT technique like LoRA
AnswersA, C, D

RLHF requires a reward model that scores model outputs against learned human preferences, supplying the scalar signal PPO maximises. Without it, no preference-based objective exists, so the fine-tuning loop cannot optimise the foundation model toward preferred responses.

Why this answer

Option A is correct because RLHF requires a reward model that has been trained on human preference data to serve as the scalar reward signal guiding policy optimization. Option C is correct because PPO (Proximal Policy Optimization) is the standard reinforcement learning algorithm used to update the policy (the language model) against the reward model while constraining updates via a KL penalty to the reference model. Option D is correct because a preference dataset containing human rankings (e.g., chosen vs. rejected responses) is the foundational input used to train the reward model in the first place.

Option B is not essential to the RLHF training workflow itself; a validation set is useful for evaluation but is not a required RLHF component. Option E is not essential because PEFT methods like LoRA are an optional efficiency technique for fine-tuning, not a required element of the RLHF pipeline.

Exam trap

MLA-C01 often tests the distinction between RLHF's essential components (preference data, reward model, PPO) and optional optimizations like PEFT or evaluation datasets, causing candidates to select non-essential items.

77
MCQeasy

A team wants to fine-tune a pre-trained Hugging Face transformer model for text classification using SageMaker. They have a custom training script. Which SageMaker estimator should they use?

A.SageMaker generic estimator with a custom container
B.SageMaker Hugging Face estimator
C.SageMaker PyTorch estimator
D.SageMaker TensorFlow estimator
AnswerB

The Hugging Face estimator ships pre-built containers with transformers, PyTorch and TensorFlow libraries, so a custom training script written against the Hugging Face API runs without dependency work. It satisfies the requirement to fine-tune a pre-trained transformer for text classification while still accepting the team's own script.

Why this answer

The SageMaker Hugging Face estimator is purpose-built for Hugging Face transformer models: it uses the official Hugging Face Deep Learning Containers, supports the transformers, datasets, and accelerate libraries, and accepts hyperparameters like epoch, learning_rate, and model_name_or_path. It is the natural choice for fine-tuning a pre-trained Hugging Face model with a custom training script.

Exam trap

MLA-C01 often tests whether candidates know that Hugging Face models have a dedicated SageMaker estimator — the trap is picking the PyTorch estimator because transformers are PyTorch-based, but the PyTorch DLC does not ship with Hugging Face libraries by default.

How to eliminate wrong answers

Option A is wrong because a generic estimator with a custom container requires you to build and maintain your own Docker image with all Hugging Face dependencies — unnecessary overhead when an official estimator exists. Option C is wrong because the PyTorch estimator uses the PyTorch DLC, which does not include Hugging Face libraries by default; you would have to install them via a requirements.txt, adding complexity. Option D is wrong because the TensorFlow estimator is for TensorFlow models and does not natively support Hugging Face transformers, which are PyTorch-first.

78
MCQmedium

A machine learning engineer is preparing a training dataset stored in Amazon S3 for a SageMaker training job. The data is in CSV format, and the engineer wants to ensure that the training job can access the data efficiently and securely. The S3 bucket is in the same AWS Region as the SageMaker training job. The engineer needs to provide the training job with the necessary permissions to read the data. Which of the following is the MOST secure and appropriate way to grant the training job access to the S3 bucket?

A.Store AWS credentials (access key and secret key) in the training script and use them to access S3 directly.
B.Create an IAM role with a policy that allows s3:GetObject on the specific S3 bucket and attach it to the SageMaker training job.
C.Make the S3 bucket public and allow the training job to read the data without authentication.
D.Generate a presigned URL for each object in the S3 bucket and pass the URLs as hyperparameters to the training job.
AnswerB

This approach follows the principle of least privilege by granting only the necessary s3:GetObject permission on the specific bucket. The IAM role is assumed by the SageMaker training job, allowing secure access without embedding credentials. It is the recommended practice for SageMaker training jobs to access S3 data.

Why this answer

The correct approach is to create an IAM role with a policy that grants s3:GetObject permission on the specific S3 bucket and attach it to the SageMaker training job. This follows the principle of least privilege and allows the training job to access the data securely without embedding credentials. It is the standard method for granting SageMaker training jobs access to S3 data.

Exam trap

The trap here is assuming that presigned URLs or embedded credentials are acceptable for long-running training jobs, when they introduce security and reliability risks.

79
MCQeasy

A data scientist is using SageMaker built-in XGBoost algorithm for a binary classification task. Which objective metric is MOST appropriate for SageMaker Automatic Model Tuning to maximize?

A.validation:mae
B.validation:rmse
C.validation:ndcg
D.validation:auc
AnswerD

For binary classification, `validation:auc` maximises the area under the ROC curve, giving a threshold-independent measure of class separation. This satisfies the stem's requirement for the most appropriate tuning objective, since AUC handles imbalanced binary labels better than accuracy and is natively supported by SageMaker Automatic Model Tuning.

Why this answer

For binary classification, AUC (Area Under the ROC Curve) is the standard evaluation metric because it measures the model's ability to discriminate between the two classes across all classification thresholds, independent of class balance. SageMaker's built-in XGBoost exposes validation:auc as the objective metric for binary classification tuning. MAE and RMSE are regression metrics, and NDCG is a ranking metric, so none of them fit binary classification.

Exam trap

The trap is mixing regression metrics (MAE, RMSE) and ranking metrics (NDCG) into a binary classification question — candidates who do not map metric families to task types will pick a plausible-sounding but wrong objective.

How to eliminate wrong answers

Option A is wrong because validation:mae (mean absolute error) is a regression metric measuring average absolute prediction error, which is meaningless for binary class labels. Option B is wrong because validation:rmse (root mean squared error) is also a regression metric, penalizing large errors in continuous predictions, not classification correctness. Option C is wrong because validation:ndcg (normalized discounted cumulative gain) is a ranking-quality metric used for learning-to-rank tasks, not binary classification.

80
MCQeasy

A data scientist wants to track the training and validation accuracy of a SageMaker training job over time. They need to visualize these metrics in Amazon CloudWatch. Which action should they take?

A.Configure the training job to write metrics directly to Amazon CloudWatch using the AWS SDK for Python (Boto3) in the training script.
B.Use the SageMaker estimator's metric_definitions parameter to specify regex patterns that extract metrics from the training logs.
C.Write the metrics to a file in the output S3 bucket and create a CloudWatch dashboard from the file.
D.Enable SageMaker Debugger and configure rules to emit metrics to CloudWatch.
AnswerB

SageMaker automatically streams training logs to CloudWatch Logs. By defining metric_definitions with regex patterns, SageMaker extracts the specified metrics from the logs and publishes them as CloudWatch metrics. This enables real-time monitoring and visualization without custom code.

Why this answer

The metric_definitions parameter in the SageMaker estimator allows you to define regex patterns that extract metric values from the training logs. SageMaker then automatically publishes these as CloudWatch metrics, enabling visualization. This is the recommended and simplest approach for tracking standard training metrics.

Exam trap

The trap here is overcomplicating the solution by considering custom code or other SageMaker features, when the built-in metric extraction is sufficient.

81
Multi-Selecthard

A data scientist is using SageMaker Experiments to track multiple training runs. They want to compare runs based on the objective metric and visualize performance. Which THREE steps should they perform? (Choose THREE.)

Select 3 answers
A.Deploy the best model to an endpoint
B.Use SageMaker Studio Experiments UI to list and compare trials
C.Log hyperparameters and metrics using the SageMaker SDK
D.Create a SageMaker Experiment
E.Enable SageMaker Model Monitor for each run
AnswersB, C, D

The SageMaker Studio Experiments UI lists trials and plots objective metrics side by side, enabling direct comparison and visualisation of performance across runs. This satisfies the requirement to compare runs by objective metric and visualise their results.

Why this answer

Option D is correct because a SageMaker Experiment is the top-level container that groups related trials (training runs) so they can be tracked and compared together. Option C is correct because the SageMaker SDK (e.g., Run, Trial, Tracker, or the log_parameters/log_metric calls) is how hyperparameters and objective metrics get recorded for each run, which is required before any comparison can be made. Option B is correct because the SageMaker Studio Experiments UI lets the data scientist list trials and compare them visually by objective metric, directly satisfying the stated goal of comparing runs and visualizing performance.

Option A is not required because deploying the best model to an endpoint is an inference/hosting step, not part of tracking or comparing experiments. Option E is not required because Model Monitor is used for detecting drift and data quality issues on deployed endpoints, not for logging or comparing training-run metrics.

Exam trap

MLA-C01 often tests whether candidates conflate Experiment tracking with deployment or monitoring, causing them to select Model Monitor or endpoint deployment as part of the comparison workflow.

82
MCQeasy

A data scientist wants to quickly build a binary classification model without writing any code. Which SageMaker feature is MOST suitable?

A.SageMaker Debugger
B.SageMaker Model Monitor
C.SageMaker Ground Truth
D.SageMaker Autopilot
AnswerD

SageMaker Autopilot automates model selection, feature engineering and hyperparameter tuning, then generates candidate models with full visibility, satisfying the no-code requirement. It handles binary classification directly, so the data scientist obtains a deployable model without authoring training scripts.

Why this answer

SageMaker Autopilot is a no-code AutoML feature that automatically explores data, selects algorithms, tunes hyperparameters, and builds the best model for a given tabular dataset, including binary classification. It requires no model-building code from the user, making it the most suitable choice. The other options are operational/monitoring or labeling tools, not model-building features.

Exam trap

The trap is that several SageMaker features have 'model' in their name (Model Monitor, Debugger), so candidates associate them with model building — but only Autopilot actually creates models without code.

How to eliminate wrong answers

Option A is wrong because SageMaker Debugger is a training-time debugging and profiling tool that inspects tensors and metrics during training — it does not build models. Option B is wrong because SageMaker Model Monitor detects data drift and quality issues on deployed models; it is a post-deployment monitoring service, not a model builder. Option C is wrong because SageMaker Ground Truth is a data labeling service for creating training datasets, not a model-building feature.

83
Multi-Selecthard

A data scientist is training a model with SageMaker and needs to reduce the cost of a long-running training job that can tolerate interruptions. The job uses a custom training script and reads data from Amazon S3. The data scientist wants the job to resume from the last saved state if the underlying compute is reclaimed. (Choose two.)

Select 2 answers
A.Set the estimator's enable_spot_training parameter to true and rely on SageMaker to automatically restart the job from the beginning after an interruption.
B.Increase the volume_size parameter so that the container's local disk can hold the entire dataset and all checkpoints during training.
C.Configure the estimator to use managed spot training and set max_wait time greater than max_run time.
D.Enable network isolation on the estimator to prevent the training container from accessing the internet during spot interruptions.
E.Set the checkpoint_s3_uri parameter on the estimator to an Amazon S3 location where the training script saves checkpoints.
AnswersC, E

Managed spot training uses Amazon EC2 Spot Instances, which are cheaper but can be interrupted. Setting max_wait greater than max_run allows the job to wait for capacity and to resume after an interruption within the overall wait window. This directly addresses cost reduction while tolerating interruptions, and it is a required configuration for the resume behavior to be useful.

Why this answer

Managed spot training lowers cost by using Spot Instances, and checkpointing to Amazon S3 lets a restarted job resume from the last saved state. Setting max_wait above max_run gives the job time to wait for capacity and to recover from interruptions. Network isolation, larger volumes, and restarting from scratch do not support the resume requirement and can increase cost or break access.

Exam trap

The trap here is assuming spot training alone preserves progress, when checkpointing to Amazon S3 is what enables resuming after an interruption.

84
MCQhard

A data scientist is training a model using SageMaker and wants to use spot instances to reduce costs. The training job is checkpointed every 5 minutes. However, the job gets interrupted frequently and never completes. What is the MOST likely cause?

A.The checkpoint interval is too long relative to the interruption frequency
B.The checkpoint S3 URI is incorrect
C.The instance type is too small for the training job
D.The job is configured with too few max retries
AnswerA

Spot capacity reclaims instances faster than the five-minute checkpoint cadence, so each interruption discards up to five minutes of progress before the next save. Shortening the checkpoint interval relative to interruption frequency satisfies the stem's requirement that the job actually complete.

Why this answer

SageMaker spot instances can be interrupted with little notice. If the checkpoint interval (5 minutes) is longer than the average time between interruptions, the job loses more progress than it saves, so it never completes. The most likely cause is that the checkpoint interval is too long relative to the interruption frequency, preventing effective resume.

Exam trap

MLA-C01 often tests whether candidates blame configuration errors (S3 URI, retries) instead of the fundamental trade-off between checkpoint frequency and interruption rate — the trap is missing that frequent interruptions require more frequent checkpoints.

How to eliminate wrong answers

Option B is wrong because an incorrect S3 URI for checkpoints would cause checkpoint writes to fail consistently, not intermittent interruptions; the job would error out rather than just never complete. Option C is wrong because an undersized instance would cause slow training or OOM errors, not frequent spot interruptions. Option D is wrong because too few max retries would cause the job to stop after a few interruptions, but the symptom 'never completes' with frequent interruptions points to checkpointing inefficiency, not retry count.

85
Multi-Selectmedium

A machine learning engineer is using SageMaker Autopilot for AutoML. Which TWO outputs does Autopilot produce?

Select 2 answers
A.A hyperparameter tuning job summary
B.An ensemble of candidate models
C.A data labeling pipeline
D.A single optimal model
E.An explainability report
AnswersB, E

Autopilot automatically selects and combines the best-performing pipelines into a single ensemble model, which typically outperforms any individual candidate. This satisfies the requirement for a deployable artefact alongside the generated notebooks and leaderboard, rather than just raw training data or a single fixed algorithm.

Why this answer

SageMaker Autopilot produces an ensemble of candidate models (B), because it automatically explores multiple algorithms and hyperparameter configurations and then combines the best-performing candidates into an ensemble for deployment. It also produces an explainability report (E), which provides feature importance and model insights so users can understand how the model makes predictions. Autopilot does not output a hyperparameter tuning job summary (A); while it performs hyperparameter optimization internally, the deliverable is not a tuning job summary.

It does not create a data labeling pipeline (C), since Autopilot assumes labeled tabular data and does not manage annotation workflows. It also does not produce only a single optimal model (D), because its output includes multiple candidates and an ensemble rather than just one model.

Exam trap

The trap is thinking Autopilot outputs a single model or a tuning job summary; candidates must remember it produces an ensemble and an explainability report as key artifacts.

86
MCQmedium

A team is fine-tuning a Hugging Face BERT model for text classification using SageMaker. They want to use the Hugging Face estimator for convenience. Which parameter must be set to use a custom training script?

A.framework_version
B.instance_type
C.hyperparameters
D.entry_point
AnswerD

The entry_point parameter specifies the path to the custom training script within the source directory, which the Hugging Face estimator then executes. This satisfies the stem's requirement of using a custom training script rather than the default built-in one.

Why this answer

The entry_point parameter specifies the path to the custom training script that the Hugging Face estimator will execute inside the container. Without it, the estimator has no user-supplied Python file to run, so it cannot perform custom fine-tuning logic. framework_version, instance_type, and hyperparameters configure the environment but do not point to the training code itself.

Exam trap

MLA-C01 often tests whether candidates confuse configuration parameters (framework_version, instance_type, hyperparameters) with the parameter that actually supplies the training code, so the trap is picking a familiar estimator argument that does not point to the script.

How to eliminate wrong answers

Option A is wrong because framework_version only selects the Hugging Face/PyTorch/TensorFlow container image version and does not identify any user script. Option B is wrong because instance_type only provisions compute capacity and has no bearing on which training code runs. Option C is wrong because hyperparameters are passed as arguments to the script but do not define the script's location or existence.

87
MCQmedium

A machine learning engineer is training a TensorFlow model using SageMaker with distributed training. They need to implement data parallelism across multiple GPUs. Which SageMaker feature should they use to distribute the training?

A.SageMaker Distributed Data Parallelism
B.SageMaker Automatic Model Tuning
C.SageMaker Debugger
D.SageMaker Model Parallelism
AnswerA

SageMaker Distributed Data Parallelism splits each mini-batch across GPUs and uses AllReduce for gradient synchronisation, scaling TensorFlow training across multiple GPUs. It satisfies the stem's requirement for data parallelism specifically, unlike model parallelism, which partitions the model itself rather than replicating it per device.

Why this answer

SageMaker Distributed Data Parallelism is the correct choice because it is purpose-built to distribute training across multiple GPUs by replicating the model on each GPU and splitting the input data across them. It uses the AllReduce algorithm with custom communication optimizations (e.g., balanced fusion of gradients) to synchronize gradients efficiently, reducing communication overhead compared to standard Horovod or torch.distributed. This feature is integrated into the SageMaker training toolkit and works with TensorFlow, PyTorch, and MXNet, making it the right tool for data parallelism.

Exam trap

MLA-C01 often tests the distinction between data parallelism and model parallelism, and candidates may confuse SageMaker Distributed Data Parallelism with SageMaker Model Parallelism, or mistakenly think that Automatic Model Tuning or Debugger can distribute training.

How to eliminate wrong answers

Option B is wrong because SageMaker Automatic Model Tuning (AMT) is a hyperparameter optimization service that searches for the best hyperparameters, not a distributed training mechanism. Option C is wrong because SageMaker Debugger is a monitoring and debugging tool that captures tensors and metrics during training, but it does not distribute training across GPUs. Option D is wrong because SageMaker Model Parallelism splits a single model across multiple GPUs (model parallelism), which is used when a model is too large to fit on one GPU, not for data parallelism where the model is replicated and data is split.

88
MCQmedium

A team is fine-tuning a large language model (LLM) using SageMaker and wants to reduce memory footprint during training. Which technique should they use?

A.Use LoRA (Low-Rank Adaptation) with fp32 precision
B.Use QLoRA (Quantized Low-Rank Adaptation) with 4-bit quantization
C.Use SageMaker Model Parallelism with tensor parallelism
D.Full fine-tuning on a p3.16xlarge instance
AnswerB

QLoRA quantises the frozen base model to 4-bit and trains only low-rank adapters, sharply reducing GPU memory versus full or standard LoRA fine-tuning. This directly satisfies the memory-footprint constraint while retaining most task accuracy for the LLM fine-tuning job.

Why this answer

QLoRA combines 4-bit quantization of the base model weights (via NF4) with LoRA adapters, dramatically reducing GPU memory during fine-tuning while preserving accuracy close to full fine-tuning. The 4-bit quantization is the key memory-saving mechanism, since the frozen base weights consume ~4x less memory than fp16/fp32. LoRA alone reduces optimizer/gradient memory but still loads the base model in higher precision, so QLoRA is the correct answer for minimizing memory footprint.

Exam trap

MLA-C01 often tests the distinction between parameter-efficient fine-tuning (LoRA) and memory-optimized fine-tuning (QLoRA), tricking candidates into selecting plain LoRA when the question specifically emphasizes reducing memory footprint.

How to eliminate wrong answers

Option A is wrong because LoRA with fp32 precision still loads the full base model in 32-bit, so memory savings come only from the small adapter parameters — the dominant memory cost (base weights) is unchanged. Option C is wrong because model parallelism with tensor parallelism shards the model across GPUs to fit larger models, but it does not reduce per-GPU memory footprint the way quantization does and adds communication overhead. Option D is wrong because full fine-tuning on a p3.16xlarge (8x V100 16GB) requires storing full gradients, optimizer states, and activations for all parameters, which is the most memory-intensive approach, not a reduction technique.

89
MCQeasy

A data scientist is using SageMaker Automatic Model Tuning to find the best hyperparameters for a model. They want to reduce the total tuning time for a given number of training jobs. Which tuning strategy should they choose?

A.Hyperband
B.Grid search
C.Random search
D.Bayesian optimization
AnswerA

Hyperband allocates resources adaptively, terminating poorly performing trials early and reallocating budget to promising configurations. This reduces total tuning time for a fixed number of training jobs, unlike grid or random search, which run every trial to completion.

Why this answer

Hyperband is an early-stopping, multi-fidelity tuning strategy that allocates resources to promising configurations and terminates poor performers early, dramatically reducing total tuning time. It is especially effective when many training jobs are needed but only a few hyperparameter combinations are worth full training. This makes it the best choice when the goal is to reduce total tuning time for a fixed number of jobs.

Exam trap

MLA-C01 often tests the distinction between sample efficiency (Bayesian) and time efficiency (Hyperband), tricking candidates into choosing Bayesian optimization when the question emphasizes reducing total tuning time.

How to eliminate wrong answers

Option B is wrong because grid search exhaustively trains every combination, which is the slowest strategy and scales poorly with dimensionality. Option C is wrong because random search samples combinations without early stopping, so it does not reduce total time as aggressively as Hyperband. Option D is wrong because Bayesian optimization is sample-efficient in terms of finding good hyperparameters but does not inherently reduce wall-clock tuning time for a given number of jobs the way Hyperband's early stopping does.

90
MCQmedium

A data scientist is using SageMaker Automatic Model Tuning with Hyperband. They want to stop poorly performing trials early to save resources. Which strategy does Hyperband use?

A.Grid search
B.Random search
C.Successive Halving
D.Bayesian optimization
AnswerC

Successive Halving runs many configurations with small resource budgets, then repeatedly discards the worst half and doubles resources for survivors. Hyperband wraps this in multiple brackets with varying budgets, so poor trials terminate early and compute is redirected to promising candidates.

Why this answer

Hyperband is a hyperparameter tuning strategy that uses Successive Halving to allocate resources efficiently. It starts many trials with small resource budgets, then iteratively promotes the best-performing trials to larger budgets while stopping poor performers early, saving compute.

Exam trap

MLA-C01 often tests the association between Hyperband and Successive Halving; candidates may confuse it with Bayesian optimization because both are advanced tuning methods.

How to eliminate wrong answers

Option A is wrong because grid search exhaustively tries all combinations and does not stop trials early. Option B is wrong because random search samples hyperparameters randomly but does not use early stopping based on performance. Option D is wrong because Bayesian optimization builds a probabilistic model to choose promising hyperparameters, but it is not the strategy Hyperband uses for early stopping.

91
MCQmedium

A machine learning team at a bank is training a binary classification model using SageMaker's built-in XGBoost algorithm on a dataset with 20 million rows and 300 features. They need to reduce training time while maintaining model accuracy. The data is stored in Amazon S3 as CSV files. Which approach should they take to speed up training?

A.Use SageMaker Processing to preprocess the data into a single large CSV file and then train with File mode.
B.Use SageMaker Pipe mode with the RecordIO protobuf format instead of File mode with CSV.
C.Increase the number of instances in the training cluster and enable distributed training with the parameter_distribution parameter set to 'fully'.
D.Convert the CSV files to TFRecord format and use File mode to load the data.
AnswerB

Pipe mode streams data directly from S3 to the training container, eliminating the need to download the full dataset to disk. RecordIO protobuf is a compact binary format that reduces I/O overhead and allows efficient shuffling. For large datasets with many features, this significantly speeds up training and reduces disk usage, while maintaining accuracy by feeding all data.

Why this answer

SageMaker's built-in XGBoost can consume data in Pipe mode with RecordIO protobuf, which streams data directly from S3 and reduces I/O overhead, leading to faster training on large datasets. File mode with CSV requires downloading all data to disk, which is slower. Distributed training options for XGBoost are limited and not configured via parameter_distribution.

Thus, switching to Pipe mode with RecordIO is the recommended approach.

Exam trap

The trap here is assuming that simply adding more instances or preprocessing data will speed up training, when the bottleneck is often data loading from S3.

92
Multi-Selectmedium

A machine learning engineer is using SageMaker Debugger to monitor a training job and wants to detect issues early. The engineer wants to receive alerts when the training job is likely to fail due to vanishing gradients and when the loss is not decreasing. Which two actions should the engineer take to achieve this? (Choose two.)

Select 2 answers
A.Configure a built-in rule for vanishing gradients, such as the VanishingGradient rule.
B.Use SageMaker Clarify to detect bias in the gradients.
C.Set up a custom rule using a Python script that analyzes gradients and loss.
D.Enable Debugger's LossNotDecreasing rule to monitor the loss curve.
E.Configure the training job to save all tensors to Amazon S3 for manual analysis.
AnswersA, D

SageMaker Debugger provides built-in rules like VanishingGradient that analyze tensor outputs during training to detect when gradients become very small. This rule can trigger an alert or stop the training job if vanishing gradients are detected, allowing the engineer to address the issue early. It is specifically designed for this purpose and requires minimal configuration.

Why this answer

The VanishingGradient and LossNotDecreasing built-in rules in SageMaker Debugger are designed to automatically detect the specified training issues and trigger alerts. They are easy to configure and provide real-time monitoring, which is exactly what the engineer needs. Custom rules or other services would require more effort or are not suited for this purpose.

Exam trap

The trap here is overlooking the availability of built-in rules and instead considering custom development or unrelated services like Clarify.

93
MCQmedium

A company wants to use SageMaker Autopilot for a regression problem. They require an explainability report that shows feature importance globally. Which Autopilot feature should they enable?

A.AutoML candidate generation
B.Ensembling mode
C.Hyperparameter optimization
D.Explainability report generation
AnswerD

Enabling explainability report generation makes Autopilot produce a SageMaker Clarify report alongside each candidate model, quantifying global feature importance via SHAP values for the regression task. This directly satisfies the stated requirement for a global feature-importance explainability report, which Autopilot does not generate by default.

Why this answer

SageMaker Autopilot's explainability report generation feature produces SHAP-based feature importance values that quantify each input feature's global contribution to model predictions. Enabling it during the AutoML job causes Autopilot to emit a model explainability report alongside the best candidate, satisfying the requirement for a global feature-importance view. The other options relate to candidate exploration, model combination, or tuning, none of which produce explainability artifacts.

Exam trap

MLA-C01 often tests whether candidates confuse model-quality features (ensembling, HPO) with model-interpretability features (Clarify/explainability), so the trap is picking a tuning option when the question asks for explainability.

How to eliminate wrong answers

Option A is wrong because AutoML candidate generation only explores and trains pipeline candidates; it does not produce feature-importance reports. Option B is wrong because ensembling mode combines multiple models to improve accuracy, not to explain feature contributions. Option C is wrong because hyperparameter optimization tunes training parameters and has no bearing on generating explainability output.

94
MCQeasy

A machine learning engineer needs to reduce costs when training a large model on SageMaker. They are willing to accept potential interruptions and have checkpointing enabled. Which instance purchasing option should they use?

A.Spot instances
B.Reserved instances
C.Dedicated hosts
D.On-demand instances
AnswerA

Spot instances exploit unused EC2 capacity at steep discounts, satisfying the cost-reduction constraint. Because the engineer accepts interruptions and checkpointing is enabled, training resumes from the last saved state after a two-minute interruption notice. This suits fault-tolerant workloads, unlike On-Demand or Reserved instances, which guarantee capacity at higher prices.

Why this answer

Spot instances offer significant cost savings (up to 60-90%) compared to on-demand, but can be reclaimed by AWS with a 2-minute notice. Checkpointing allows resuming training from the last saved state, making spot instances suitable.

95
Multi-Selecthard

A machine learning engineer is evaluating a binary classification model that predicts customer churn. The model achieves 95% accuracy, but the engineer suspects class imbalance is causing a misleading metric. Which THREE evaluation steps should the engineer perform to properly assess the model? (Choose THREE.)

Select 3 answers
A.Calculate RMSE
B.Calculate precision, recall, and F1-score
C.Compute Mean Absolute Error (MAE)
D.Plot the ROC curve and compute AUC
E.Compute the confusion matrix
AnswersB, D, E

Precision, recall, and F1-score expose performance on the minority churn class that accuracy conceals under imbalance. Recall reveals missed churners, precision captures false alarms, and F1-score balances both, directly satisfying the need to assess the model beyond a misleading 95% accuracy figure.

Why this answer

Option B is correct because precision, recall, and F1-score are threshold-based classification metrics that reveal how well the model identifies the minority churn class, unlike accuracy which can be inflated by class imbalance. Option D is correct because the ROC curve plots True Positive Rate against False Positive Rate across thresholds and AUC summarizes discriminative ability independently of a single threshold, making it robust for imbalanced binary classification. Option E is correct because the confusion matrix breaks predictions into TP, TN, FP, and FN counts, exposing exactly how many churners are missed or falsely flagged and serving as the basis for precision, recall, and F1.

Options A and C are incorrect because RMSE and MAE are regression error metrics measuring continuous prediction deviations, not appropriate for evaluating a binary classification model's class-imbalanced performance.

Exam trap

MLA-C01 often tests the misconception that high accuracy implies a good model, especially with imbalanced datasets; candidates must recognize that accuracy is not a reliable metric for classification with class imbalance.

96
MCQmedium

A data scientist is training an object detection model using SageMaker built-in Object Detection algorithm. They want to visualize the bounding boxes on validation images after training. Which approach should they use?

A.Use SageMaker Debugger to capture output tensors
B.Write a custom inference script that saves images with bounding boxes
C.Enable SageMaker Model Monitor
D.Use SageMaker Clarify
AnswerB

The built-in Object Detection algorithm emits bounding box coordinates as JSON output, not annotated images. A custom inference script must parse those predictions and draw boxes onto the validation images, since no built-in visualisation step exists.

Why this answer

SageMaker's built-in Object Detection algorithm outputs predictions in JSON (bounding boxes, class labels, scores) via inference; to visualize boxes on validation images, you must write a custom inference script that reads those predictions and draws rectangles on the images. No built-in visualization step exists in the algorithm itself. The other options address debugging, monitoring, or explainability, not bounding-box rendering.

Exam trap

MLA-C01 often tests whether candidates assume built-in algorithms include visualization tooling, so they pick Debugger, Clarify, or Model Monitor instead of recognizing that custom inference code is required.

How to eliminate wrong answers

Option A is wrong because SageMaker Debugger captures tensors and metrics for training diagnostics, not rendered bounding-box images. Option C is wrong because Model Monitor detects data drift and quality issues in production, not training-time visualization. Option D is wrong because SageMaker Clarify produces bias and explainability reports, not bounding-box overlays on images.

97
MCQhard

An ML engineer is fine-tuning a large language model using LoRA on SageMaker. The training is converging slowly, and GPU utilization is low. The engineer suspects the bottleneck is data loading. Which action should the engineer take to improve GPU utilization?

A.Increase the batch size to maximize GPU memory usage
B.Enable checkpointing and use spot instances
C.Use SageMaker Pipe mode to stream data from S3 directly to the training instances
D.Reduce model parallelism to decrease communication overhead
AnswerC

Pipe mode streams training data directly from Amazon S3 to the instances, removing the download-and-store step that starves the GPU. This raises input throughput and GPU utilisation, addressing the data-loading bottleneck the engineer suspects during LoRA fine-tuning.

Why this answer

SageMaker Pipe mode streams training data directly from Amazon S3 to the training container over a high-throughput channel, bypassing the local disk and the download-then-read pattern of File mode. This removes the I/O bottleneck that starves the GPU when the dataset is large or when many small files cause slow reads. With faster data delivery, the GPU spends less time idle waiting for batches, raising utilization and speeding convergence.

Exam trap

The trap is treating GPU underutilization as a compute or memory problem — candidates reach for batch size or parallelism changes, but low GPU utilization with slow convergence almost always points to an input pipeline bottleneck that Pipe mode is designed to solve.

How to eliminate wrong answers

Option A is wrong because increasing batch size does not fix a data-loading bottleneck — if the input pipeline cannot feed the current batch size fast enough, a larger batch simply makes each wait longer and may cause GPU OOM without improving utilization. Option B is wrong because checkpointing and spot instances address cost and fault tolerance, not data throughput; spot instances can even be interrupted, and checkpointing adds I/O overhead rather than removing it. Option D is wrong because reducing model parallelism targets communication overhead between GPUs, which is a different bottleneck; if the GPU is idle waiting for data, changing parallelism does not help and may hurt convergence.

98
MCQmedium

A team is training a large language model and needs to split the model layers across multiple GPUs due to memory constraints. Which distributed training strategy should they use?

A.Data parallelism
B.Hyperparameter tuning
C.Autopilot
D.Model parallelism
AnswerD

Model parallelism partitions the model's layers across multiple GPUs, so each device holds only a subset of weights. This directly resolves the memory constraint described, where a single GPU cannot hold the entire large language model, unlike data parallelism which replicates the full model on every device.

Why this answer

Model parallelism splits the model's layers (or tensors) across multiple GPUs so that each GPU holds only a portion of the model's parameters, which is required when the model is too large to fit in a single GPU's memory. Data parallelism, by contrast, replicates the full model on every GPU and only partitions the training data, so it does not solve a memory-constraint problem.

Exam trap

MLA-C01 often tests the distinction between data parallelism (replicate model, split data) and model parallelism (split model, replicate data) — candidates pick data parallelism by default because it is more common, missing the memory-constraint keyword.

How to eliminate wrong answers

Option A is wrong because data parallelism replicates the entire model on each GPU and only shards the input data, so it does not reduce per-GPU memory footprint and cannot fit a model that exceeds single-GPU memory. Option B is wrong because hyperparameter tuning is an optimization/search process (e.g., SageMaker Automatic Model Tuning) and has nothing to do with distributing model layers across devices. Option C is wrong because Autopilot is a SageMaker feature that automates ML workflow decisions (feature engineering, algorithm selection, tuning), not a distributed training strategy for splitting layers across GPUs.

99
MCQmedium

A team is training a large deep learning model on SageMaker using a single ml.p3.16xlarge instance. Training is taking too long. They want to reduce time by distributing across multiple GPUs but are constrained by model size that does not fit in a single GPU memory. Which distributed training strategy should they use?

A.Data parallelism using SageMaker distributed data parallelism
B.Switch to a smaller instance type and use horizontal scaling
C.Use multiple training jobs with hyperparameter tuning
D.Model parallelism using SageMaker distributed model parallelism
AnswerD

Model parallelism shards the model's layers and parameters across multiple GPUs, so weights that cannot fit in one GPU's memory are held collectively. This directly satisfies the stem's constraint that the model does not fit in a single GPU, whereas data parallelism would replicate the full model on every device.

Why this answer

Model parallelism splits the model across multiple GPUs, which is needed when the model does not fit in a single GPU. Data parallelism replicates the model on each GPU and splits data, which requires the model to fit in each GPU's memory.

100
MCQeasy

Which SageMaker feature provides AutoML capabilities, including automatic data preprocessing, model selection, and hyperparameter tuning?

A.SageMaker Data Wrangler
B.SageMaker Automatic Model Tuning
C.SageMaker Autopilot
D.SageMaker Experiments
AnswerC

SageMaker Autopilot automates the full AutoML workflow: data preprocessing, algorithm selection, and hyperparameter tuning, then generates explainability reports. It directly satisfies the stem's requirement for automatic preprocessing, model selection, and tuning, unlike manual training jobs or built-in algorithms used alone.

Why this answer

SageMaker Autopilot is the only SageMaker feature that provides end-to-end AutoML: it automatically inspects and preprocesses raw tabular data, explores candidate algorithms, performs hyperparameter tuning, and selects the best model. It also generates explainability reports and notebooks so users can understand and reproduce the pipeline. The other options are narrower tools that address only one part of the ML lifecycle.

Exam trap

MLA-C01 often tests the distinction between SageMaker's specialized tools and its end-to-end AutoML service, so candidates may incorrectly choose Automatic Model Tuning because it also involves automation, but it only covers hyperparameter optimization, not the full AutoML pipeline.

How to eliminate wrong answers

Option A is wrong because SageMaker Data Wrangler is a data preparation and feature engineering tool that helps import, transform, and visualize data, but it does not perform model selection or hyperparameter tuning. Option B is wrong because SageMaker Automatic Model Tuning (AMT) only optimizes hyperparameters for a chosen algorithm using search strategies like Bayesian optimization; it does not handle data preprocessing or model selection. Option D is wrong because SageMaker Experiments is a tracking and organization service for ML runs, not an AutoML engine that builds models.

101
MCQeasy

Which SageMaker built-in algorithm is designed for time series forecasting?

A.Linear Learner
B.Factorisation Machines
C.DeepAR
D.BlazingText
AnswerC

DeepAR is the only SageMaker built-in algorithm purpose-built for time series forecasting, using autoregressive recurrent neural networks to model probability distributions across many related series. It satisfies the stem's forecasting constraint directly, unlike clustering, classification or regression algorithms, and supports both training from scratch and transfer learning on related datasets.

Why this answer

DeepAR is a SageMaker built-in algorithm specifically designed for time series forecasting. It uses recurrent neural networks (RNNs) to predict future values based on historical data, making it ideal for forecasting tasks such as demand prediction or sales forecasting.

Exam trap

The trap is confusing general-purpose algorithms like Linear Learner with specialized time series algorithms. Candidates might think any regression algorithm can forecast, but DeepAR is specifically designed for sequential data.

How to eliminate wrong answers

Option A is wrong because Linear Learner is a general-purpose supervised learning algorithm for classification and regression, not specialized for time series. Option B is wrong because Factorization Machines are used for recommendation systems and sparse data, not time series forecasting. Option D is wrong because BlazingText is for text classification and word embeddings, not time series.

102
Multi-Selectmedium

A data scientist is evaluating a binary classification model. They have the confusion matrix and want to assess the model's performance comprehensively. Which THREE metrics should they consider? (Select THREE.)

Select 3 answers
A.Precision
B.RMSE
C.Recall
D.F1 score
E.R²
AnswersA, C, D

Precision measures the proportion of positive predictions that are actually correct, derived from the confusion matrix's true-positive and false-positive counts. It is one of the three complementary metrics needed to assess a binary classifier comprehensively.

Why this answer

Precision (A) is correct because it measures the proportion of positive predictions that are actually correct (TP / (TP + FP)), which is essential for evaluating a binary classifier's reliability on predicted positives. Recall (C) is correct because it measures the proportion of actual positives that were correctly identified (TP / (TP + FN)), capturing the model's ability to find all relevant cases. F1 score (D) is correct because it is the harmonic mean of precision and recall (2 × (Precision × Recall) / (Precision + Recall)), giving a balanced single metric when both false positives and false negatives matter.

RMSE (B) is not appropriate here because it is a regression error metric measuring the square root of the average squared difference between predicted and actual continuous values, not classification outcomes. R² (E) is also a regression metric that quantifies the proportion of variance explained in a continuous target, so it does not apply to a binary classification confusion matrix.

Exam trap

MLA-C01 often tests the confusion between classification and regression metrics, so candidates may incorrectly select RMSE or R² for a classification task.

103
MCQmedium

A machine learning engineer is training a model using SageMaker and wants to set up monitoring to detect if gradients become too large, which could destabilize training. Which SageMaker Debugger built-in rule should they enable?

A.DeadRelu
B.LossNotDecreasing
C.Overfit
D.ExplodingGradients
AnswerD

ExplodingGradients monitors gradient magnitudes during training and raises an alert when values exceed a threshold, directly detecting the instability described. It satisfies the requirement to catch gradients becoming too large before they destabilise training, unlike rules targeting vanishing gradients, overfitting, or loss convergence.

Why this answer

The ExplodingGradients built-in rule in SageMaker Debugger is specifically designed to detect when gradient values become excessively large during training, which can cause numerical instability and prevent convergence. It monitors the gradients tensor and triggers if the ratio of the maximum absolute gradient to the average absolute gradient exceeds a threshold (default 10.0). This directly addresses the engineer's requirement to detect destabilizing gradients.

Exam trap

MLA-C01 often tests the specific purpose of each SageMaker Debugger built-in rule, and candidates may confuse ExplodingGradients with LossNotDecreasing or DeadRelu, which address different training issues.

How to eliminate wrong answers

Option A is wrong because DeadRelu detects when ReLU activations output zero for all inputs, indicating dead neurons, not large gradients. Option B is wrong because LossNotDecreasing monitors the training loss and alerts if it fails to decrease over time, which is a symptom of poor learning, not specifically gradient explosion. Option C is wrong because Overfit detects when validation loss increases while training loss decreases, indicating overfitting, which is unrelated to gradient magnitude.

104
MCQmedium

A data scientist needs to evaluate a binary classification model. The dataset is highly imbalanced (5% positive class). Which metric is MOST appropriate for assessing model performance?

A.Precision
B.Accuracy
C.Recall
D.AUC
AnswerD

AUC measures ranking quality across all thresholds and stays insensitive to class prevalence, so the 5% positive rate does not distort it. Accuracy would mislead here, since predicting all negatives scores 95%. AUC therefore satisfies the stem's imbalanced-classification constraint by summarising separability between classes.

Why this answer

AUC (Area Under the ROC Curve) evaluates the model's ability to rank positive instances above negatives across all classification thresholds, making it robust to class imbalance because it does not depend on a single threshold or on the majority class dominating the score. With only 5% positives, accuracy and threshold-dependent metrics can be misleading, but AUC remains a reliable discriminator measure.

Exam trap

The trap is defaulting to accuracy because it is the most familiar metric — on a 95/5 imbalance, accuracy is dominated by the majority class and hides poor positive-class performance.

How to eliminate wrong answers

Option A is wrong because precision alone ignores recall and is threshold-dependent — a model can achieve high precision by predicting very few positives, which is misleading on imbalanced data. Option B is wrong because accuracy is dominated by the majority class: a model that predicts 'negative' for every instance achieves 95% accuracy on a 5% positive dataset while being useless. Option C is wrong because recall alone ignores false positives and is also threshold-dependent; high recall can be achieved by predicting everything positive, which is not a meaningful performance measure on its own.

105
MCQeasy

A data scientist wants to use SageMaker Autopilot to automatically build a regression model. The dataset contains 200 features and 50,000 rows. Which output does SageMaker Autopilot provide?

A.Only the best model without any metrics
B.A leaderboard of candidate models with metrics and explainability reports
C.A single optimal model with no further tuning
D.A Python script for manual training
AnswerB

SageMaker Autopilot explores preprocessing and algorithms automatically, then returns a leaderboard ranking candidate models by objective metric, alongside notebooks and explainability reports. This satisfies the requirement to automatically build a regression model from 200 features and 50,000 rows.

Why this answer

SageMaker Autopilot automatically explores different algorithms and hyperparameter combinations to find the best model for a given dataset. It provides a leaderboard of candidate models, each with performance metrics, and generates explainability reports that show how features influence predictions. This allows data scientists to understand and select the most suitable model for deployment.

Exam trap

MLA-C01 often tests the misconception that Autopilot only outputs a single best model, when in fact it provides a leaderboard and explainability reports, so candidates may overlook the comprehensive output.

How to eliminate wrong answers

Option A is wrong because Autopilot does not just output the best model; it provides a leaderboard with multiple candidates and their metrics. Option C is wrong because Autopilot does not produce a single optimal model without further tuning; it generates several candidates and allows further tuning. Option D is wrong because Autopilot does not generate a Python script for manual training; it automates the training process and provides notebooks for reproducibility, but the output is not a script for manual training.

106
MCQmedium

A data scientist is using SageMaker to train an XGBoost model for a regression problem. After training, they evaluate the model on a test set and get an RMSE of 10 and an R² of 0.85. Which additional metric would give the MOST insight into the model's average prediction error magnitude?

A.AUC
B.Confusion matrix
C.F1 score
D.Mean Absolute Error (MAE)
AnswerD

MAE reports the average absolute difference between predicted and actual values in the target's own units, directly quantifying typical prediction error magnitude. Unlike RMSE, which squares errors and overweights outliers, MAE gives the straightforward average error the stem requests, complementing the existing R² of 0.85.

Why this answer

MAE measures the average absolute difference between predicted and actual values in the same units as the target, so it directly answers 'how far off is the model on average?' RMSE of 10 is also in target units but is dominated by large errors due to squaring, so MAE complements it by showing typical error magnitude without outlier amplification.

Exam trap

The trap is picking a familiar classification metric (AUC, F1, confusion matrix) for a regression problem; always check whether the target is continuous before selecting metrics.

How to eliminate wrong answers

Option A is wrong because AUC is a classification metric measuring the trade-off between true positive and false positive rates across thresholds, and is undefined for continuous regression targets. Option B is wrong because a confusion matrix is a classification diagnostic (TP/FP/TN/FN counts) and has no meaning for a regression problem predicting continuous values. Option C is wrong because F1 score is the harmonic mean of precision and recall, again a classification-only metric that cannot be computed on continuous predictions.

107
MCQmedium

A machine learning engineer is preparing a training script that must run on multiple GPU instances with SageMaker. The script currently reads the entire training dataset from local disk into memory, which fails on larger datasets. The engineer wants the script to stream training data from the SageMaker training channel path without loading everything into memory. Which approach should the engineer take?

A.Use a framework-provided dataset or iterator that reads from the path in SM_CHANNEL_TRAIN and yields batches lazily during training.
B.Set the estimator's volume_size parameter large enough to hold the dataset and enable cacheutils to copy the data into the container's memory before training starts.
C.Read the dataset directly from the Amazon S3 URI passed in the SM_CHANNEL_TRAIN environment variable using the AWS SDK for Python (Boto3) inside the training loop.
D.Configure the estimator with an Amazon EFS file system and mount the dataset into the container, then use memory-mapped file access for all training samples.
AnswerA

SageMaker copies channel data to a local directory and exposes the path through the SM_CHANNEL_TRAIN environment variable. Framework integrations such as PyTorch Dataset/DataLoader, TensorFlow tf.data, or the SageMaker training toolkit input modules can read from that path lazily, so only the batches needed for each step are materialized. This avoids loading the full dataset into memory.

Why this answer

SageMaker downloads the contents of each training channel to a local path and exposes that path through the SM_CHANNEL_TRAIN environment variable. Using a framework-native dataset or iterator that reads lazily from that path streams batches on demand, so memory stays bounded regardless of dataset size. Direct S3 reads, EFS mounts, and volume resizing do not address in-memory loading.

Exam trap

The trap here is assuming SM_CHANNEL_TRAIN holds an S3 URI and that Boto3 streaming is the intended pattern, when it actually points to a local container path.

108
MCQmedium

A data scientist wants to fine-tune a large language model for a question-answering task. They want to reduce memory usage during training by using a low-rank approximation of the weight updates. Which technique should they use?

A.Full fine-tuning
B.Instruction tuning
C.LoRA
D.RLHF
AnswerC

LoRA freezes the pretrained weights and injects trainable low-rank decomposition matrices into each layer, so only these small matrices receive gradient updates. This directly satisfies the stem's constraint of reducing memory usage via a low-rank approximation of weight updates during fine-tuning.

Why this answer

LoRA (Low-Rank Adaptation) adds low-rank matrices to model weights, significantly reducing memory footprint while achieving competitive performance. QLoRA adds quantization for further reduction.

← PreviousPage 2 of 2 · 108 questions total

Ready to test yourself?

Try a timed practice session using only Mla Model Development questions.