Courseiva

CCNA Machine Learning and Deep Learning Questions

75 of 95 questions · Page 1/2 · Machine Learning and Deep Learning · Answers revealed

1
Multi-Selectmedium

A team is designing a deep learning pipeline for a computer vision task. They want to reduce overfitting. Which two techniques are specifically effective for this purpose? (Select TWO.)

Select 2 answers
A.Dropout
B.Using a smaller batch size
C.Adding more layers
D.L2 weight regularization
E.Increasing the learning rate
AnswersA, D

Dropout randomly deactivates neurons, reducing overfitting by preventing reliance on specific features.

Why this answer

Dropout is a regularization technique that randomly drops a fraction of neurons during training, which prevents the network from relying too heavily on any single neuron and forces it to learn more robust features. This reduces overfitting by introducing noise that improves generalization.

Exam trap

This exam often tests the misconception that increasing model capacity (more layers) or adjusting batch size directly reduces overfitting, when in fact these changes typically require additional regularization to be effective.

2
MCQeasy

Refer to the exhibit. What is the recall of the model?

A.0.44
B.0.80
C.0.90
D.0.99
AnswerA

Recall equals true positives divided by the sum of true positives and false negatives. Reading the confusion matrix, 0.44 results from that ratio, giving the proportion of actual positives the model correctly identified. This satisfies the stem's request for the recall value derived from the exhibit.

Why this answer

Recall is calculated as True Positives divided by (True Positives + False Negatives). From the confusion matrix in the exhibit, True Positives = 400 and False Negatives = 500, so recall = 400 / (400 + 500) = 400 / 900 ≈ 0.44. Option A is correct because this matches the computed recall value.

Exam trap

CompTIA often tests the confusion between recall and precision, so candidates mistakenly compute precision (TP/(TP+FP)) instead of recall, leading them to choose 0.80.

How to eliminate wrong answers

Option B (0.80) is wrong because it likely results from incorrectly using precision (TP/(TP+FP) = 40/50 = 0.80) instead of recall. Option C (0.90) is wrong because it may come from dividing TP by total predictions (40/100 = 0.40) or misreading the matrix, but 0.90 is not supported by any standard metric from the given values. Option D (0.99) is wrong because it is far too high and could stem from confusing recall with accuracy or ignoring the false negatives entirely.

3
MCQmedium

A machine learning engineer trains a decision tree to predict customer churn. The tree achieves 99 percent accuracy on the training set but only 68 percent on a held-out test set. The engineer wants to reduce this gap. Which single action is most likely to improve test performance?

A.Limit the tree depth and require a minimum number of samples per leaf
B.Train the tree on the test set as well so it learns the held-out distribution
C.Increase the maximum depth of the tree so it can fit more training examples
D.Add more features derived from the training set to give the tree more signal
AnswerA

Constraining depth and requiring a minimum samples per leaf restricts the tree from creating tiny, noise-driven splits. This regularization reduces variance and typically narrows the train-test gap. For a tree already at 99 percent training accuracy, these hyperparameters directly address the overfitting that causes the 31-point drop on held-out data.

Why this answer

The large train-test gap signals high variance, meaning the unconstrained decision tree memorized training noise. Pre-pruning through maximum depth and minimum samples per leaf limits how finely the tree can split, reducing variance and improving held-out accuracy without discarding the model or leaking test data.

Exam trap

The trap here is equating higher training accuracy with a better model, when a near-perfect training score alongside poor test performance is the classic signature of overfitting.

4
Multi-Selectmedium

Which TWO techniques are commonly used to handle missing data in a machine learning dataset? (Choose TWO.)

Select 2 answers
A.Normalization
B.Imputation with mean or median
C.Deletion of rows with missing values
D.One-hot encoding
E.Dimensionality reduction
AnswersB, C

Replacing missing values with mean/median is a common imputation method.

Why this answer

Imputation with mean or median is a standard technique for handling missing numerical data because it preserves the dataset size and avoids introducing bias from simply discarding rows. By replacing missing values with the central tendency of the observed data, the model can still learn patterns without losing information, though it may reduce variance slightly.

Exam trap

CompTIA often tests the distinction between data preprocessing techniques (like normalization and encoding) and actual missing data handling methods, so candidates mistakenly select normalization or one-hot encoding as solutions for missing values.

5
MCQeasy

A data scientist is building a classification model to detect fraudulent transactions. The dataset is highly imbalanced with only 1% fraudulent cases. Which approach should the scientist use to evaluate model performance most effectively?

A.F1 score
B.Accuracy
C.Recall
D.Precision
AnswerA

F1 score suits this imbalanced fraud scenario because it combines precision and recall into a single harmonic mean, so strong performance on the 1% fraudulent minority cannot be masked by the 99% legitimate majority. Accuracy would mislead here, since predicting every transaction as legitimate already yields 99%.

Why this answer

In highly imbalanced datasets like fraud detection (1% positive class), accuracy is misleading because a model that predicts all transactions as legitimate would achieve 99% accuracy yet fail to detect any fraud. The F1 score (harmonic mean of precision and recall) is the most effective metric because it balances both false positives and false negatives, providing a single score that reflects the model's ability to correctly identify the minority class without being skewed by class imbalance.

Exam trap

CompTIA often tests the misconception that accuracy is always the best metric for classification, but in imbalanced datasets, accuracy is a trap because it does not reflect performance on the minority class, leading candidates to overlook metrics like F1 score that directly address class imbalance.

How to eliminate wrong answers

Option B (Accuracy) is wrong because it is dominated by the majority class (99% legitimate transactions), so a trivial model that never predicts fraud can still achieve 99% accuracy, masking poor fraud detection performance. Option C (Recall) is wrong because it only measures the proportion of actual fraud cases correctly identified (true positives / (true positives + false negatives)), ignoring false positives; a model that flags every transaction as fraud would have perfect recall but be unusable in practice. Option D (Precision) is wrong because it only measures the proportion of predicted fraud cases that are actually fraud (true positives / (true positives + false positives)), ignoring false negatives; a model that makes very few fraud predictions but with high precision would miss many actual frauds, which is unacceptable in fraud detection.

6
MCQmedium

An e-commerce company uses a gradient boosting model to forecast daily sales. Recently, the model's predictions have become less accurate, showing a significant drop in R-squared on validation data. The data scientist checks for data drift but finds no significant changes in feature distributions. The model was trained on data from the past 24 months and is retrained monthly. Upon inspecting the feature importance, the data scientist notices that the top feature 'promotion_flag' has decreased in importance over time. What is the most likely cause of the performance degradation, and what should be done?

A.The model is overfitting to historical promotions; apply more regularization
B.Concept drift has occurred; retrain the model more frequently with recent data only, or use an online learning approach
C.The model's hyperparameters need tuning; perform a grid search
D.The promotion_flag feature is leaking future information; remove it
AnswerB

Stable feature distributions but a shifting feature-target relationship indicate concept drift, not data drift. The declining importance of promotion_flag shows the old mapping no longer holds, so retraining on recent data or adopting online learning restores accuracy.

Why this answer

The model's performance degradation, despite no data drift in feature distributions, is likely due to concept drift—the relationship between features and the target variable has changed over time. The decreasing importance of 'promotion_flag' suggests that promotions no longer influence sales as they once did. Retraining more frequently with recent data or using online learning can help the model adapt to the new concept.

Exam trap

The trap is that candidates may confuse concept drift with data drift or overfitting, but the key clue is the change in feature importance over time without changes in feature distributions, pointing to concept drift rather than other issues.

How to eliminate wrong answers

Option A is wrong because overfitting would typically show good performance on training data but poor on validation; here, the issue is a drop in validation R-squared over time, not necessarily overfitting. Option C is wrong because hyperparameter tuning would not address a change in the underlying data relationship; it might yield marginal improvements but not solve concept drift. Option D is wrong because feature leakage would cause overly optimistic performance during training, not a gradual degradation; also, 'promotion_flag' is a legitimate feature, and removing it without evidence of leakage is not appropriate.

7
Multi-Selectmedium

A data science team is preparing a dataset of customer support tickets to train a supervised model that routes each ticket to the correct department. They have 40,000 tickets labeled with one of eight departments. Which TWO preprocessing steps are most appropriate before training? (Choose two.)

Select 2 answers
A.Split the data into training, validation, and test sets before fitting any preprocessing.
B.Delete all tickets shorter than 20 words to reduce noise.
C.Standardize the department labels to zero mean and unit variance.
D.Tokenize the ticket text and convert it to numerical vectors using an embedding or TF-IDF representation.
E.Apply one-hot encoding to the raw ticket text strings.
AnswersA, D

Holding out validation and test data before fitting vectorizers or scalers prevents information leakage from the evaluation sets into training. If TF-IDF vocabulary or normalization statistics are learned on the full dataset, reported performance becomes optimistic and unreliable. This step ensures the routing model is evaluated on genuinely unseen tickets.

Why this answer

Text classification requires converting raw strings into numeric features and preserving an honest evaluation split. Tokenization with TF-IDF or embeddings supplies the numerical representation, while splitting before fitting preprocessing prevents leakage. Together these steps prepare the ticket data so the routing model can be trained and fairly assessed on unseen examples.

Exam trap

The trap here is fitting the vectorizer on the entire dataset before splitting, which leaks vocabulary and statistics from the test set into training and inflates reported accuracy.

8
MCQhard

A data scientist is building a model to predict the likelihood of a patient having a rare disease. The dataset is highly imbalanced, with only 2% of patients having the disease. The data scientist trains a logistic regression model and achieves 98% accuracy, but the model predicts 'no disease' for all patients. Which evaluation metric should the data scientist use to better assess the model's performance?

A.Mean Squared Error (MSE)
B.Precision-Recall AUC (Area Under the Curve)
C.Accuracy
D.R-squared (R²)
AnswerB

Precision-Recall AUC is particularly useful for imbalanced datasets because it focuses on the positive class. It plots precision against recall at various thresholds, providing a comprehensive view of the model's ability to identify the rare disease without being overwhelmed by the large number of true negatives. This metric helps assess how well the model distinguishes the minority class.

Why this answer

Precision-Recall AUC is designed for imbalanced classification problems, focusing on the performance of the positive class. It provides a more informative picture than accuracy, which can be misleadingly high when the negative class dominates. This metric helps the data scientist understand the trade-off between precision and recall for the rare disease detection.

Exam trap

The trap here is relying on accuracy as the primary metric for imbalanced datasets, which can hide poor performance on the minority class.

9
MCQmedium

A deep learning model for sentiment analysis uses a softmax output layer. The hidden layers currently use tanh activation. Which activation function should replace tanh to mitigate vanishing gradients in deeper networks?

A.Sigmoid
B.Softmax
C.ReLU
D.Linear
AnswerC

ReLU outputs the input directly for positive values, giving a derivative of one and avoiding the saturation that drives tanh's gradient toward zero in deep stacks. This preserves gradient magnitude during backpropagation, directly mitigating the vanishing gradient problem in the deeper network.

Why this answer

ReLU (Rectified Linear Unit) is correct because it outputs zero for negative inputs and a positive linear slope for positive inputs, which avoids the saturation problem of tanh. In deeper networks, tanh gradients can vanish as activations approach ±1, slowing or halting learning. ReLU's non-saturating nature keeps gradients flowing for positive inputs, mitigating the vanishing gradient problem.

Exam trap

Candidates often mistakenly believe that any non-linear activation works equally well in deep networks, but the trap is that they may choose sigmoid because it is non-linear, ignoring its saturation-induced vanishing gradient problem in deeper architectures.

How to eliminate wrong answers

Option A is wrong because sigmoid also saturates at 0 and 1, causing vanishing gradients in deep networks, similar to tanh. Option B is wrong because softmax is typically used in the output layer for multi-class classification, not in hidden layers, and it saturates across classes, exacerbating vanishing gradients. Option D is wrong because a linear activation would collapse the network into a single linear transformation, removing the non-linearity needed to learn complex patterns in sentiment analysis.

10
MCQhard

A data scientist is training a random forest model on a large dataset and notices that the model is overfitting. Which hyperparameter adjustment is most likely to reduce overfitting?

A.Increase the maximum features
B.Decrease the maximum depth of trees
C.Decrease the minimum samples split
D.Increase the number of trees
AnswerB

Maximum depth controls how many splits each tree can make. Reducing it constrains tree complexity, limiting the model's ability to memorise training noise, which lowers variance and reduces overfitting. Increasing depth or adding trees would typically worsen the overfitting observed.

Why this answer

Decreasing the maximum depth of trees limits how deep each decision tree can grow, which reduces the model's capacity to learn overly specific patterns from the training data. This directly combats overfitting by enforcing simpler trees that generalize better to unseen data.

Exam trap

A common mistake is assuming that increasing the number of trees always reduces overfitting, but in random forests, while more trees reduce variance through averaging, they do not address the root cause of overfitting in individual trees. Decreasing maximum depth directly limits tree complexity.

How to eliminate wrong answers

Option A is wrong because increasing the maximum features (the number of features considered at each split) actually increases tree diversity and can reduce overfitting in some cases, but it is not the most direct adjustment for overfitting—it can also increase variance if set too high. Option C is wrong because decreasing the minimum samples split (the minimum number of samples required to split an internal node) allows trees to split on smaller subsets, which increases model complexity and exacerbates overfitting. Option D is wrong because increasing the number of trees generally improves stability and reduces variance due to averaging, but it does not directly reduce overfitting; in fact, more trees can sometimes memorize noise if individual trees are already overfit.

11
MCQhard

A company is building a computer vision system to detect defects in manufactured parts. They have 10,000 labeled images per class (defective and non-defective). They want to achieve high accuracy with limited computational resources. Which deep learning architecture and approach is most appropriate?

A.Train a custom CNN from scratch with many layers
B.Use a decision tree ensemble
C.Use a pre-trained VGG16 and fine-tune the last few layers
D.Use an RNN to process image sequences
AnswerC

Transfer learning with a pre-trained VGG16 exploits features already learned from large image datasets, so fine-tuning only the final layers reaches high defect-detection accuracy with far less training data and compute than training a deep network from scratch.

Why this answer

Using a pre-trained VGG16 and fine-tuning the last few layers leverages transfer learning, which is ideal when you have a moderate-sized labeled dataset (10,000 images per class) and limited computational resources. The pre-trained model already captures general visual features from ImageNet, so only the task-specific layers need to be trained, reducing training time and resource requirements while still achieving high accuracy.

Exam trap

CompTIA AI often tests the misconception that more layers or training from scratch always yields better accuracy, when in reality transfer learning with a pre-trained model is the most practical choice for moderate datasets and limited compute.

How to eliminate wrong answers

Option A is wrong because training a custom CNN from scratch with many layers would require a very large dataset (typically millions of images) and extensive computational resources to converge, and with only 10,000 images per class, it risks overfitting and poor generalization. Option B is wrong because decision tree ensembles are not designed for high-dimensional image data; they lack the spatial feature extraction capabilities of CNNs and would perform poorly on raw pixel inputs. Option D is wrong because RNNs are designed for sequential data (e.g., time series, text) and are not suitable for static image classification; they would ignore spatial structure and be computationally inefficient for this task.

12
MCQmedium

A team is developing a recommendation system for an e-commerce platform. They want to use collaborative filtering but are concerned about cold-start problems for new users. Which approach would best mitigate the cold-start problem?

A.Incorporate user demographic features as side information
B.Increase the number of latent factors in matrix factorization
C.Use a popularity-based baseline for all recommendations
D.Use only item-based collaborative filtering
AnswerA

Demographic side information lets the system infer preferences for users lacking interaction history, so recommendations are generated from attributes rather than behaviour. This directly mitigates the cold-start constraint, where collaborative filtering alone cannot compute similarity for new users.

Why this answer

Incorporating user demographic features as side information allows the collaborative filtering model to generate initial recommendations for new users based on their demographic profile, effectively addressing the cold-start problem. This approach uses content-based features to bootstrap the recommendation process until sufficient user interaction data is collected.

Exam trap

The CompTIA AI+ exam often tests the misconception that increasing model complexity or switching between collaborative filtering variants alone can solve the cold-start problem, when in fact the solution requires incorporating auxiliary data (side information) to bootstrap recommendations for new users.

How to eliminate wrong answers

Option B is wrong because increasing the number of latent factors in matrix factorization does not solve the cold-start problem; it only increases model complexity and may lead to overfitting without addressing the lack of user interaction data. Option C is wrong because using a popularity-based baseline for all recommendations ignores personalization entirely, which defeats the purpose of collaborative filtering and does not leverage user-specific signals even when they become available. Option D is wrong because using only item-based collaborative filtering still requires some user interaction data to compute item similarities; it does not inherently handle new users with no history.

13
MCQhard

A team trains a recurrent neural network to translate sentences averaging 60 words. During evaluation they notice that translations of the final words in long sentences are frequently wrong, while the opening words are translated accurately. Which architectural change best addresses this behavior?

A.Lower the learning rate and train for more epochs
B.Replace the recurrent layers with an attention-based transformer encoder-decoder
C.Increase the batch size used during training
D.Increase the number of recurrent layers stacked in the encoder
AnswerB

The described degradation at the end of long sequences is characteristic of vanishing gradients and limited memory in plain recurrent networks. Self-attention lets every output position attend directly to every input position in constant path length, so information from early and late tokens remains accessible. Transformers are the standard architecture for long-sequence translation and directly resolve this failure.

Why this answer

Recurrent networks compress the entire source sentence into a fixed-size hidden state and suffer vanishing gradients across many timesteps, which disproportionately harms recall of distant tokens. An attention-based transformer removes the sequential bottleneck, allowing direct token-to-token connections and preserving context for the final words of long sentences.

Exam trap

The trap here is assuming more layers, larger batches, or longer training can overcome a positional memory problem, when the limitation is architectural rather than an optimization or capacity issue.

14
MCQhard

A media company uses a natural language processing (NLP) model to classify news articles into topics. The model was trained on articles from 2015-2018. In 2023, the model's F1 score drops significantly. The data scientists find that the word embeddings no longer capture the meaning of some terms (e.g., 'covid', 'metaverse'). The model uses static word embeddings (Word2Vec) trained on the original corpus. Which solution BEST addresses the observed degradation? A. Replace static embeddings with contextual embeddings from a transformer model like BERT, then fine-tune the classifier. B. Retrain the static Word2Vec embeddings on a larger corpus from 2023. C. Apply data augmentation to the original training data by replacing words with synonyms. D. Increase the dimensionality of the static embeddings.

A.Retrain the static Word2Vec embeddings on a larger corpus from 2023.
B.Increase the dimensionality of the static embeddings.
C.Replace static embeddings with contextual embeddings from a transformer model like BERT, then fine-tune the classifier.
D.Apply data augmentation to the original training data by replacing words with synonyms.
AnswerC

Contextual embeddings dynamically represent words based on context, handling semantic shift effectively.

Why this answer

Contextual embeddings (e.g., BERT) capture meaning based on context, adapting to new uses of words like 'covid' meaning pandemic. Fine-tuning the classifier on new data would update the model. Option A (retraining static embeddings) might capture new word senses but still assigns a single vector per word, missing context.

Option B (increasing dimensionality) does not address the semantic shift. Option D (data augmentation) does not introduce new word meanings.

15
Multi-Selecteasy

A machine learning engineer is preparing to train a deep neural network for image classification. To avoid overfitting, which TWO techniques should the engineer apply? (Select TWO.)

Select 2 answers
A.Use dropout regularization.
B.Use data augmentation.
C.Increase the number of layers.
D.Remove all non-linear activation functions.
E.Reduce the training dataset size.
AnswersA, B

Dropout is a regularization technique that helps prevent overfitting by randomly dropping units.

Why this answer

Dropout regularization is a technique that randomly drops a fraction of neurons during training, which prevents the network from relying too heavily on any single neuron and reduces co-adaptation. This acts as a form of ensemble learning and significantly reduces overfitting by improving generalization.

Exam trap

The CompTIA AI+ exam often tests the misconception that increasing model complexity (like adding layers) or reducing data helps with overfitting, when in reality these actions worsen it, while regularization and data augmentation are the correct countermeasures.

16
MCQhard

A machine learning engineer is troubleshooting a recurrent neural network that fails to learn long-range dependencies in sequential data. The gradients are computed using backpropagation through time. Which phenomenon is most likely occurring, and what architectural change would best address it?

A.Underfitting; increase the number of time steps
B.Vanishing gradients; use LSTM or GRU units
C.Exploding gradients; apply gradient clipping
D.Overfitting; reduce the number of layers
AnswerB

Backpropagation through time multiplies many small Jacobian terms, so gradients shrink exponentially across long sequences, preventing the network from learning distant dependencies. LSTM and GRU units introduce gated cell states that carry information along a near-constant error path, preserving gradient magnitude over long ranges.

Why this answer

In standard RNNs, backpropagation through time (BPTT) multiplies gradients across many time steps, causing them to shrink exponentially (vanishing gradients). This prevents the network from learning long-range dependencies. LSTM or GRU units introduce gating mechanisms that preserve gradient flow over many time steps, directly solving this problem.

Exam trap

The trap here is that candidates may confuse 'failure to learn long-range dependencies' with exploding gradients, but the correct clue is the inability to capture distant patterns, not training instability or NaN losses.

How to eliminate wrong answers

Option A is wrong because increasing the number of time steps would exacerbate the vanishing gradient problem, not fix it; underfitting is not the core issue here. Option C is wrong because exploding gradients cause large gradient values and training instability, but the described symptom (failure to learn long-range dependencies) is classic vanishing gradients, not exploding gradients; gradient clipping addresses exploding gradients, not vanishing ones. Option D is wrong because overfitting is characterized by high variance and poor generalization, not by an inability to learn long-range patterns; reducing layers would reduce capacity and could worsen underfitting, not solve the gradient propagation issue.

17
MCQeasy

A data scientist trains a linear regression model on housing prices. The training error is low, but test error is high. What is the most likely issue?

A.Overfitting
B.Multicollinearity
C.Data leakage
D.Underfitting
AnswerA

Overfitting occurs when the model captures noise and training-specific patterns, fitting training data closely while generalising poorly. The gap between low training error and high test error is the defining symptom, indicating the model has learned idiosyncrasies rather than the underlying relationship.

Why this answer

Low training error combined with high test error is the classic symptom of overfitting, where the model has memorized the training data, including its noise, rather than learning the underlying patterns. This causes the model to perform poorly on unseen data, which is exactly what the high test error indicates.

Exam trap

CompTIA often tests the distinction between overfitting and underfitting by presenting a scenario where training error is low but test error is high, leading candidates to mistakenly choose underfitting because they focus only on the high test error without considering the low training error.

How to eliminate wrong answers

Option B (Multicollinearity) is wrong because while it can inflate the variance of coefficient estimates and make them unstable, it does not typically cause a large gap between low training error and high test error; the model can still fit the training data well. Option C (Data leakage) is wrong because data leakage usually results in overly optimistic performance on both training and test sets during evaluation, not a high test error after low training error. Option D (Underfitting) is wrong because underfitting produces high error on both the training and test sets, not low training error with high test error.

18
MCQeasy

A company uses linear regression to predict sales based on advertising spend. The model's residuals show a pattern of increasing variance as spend increases. Which assumption of linear regression is violated?

A.Normality
B.Homoscedasticity
C.Linearity
D.Independence
AnswerB

Increasing variance in residuals as advertising spend rises directly breaches homoscedasticity, which requires constant error variance across all predictor values. The stem's fan-shaped residual pattern is the textbook signature of heteroscedasticity, so this assumption is the one violated.

Why this answer

The pattern of increasing residual variance with higher advertising spend violates the assumption of homoscedasticity, which requires constant variance of errors across all levels of the independent variable. In linear regression, heteroscedasticity like this can lead to inefficient coefficient estimates and unreliable confidence intervals, often detected via a Breusch-Pagan test or residual plot analysis.

Exam trap

CompTIA AI exams often test the distinction between homoscedasticity and normality, trapping candidates who confuse residual variance patterns with residual distribution shape, especially when the question describes a 'fan' or 'cone' shape in the residual plot.

How to eliminate wrong answers

Option A is wrong because normality refers to the distribution of residuals being approximately normal, not the variance pattern; heteroscedasticity does not directly violate normality. Option C is wrong because linearity assumes a straight-line relationship between spend and sales, which is not indicated by changing variance; the residual pattern here concerns spread, not curvature. Option D is wrong because independence assumes errors are uncorrelated with each other, typically violated in time-series data or clustered samples, not by variance changes across the predictor range.

19
Multi-Selectmedium

An AI team is preparing a support-vector machine to classify handwritten digits. Before training, they want to apply preprocessing steps that help the linear kernel separate the classes more effectively and improve generalization. Which two steps are most appropriate? (Choose two.)

Select 2 answers
A.Remove all pixels with zero variance across the dataset
B.Increase the number of support vectors by loosening the margin
C.Apply a nonlinear kernel such as RBF while claiming to keep the linear kernel
D.Tune the C regularization parameter with cross-validation
E.Scale each pixel feature to a common range such as 0 to 1
AnswersD, E

C controls the trade-off between maximizing the margin and penalizing misclassification. Too large a C produces a narrow margin that fits noise and overfits; too small a C underfits. Because the appropriate value depends on the dataset, tuning C with cross-validation on the scaled features selects a model that generalizes better. This is a core step in preparing an SVM for digit classification and directly targets generalization.

Why this answer

SVMs are sensitive to feature scale and to the C hyperparameter, so scaling pixel features and tuning C via cross-validation are the two steps that most directly improve linear-kernel separation and generalization. Removing constant pixels is trivial cleanup, inflating support vectors is a misunderstanding, and switching kernels abandons the linear setup the team specified. Together, scaling and C tuning form the standard SVM preparation workflow.

Exam trap

The trap here is focusing on kernel or support-vector mechanics while overlooking that unscaled features and an untuned C parameter are the most common reasons an SVM underperforms.

20
MCQhard

A company uses a neural network for fraud detection. The dataset has 99% legitimate, 1% fraudulent. The model achieves 99% accuracy but fails to detect most frauds. Which metric should they focus on?

A.Precision
B.F1-score
C.Recall
D.AUC-ROC
AnswerC

Recall measures the proportion of actual frauds correctly identified, directly addressing the 99:1 class imbalance where accuracy is misleading. Optimising recall reduces false negatives, ensuring the minority fraudulent cases are detected rather than ignored by a model that predicts "legitimate" almost always.

Why this answer

Recall (sensitivity) measures the proportion of actual positives correctly identified. In this fraud detection scenario with 99% legitimate and 1% fraudulent transactions, a 99% accuracy can be achieved by simply predicting all transactions as legitimate, which yields 0% recall for the fraud class. Focusing on recall ensures the model captures the majority of fraudulent cases, addressing the critical failure to detect fraud despite high accuracy.

Exam trap

The AI0-001 exam often tests the misconception that high accuracy implies good model performance, especially in imbalanced datasets, leading candidates to overlook recall as the critical metric for detecting rare events like fraud.

How to eliminate wrong answers

Option A is wrong because precision measures the proportion of predicted positives that are actually positive; while important for avoiding false alarms, it does not directly address the failure to detect fraud (false negatives). Option B is wrong because F1-score is the harmonic mean of precision and recall; although it balances both, the primary issue here is low recall, so focusing on recall directly is more appropriate. Option D is wrong because AUC-ROC measures the model's ability to distinguish between classes across all thresholds, but it can be misleadingly high even when recall for the minority class is poor, especially in imbalanced datasets; it does not directly target the failure to detect fraud.

21
MCQmedium

While training a deep neural network, the loss function fails to converge and oscillates wildly. Which adjustment is most likely to stabilize training?

A.Increase the number of hidden layers
B.Decrease the batch size
C.Reduce the learning rate
D.Use a test set
AnswerC

An excessively large learning rate causes the optimiser to overshoot minima, producing the wild oscillation described. Reducing it shrinks each weight update, letting the loss descend smoothly toward convergence instead of bouncing across the loss surface.

Why this answer

When the loss function oscillates wildly and fails to converge, it typically indicates that the learning rate is too high, causing the optimizer to overshoot the minima. Reducing the learning rate allows the gradient descent updates to take smaller, more stable steps, which helps the loss converge smoothly. This is a fundamental hyperparameter tuning step in deep learning training.

Exam trap

CompTIA often tests the misconception that increasing model complexity (more layers) or using more data (test set) directly fixes training instability, when in fact the learning rate is the primary culprit for oscillation and non-convergence.

How to eliminate wrong answers

Option A is wrong because increasing the number of hidden layers adds more parameters and non-linearity, which can exacerbate instability and overfitting, not stabilize training. Option B is wrong because decreasing the batch size increases the variance in gradient estimates, which often leads to noisier updates and can worsen oscillation, not reduce it. Option D is wrong because using a test set is for evaluating generalization performance after training, not for stabilizing the training process itself.

22
MCQhard

An organization has a dataset with categorical features having high cardinality (e.g., ZIP codes). They plan to use a tree-based model. Which encoding method is most appropriate?

A.Label encoding
B.One-hot encoding
C.Target encoding (mean encoding)
D.Frequency encoding
AnswerC

Target encoding maps categories to the mean target, preserving predictive information compactly.

Why this answer

Target encoding (mean encoding) replaces each category with the mean of the target variable for that category, which works well with tree-based models on high-cardinality features because it captures the predictive signal without exploding the feature space. This method avoids the dimensionality explosion of one-hot encoding and the arbitrary ordering of label encoding, making it the most appropriate choice for high-cardinality categorical features in tree-based models.

Exam trap

CompTIA often tests the misconception that one-hot encoding is always the safest choice for categorical data, but the trap here is that high cardinality makes one-hot encoding impractical, and candidates overlook target encoding as a cardinality-efficient alternative.

How to eliminate wrong answers

Option A is wrong because label encoding assigns arbitrary integer values to categories, which introduces an artificial ordinal relationship that tree-based models can misinterpret, leading to suboptimal splits. Option B is wrong because one-hot encoding creates a binary column for each unique category, which with high cardinality (e.g., thousands of ZIP codes) results in an extremely high-dimensional sparse feature matrix that degrades model performance and increases computational cost. Option D is wrong because frequency encoding replaces categories with their occurrence counts, which loses the relationship with the target variable and can be misleading when frequency does not correlate with the target, unlike target encoding which directly uses the target mean.

23
MCQeasy

A data scientist is training a binary classification model to detect fraudulent transactions. The dataset is highly imbalanced with 99% legitimate and 1% fraudulent. Which evaluation metric should be prioritized to assess model performance?

A.Accuracy
B.F1-score
C.Mean Squared Error
D.Log Loss
AnswerB

With 99% legitimate transactions, accuracy is misleading because a model predicting all legitimate scores 99%. F1-score balances precision and recall on the minority fraudulent class, exposing poor detection that accuracy hides, making it the appropriate metric for this imbalanced binary classification task.

Why this answer

In a highly imbalanced dataset (99% legitimate, 1% fraudulent), accuracy is misleading because a model that predicts all transactions as legitimate would achieve 99% accuracy without detecting any fraud. The F1-score combines precision and recall into a single metric, making it the preferred choice for evaluating binary classification performance on imbalanced data, as it penalizes both false positives and false negatives equally.

Exam trap

The trap here is that candidates often default to accuracy as the primary metric, not realizing that in highly imbalanced scenarios, accuracy can be artificially high and meaningless, while the F1-score reveals the true performance on the minority class.

How to eliminate wrong answers

Option A is wrong because accuracy is not suitable for imbalanced datasets; a naive model predicting the majority class can achieve high accuracy while failing to detect any fraudulent transactions. Option C is wrong because Mean Squared Error (MSE) is a regression metric used for continuous outputs, not for binary classification tasks. Option D is wrong because Log Loss measures the probabilistic confidence of predictions and, while useful, does not directly account for class imbalance in the same way the F1-score does; it can be dominated by the majority class's probabilities.

24
MCQeasy

A team is implementing a machine learning pipeline to classify images for a defect detection system. They are considering using a pre-trained convolutional neural network (CNN) and fine-tuning it on their small dataset. What is the primary advantage of transfer learning in this scenario?

A.It ensures the model is not biased toward the original dataset
B.It eliminates the need for data preprocessing
C.It allows the model to leverage learned features from a large dataset, reducing training time and required data
D.It reduces the risk of overfitting by using a larger model
AnswerC

Fine-tuning reuses convolutional filters already trained on millions of images, so the small defect dataset only needs to adjust higher layers. This cuts training time and data volume while retaining robust feature extraction, directly addressing the small-dataset constraint in the stem.

Why this answer

Transfer learning allows the team to start with a pre-trained CNN (e.g., trained on ImageNet) that has already learned general features like edges, textures, and shapes from a massive dataset. By fine-tuning only the later layers on their small defect dataset, they dramatically reduce training time and the amount of labeled data needed, while still achieving high accuracy.

Exam trap

The trap here is that candidates may think transfer learning eliminates all bias or preprocessing needs (options A and B), or mistakenly believe a larger model inherently reduces overfitting (option D), when in fact the core benefit is leveraging pre-learned features to reduce data and training time.

How to eliminate wrong answers

Option A is wrong because transfer learning does not eliminate bias from the original dataset; in fact, it intentionally leverages that bias (learned features) as a starting point, and fine-tuning may still carry some original dataset bias. Option B is wrong because transfer learning does not eliminate the need for data preprocessing; images must still be resized, normalized, and augmented to match the pre-trained model's input requirements. Option D is wrong because using a larger model (e.g., deeper CNN) actually increases the risk of overfitting on a small dataset, not reduces it; transfer learning mitigates overfitting by providing a strong feature initialization, not by using a larger model.

25
MCQmedium

An organization wants to automate the detection of defective products on an assembly line using computer vision. They have a limited number of labeled images for defective items. Which approach would be most effective?

A.Use a support vector machine with handcrafted features
B.Train a convolutional neural network from scratch on the limited data
C.Synthesize additional defective images using GANs
D.Use transfer learning with a pre-trained model like ResNet and fine-tune on the defect data
AnswerD

Transfer learning reuses features learned from large datasets like ImageNet, so a pre-trained ResNet needs only a small labelled defect set for fine-tuning. This directly addresses the limited labelled images constraint, unlike training from scratch which would overfit.

Why this answer

Transfer learning with a pre-trained model like ResNet is most effective because it leverages features learned from large datasets (e.g., ImageNet) and adapts them to the defect detection task with limited labeled data. Fine-tuning only the later layers preserves general visual features while specializing for defect classification, avoiding overfitting that would occur with a small dataset.

Exam trap

CompTIA often tests the misconception that more data (via GANs) is always better, or that starting from scratch is necessary for a new task, when in reality transfer learning is the standard solution for small datasets in computer vision.

How to eliminate wrong answers

Option A is wrong because handcrafted features with SVM require domain expertise to design and often fail to capture the complex, high-dimensional patterns in defect images, leading to poor generalization. Option B is wrong because training a CNN from scratch on limited data causes severe overfitting, as deep networks have millions of parameters that cannot be reliably learned from a small sample. Option C is wrong because while GANs can synthesize images, generating realistic and diverse defective samples that accurately represent real defects is extremely difficult and may introduce artifacts, making the model unreliable for production.

26
MCQeasy

A hospital wants to deploy a machine learning model to predict patient readmission risk within 30 days. They have a dataset with 10,000 records, 70 features including demographics, lab results, and past admissions. The target variable is binary (readmitted or not). The data scientist trains a logistic regression model and achieves an AUC of 0.85 on the test set. However, the hospital's clinicians require interpretability of predictions to trust the model. Which action should the data scientist take to ensure the model meets the interpretability requirement while maintaining performance?

A.Reduce the number of features to 10 using PCA and retrain the logistic regression
B.Replace logistic regression with a random forest model and use feature importance plots
C.Train a deep neural network and apply LIME or SHAP for explanations
D.Use the logistic regression model as is, since it is inherently interpretable with coefficients
AnswerD

Logistic regression produces coefficients that quantify each feature's contribution to the predicted readmission probability, satisfying the clinicians' interpretability constraint directly. Its AUC of 0.85 already meets performance expectations, so no trade-off or surrogate explainability tooling is needed.

Why this answer

Logistic regression is a linear model whose predictions are computed as a weighted sum of input features passed through a sigmoid function. The coefficients directly represent the log-odds change per unit increase in each feature, making the model inherently interpretable without any post-hoc explanation tools. Since the model already achieves an AUC of 0.85, which meets performance requirements, no architectural change is needed.

The data scientist should retain the logistic regression model and present the coefficients (and odds ratios) to clinicians to satisfy the interpretability requirement.

Exam trap

AI0-001 often tests the misconception that more complex models with post-hoc explanation tools (like SHAP or LIME) are necessary for interpretability, when in fact inherently interpretable models like logistic regression should be preferred when they meet performance requirements.

How to eliminate wrong answers

Option A is wrong because PCA transforms features into uncorrelated principal components that are linear combinations of the original variables, destroying the direct interpretability of coefficients in terms of original clinical features (e.g., 'age' or 'lab result X'). Option B is wrong because random forest, while offering feature importance plots, is a non-linear ensemble whose individual predictions are not directly interpretable; feature importance only shows global importance, not per-patient reasoning, and replacing logistic regression may not maintain the same AUC without extensive tuning. Option C is wrong because deep neural networks are highly non-linear and require post-hoc explanation methods like LIME or SHAP, which are approximations and may not provide the faithful, stable interpretability clinicians need; also, training a deep network on 10,000 records risks overfitting and may not improve performance.

27
MCQeasy

A data scientist needs to predict whether a customer will churn based on historical data containing features like account age, monthly charges, and support tickets. The target variable is binary (churn or not). Which type of machine learning algorithm should be used?

A.Linear regression
B.Logistic regression
C.K-means clustering
D.Principal component analysis
AnswerB

Logistic regression estimates the probability of a binary outcome via a sigmoid function, outputting a class label for churn or no churn. This satisfies the stem's constraint of a binary target variable, unlike linear regression, which predicts continuous values.

Why this answer

Logistic regression is the correct choice because it is specifically designed for binary classification tasks, such as predicting whether a customer will churn (yes/no). It models the probability of the binary outcome using a logistic (sigmoid) function, making it suitable for this supervised learning problem with a categorical target variable.

Exam trap

CompTIA AI often tests the distinction between regression and classification algorithms, trapping candidates who confuse linear regression (continuous output) with logistic regression (binary output) due to the misleading similarity in names.

How to eliminate wrong answers

Option A is wrong because linear regression predicts a continuous numeric output, not a binary class label, and would produce values outside the [0,1] range, making it unsuitable for classification. Option C is wrong because K-means clustering is an unsupervised learning algorithm used for grouping unlabeled data into clusters, not for predicting a known binary target. Option D is wrong because principal component analysis (PCA) is a dimensionality reduction technique used for feature extraction or noise reduction, not for making predictions on a target variable.

28
MCQhard

A deep learning model for natural language processing uses a recurrent neural network (RNN) to process long sequences. The gradients vanish after many time steps. Which architectural change is most effective to mitigate this problem?

A.Add dropout regularization
B.Use a larger learning rate
C.Replace the RNN cells with Long Short-Term Memory (LSTM) units
D.Increase the number of hidden layers
AnswerC

LSTM units add gating mechanisms and a cell state that preserve gradient flow across many time steps, directly counteracting the vanishing gradients an RNN suffers on long sequences. Replacing the RNN cells satisfies the stem's constraint, whereas simply adding layers or changing activation functions would not address the underlying decay.

Why this answer

LSTMs are specifically designed with a gating mechanism (input, forget, and output gates) and a cell state that allows gradients to flow unchanged over many time steps, directly addressing the vanishing gradient problem in standard RNNs. This architectural change preserves long-range dependencies in sequences, which is critical for tasks like language modeling or machine translation.

Exam trap

CompTIA AI often tests the misconception that regularization or hyperparameter tuning (like learning rate) can fix architectural gradient problems, but the correct answer always targets the root cause—here, the LSTM's gated structure that directly mitigates vanishing gradients.

How to eliminate wrong answers

Option A is wrong because dropout regularization randomly drops units during training to prevent overfitting, but it does not solve the vanishing gradient problem—it can even exacerbate gradient issues by reducing signal flow. Option B is wrong because increasing the learning rate can cause the gradients to explode or the loss to diverge, and it does not address the fundamental issue of gradients shrinking to zero over time. Option D is wrong because adding more hidden layers increases model depth, which typically worsens the vanishing gradient problem in standard RNNs due to repeated multiplication of small gradients through additional layers.

29
MCQmedium

A data scientist is training a deep neural network for sentiment analysis. The training loss decreases steadily but the validation loss starts to increase after 10 epochs. What is the most likely cause and best corrective action?

A.Underfitting; increase model complexity
B.Vanishing gradients; use ReLU activation
C.Data leakage; shuffle data before splitting
D.Overfitting; apply dropout and early stopping
AnswerD

Diverging training and validation loss after 10 epochs signals overfitting: the model memorises training data and generalises poorly. Dropout regularises the network, while early stopping halts training at the point validation loss begins rising, restoring generalisation.

Why this answer

The scenario describes a classic case of overfitting: the training loss decreases steadily, indicating the model is learning the training data well, but the validation loss increases after 10 epochs, meaning the model is memorizing noise and patterns specific to the training set rather than generalizing. The best corrective action is to apply dropout (which randomly drops neurons during training to reduce co-adaptation) and early stopping (which halts training when validation performance degrades), both of which are standard regularization techniques for deep neural networks.

Exam trap

CompTIA often tests the distinction between underfitting and overfitting by describing a diverging validation loss after initial improvement, leading candidates to mistakenly choose underfitting or vanishing gradients when the key indicator is the validation loss increase after a period of good training loss reduction.

How to eliminate wrong answers

Option A is wrong because underfitting would cause both training and validation loss to remain high or plateau, not a decreasing training loss with increasing validation loss; increasing model complexity would worsen overfitting, not fix it. Option B is wrong because vanishing gradients typically cause the training loss to stagnate or decrease very slowly, not a steady decrease followed by validation loss increase; ReLU activation helps mitigate vanishing gradients but does not address the overfitting pattern described. Option C is wrong because data leakage would cause both training and validation metrics to be artificially high from the start, not a divergence after 10 epochs; shuffling data before splitting is a best practice but does not correct overfitting that has already occurred.

30
MCQmedium

A team trains a decision tree on a customer churn dataset with 40 features. The unpruned tree reaches 100% accuracy on the training set but only 68% on a held-out validation set. The team wants to reduce this gap without changing the algorithm. Which action is most appropriate?

A.Increase the maximum tree depth
B.Remove the validation set and report training accuracy
C.Add more training examples without changing hyperparameters
D.Apply post-pruning with cost-complexity pruning
AnswerD

Cost-complexity pruning removes subtrees that add little predictive value, trading a small amount of training accuracy for better generalization. The parameter alpha controls the penalty for tree size and is tuned via cross-validation. Since the unpruned tree perfectly memorizes training data, pruning directly targets the overfitting gap and usually raises validation accuracy. It keeps the same algorithm, satisfying the team's constraint.

Why this answer

A tree with perfect training accuracy and much lower validation accuracy is overfitting, so the fix is to constrain its complexity. Cost-complexity pruning is the standard, algorithm-preserving remedy: it prunes branches that contribute little to reducing impurity and selects the penalty strength by cross-validation. The other actions either worsen overfitting, address it only indirectly, or destroy the ability to detect it.

Exam trap

The trap here is equating a perfect training score with a better model, when a 100% training accuracy alongside weak validation accuracy is a classic overfitting signal that calls for pruning rather than more capacity.

31
MCQeasy

Refer to the exhibit. The training log shows losses and accuracies over 5 epochs. What is the most likely problem?

A.Data leakage
B.Overfitting
C.Underfitting
D.Vanishing gradient
AnswerB

Overfitting fits because the log shows training loss still falling while validation loss rises after early epochs, meaning the model memorises training data rather than generalising. The widening gap between training and validation accuracy is the diagnostic constraint the stem's exhibit presents.

Why this answer

The training log shows high training accuracy (e.g., 99%) but low validation accuracy (e.g., 60%) across epochs, with the validation loss increasing after an initial drop. This divergence indicates the model has memorized the training data rather than learning generalizable patterns, which is the hallmark of overfitting.

Exam trap

CompTIA often tests the distinction between overfitting and underfitting by showing a training log where training accuracy is high but validation accuracy is low, leading candidates to mistakenly think the model is 'learning well' when it is actually memorizing.

How to eliminate wrong answers

Option A is wrong because data leakage would cause both training and validation accuracies to be artificially high and closely aligned, not a growing gap. Option C is wrong because underfitting would show low accuracy on both training and validation sets, not high training accuracy. Option D is wrong because vanishing gradient typically manifests as slow or stalled learning (flat loss curves) across all data splits, not a divergence between training and validation performance.

32
MCQmedium

Refer to the exhibit. What is the most likely issue and what action should be taken?

A.Learning rate is too low; increase it
B.Underfitting; increase model complexity
C.Overfitting; apply early stopping around epoch 15
D.Data imbalance; use class weights
AnswerC

The exhibit shows training loss continuing to fall while validation loss rises after roughly epoch 15, the classic divergence signature of overfitting. Early stopping at that point halts training before the model memorises noise, preserving generalisation. Regularisation or more data would also help, but stopping is the direct fix.

Why this answer

The training loss continues to decrease while the validation loss starts to increase after approximately epoch 15, which is a classic sign of overfitting. The model is memorizing the training data rather than generalizing, so applying early stopping around epoch 15 would prevent further divergence and preserve the best validation performance.

Exam trap

CompTIA often tests the distinction between overfitting and underfitting by showing loss curves where training loss continues to drop while validation loss rises, tricking candidates into thinking the model needs more training or a lower learning rate.

How to eliminate wrong answers

Option A is wrong because a low learning rate would cause both training and validation loss to decrease very slowly or plateau, not diverge with validation loss rising. Option B is wrong because underfitting would show both training and validation loss remaining high and not decreasing, whereas here training loss is still dropping. Option D is wrong because data imbalance typically causes poor performance on the minority class across both training and validation sets, not a divergence in loss curves after a certain epoch.

33
MCQmedium

A team is deploying a sentiment classifier and notices that the model outputs probabilities such as 0.83 for the positive class, but the actual positive rate among examples scored near 0.83 is only about 0.55. Stakeholders need the scores to reflect true likelihoods. Which action should the team take?

A.Replace the classifier with a k-nearest neighbors model, which produces inherently calibrated probabilities.
B.Raise the decision threshold from 0.5 to 0.83 so only highly confident examples are labeled positive.
C.Retrain the classifier with a lower learning rate and more epochs to improve probability estimates.
D.Apply a calibration method such as Platt scaling or isotonic regression to the model's output scores.
AnswerD

The stated problem is miscalibration: predicted probabilities do not match observed frequencies. Platt scaling fits a logistic transform on held-out scores, and isotonic regression fits a monotonic step function; both map raw scores to calibrated probabilities. This directly addresses the mismatch between 0.83 predicted and roughly 0.55 observed, without changing the model's ranking of examples.

Why this answer

The scenario describes a reliability problem: scores near 0.83 correspond to an actual positive rate near 0.55, so the probabilities are overconfident. Calibration techniques such as Platt scaling and isotonic regression learn a mapping from raw scores to empirical frequencies using a held-out set, producing probabilities that better reflect true likelihoods without changing the model's ranking ability.

Exam trap

The trap here is responding to a miscalibration complaint by changing the decision threshold, which alters classifications rather than the meaning of the probability scores.

34
Multi-Selecthard

Which TWO are key differences between Convolutional Neural Networks (CNN) and Recurrent Neural Networks (RNN)?

Select 2 answers
A.CNNs are designed for sequential data; RNNs for spatial data
B.RNNs have internal memory; CNNs do not
C.CNNs can handle variable-length inputs; RNNs require fixed-size inputs
D.CNNs use backpropagation; RNNs do not
E.CNNs use weight sharing across spatial dimensions; RNNs share weights across time steps
AnswersB, E

RNNs maintain a hidden state for temporal memory; CNNs are feedforward.

Why this answer

RNNs possess a hidden state that acts as internal memory, allowing them to retain information from previous time steps, which is essential for processing sequential data. In contrast, CNNs lack this internal memory mechanism; they process inputs independently without maintaining a state across different inputs, making them unsuitable for tasks requiring temporal context.

Exam trap

CompTIA often tests the misconception that CNNs and RNNs are distinguished by their training algorithms or input size requirements, when the core difference lies in their architectural design—specifically, internal memory and weight sharing mechanisms.

35
MCQhard

A company deploys a deep learning model for real-time object detection in autonomous vehicles. The model was trained on high-end GPUs but needs to run on edge devices with limited computational resources. Which technique is most effective for reducing model size and inference latency while maintaining acceptable accuracy?

A.Hyperparameter tuning
B.Batch normalization
C.Dropout
D.Quantization
AnswerD

Quantization reduces numerical precision of weights and activations, typically from 32-bit floats to 8-bit integers, shrinking memory footprint and enabling faster integer arithmetic on edge hardware. This directly satisfies the constraint of limited computational resources while preserving acceptable detection accuracy.

Why this answer

Quantization reduces the precision of the model's weights and activations (e.g., from 32-bit floating point to 8-bit integers), which significantly decreases model size and speeds up inference on edge devices with limited computational resources. This technique directly addresses the constraints of edge deployment while often maintaining acceptable accuracy through careful calibration.

Exam trap

CompTIA AI exams often test the misconception that regularization techniques like dropout or batch normalization can reduce model size or inference latency, when in fact they are training-phase optimizations that do not directly address edge deployment constraints.

How to eliminate wrong answers

Option A is wrong because hyperparameter tuning optimizes training settings (e.g., learning rate, batch size) to improve model convergence, but it does not directly reduce model size or inference latency on edge devices. Option B is wrong because batch normalization normalizes layer inputs during training to stabilize and accelerate training, but it adds computational overhead during inference and does not reduce model size or latency. Option C is wrong because dropout is a regularization technique that randomly drops neurons during training to prevent overfitting, but it is typically disabled during inference and does not reduce model size or inference latency.

36
Multi-Selectmedium

Which THREE are common activation functions used in neural networks? (Choose THREE.)

Select 3 answers
A.ReLU
B.Softmax
C.Sigmoid
D.Linear
E.Tanh
AnswersA, C, E

Rectified Linear Unit is widely used in hidden layers.

Why this answer

ReLU (Rectified Linear Unit) is a common activation function in neural networks because it introduces non-linearity while being computationally efficient. It outputs the input directly if positive, otherwise zero, which helps mitigate the vanishing gradient problem compared to sigmoid or tanh. This makes it a default choice for hidden layers in many deep learning architectures.

Exam trap

CompTIA often tests the distinction between activation functions used in hidden layers versus output layers, so candidates mistakenly select Softmax as a general activation function when it is only appropriate for the final layer in classification tasks.

37
Multi-Selecteasy

Which TWO are evaluation metrics for classification problems? (Choose two.)

Select 2 answers
A.Precision
B.Mean Absolute Error
C.R-squared
D.Mean Squared Error
E.Recall
AnswersA, E

Correct: Precision is a classification metric.

Why this answer

Precision is a classification metric that measures the proportion of true positive predictions among all positive predictions made by the model. It is calculated as TP / (TP + FP) and is critical when the cost of false positives is high, such as in spam detection or fraud alert systems.

Exam trap

CompTIA often tests the distinction between classification and regression metrics, and the trap here is that candidates may mistakenly select Mean Absolute Error or Mean Squared Error because they are common evaluation metrics, but they are exclusively used for regression problems, not classification.

38
MCQeasy

A machine learning engineer needs to choose an algorithm for grouping customers into segments based on purchasing behavior without any labels. Which algorithm should the engineer use?

A.K-means clustering
B.Random forest classifier
C.Linear regression
D.Support vector machine
AnswerA

K-means clustering partitions unlabelled data into k groups by minimising within-cluster variance, directly satisfying the stem's requirement to segment customers without labels. Supervised alternatives need target labels, which are absent here, so K-means fits the unsupervised grouping constraint.

Why this answer

K-means clustering is an unsupervised learning algorithm that groups unlabeled data into clusters based on feature similarity, making it ideal for segmenting customers by purchasing behavior without predefined labels. It partitions data into K clusters by minimizing within-cluster variance, which directly addresses the requirement of discovering natural groupings in the data.

Exam trap

The AI0-001 exam often tests the distinction between supervised and unsupervised learning, and the trap here is that candidates may confuse clustering with classification, picking a supervised algorithm like Random Forest or SVM because they think of 'grouping' as a classification task.

How to eliminate wrong answers

Option B (Random forest classifier) is wrong because it is a supervised ensemble method that requires labeled training data to classify instances, not suitable for unlabeled customer segmentation. Option C (Linear regression) is wrong because it is a supervised regression algorithm used to predict continuous values from labeled data, not for grouping unlabeled data. Option D (Support vector machine) is wrong because it is a supervised classification algorithm that requires labeled data to find a separating hyperplane, and cannot perform unsupervised clustering without modifications.

39
MCQmedium

A machine learning engineer is evaluating a classifier on a dataset with 1,000 examples where only 30 are positive. The model predicts the negative class for almost every example. The team reports 97% accuracy and claims success. Which metric should the engineer introduce to reveal the model's poor performance on the positive class?

A.Recall for the positive class
B.Silhouette score
C.Mean squared error
D.R-squared
AnswerA

Recall, or sensitivity, is the fraction of actual positives the model correctly identifies. With only 30 positives and a model that predicts negative almost always, recall will be near zero even though accuracy is 97%. Reporting positive-class recall immediately exposes that the classifier misses nearly all fraud, disease, or defect cases. It is the metric that directly measures performance on the minority class the team cares about.

Why this answer

Accuracy is misleading under severe class imbalance because a trivial majority-class predictor scores high. Recall on the positive class directly measures how many of the 30 true positives were captured, exposing the near-zero detection rate. Recall is the appropriate complement to accuracy when the cost of false negatives is high, such as fraud, medical screening, or safety defects, and it guides the team toward resampling, class weighting, or threshold adjustment.

Exam trap

The trap here is trusting overall accuracy on an imbalanced dataset, when a model that always predicts the majority class can score very high while being useless for the minority class.

40
MCQhard

A financial institution is developing a fraud detection model using historical transaction data. The dataset contains over 10 million records, but only 0.01% of transactions are fraudulent. The current model uses a neural network trained with standard cross-entropy loss, and the team applies random undersampling of the majority class to create a balanced training set. However, the model still produces a high number of false positives (legitimate transactions flagged as fraud) and misses approximately 30% of actual fraud cases. The business requires that at least 95% of frauds be caught, and the false positive rate must be below 1% to avoid overwhelming fraud analysts. The team has limited resources to collect additional data and cannot change the model architecture significantly. Which approach should the team take to best meet the business requirements?

A.Use cost-sensitive learning by assigning a higher misclassification cost to the fraud class.
B.Apply feature selection to remove noisy predictors and then retrain the current model.
C.Switch to an anomaly detection algorithm such as Isolation Forest or One-Class SVM.
D.Collect more transaction data, especially fraudulent examples, to naturally balance the classes.
AnswerA

This directly penalizes false negatives more, encouraging the model to catch more frauds while maintaining a low false positive rate through tuning.

Why this answer

Cost-sensitive learning directly addresses the business requirement by penalizing fraud misclassifications more heavily, which shifts the decision threshold to favor recall on the fraud class while still allowing the team to tune the trade-off between false positives and false negatives. Because the team cannot collect more data or change the architecture significantly, adjusting the loss function's class weights is the most practical lever to hit the 95% recall and <1% FPR targets. It also avoids the information loss caused by random undersampling, which discards 99.99% of legitimate transactions and distorts the true class distribution.

Exam trap

The trap is assuming that balancing the dataset (undersampling) or switching to anomaly detection solves imbalance, when the real issue is the asymmetric misclassification cost—candidates overlook cost-sensitive learning as the direct lever for meeting recall/FPR targets.

How to eliminate wrong answers

Option B is wrong because feature selection may remove signal and does not address the fundamental class imbalance or the asymmetric cost of errors; it is unlikely to move recall from 70% to 95% on its own. Option C is wrong because switching to Isolation Forest or One-Class SVM reframes the problem as unsupervised anomaly detection, which typically sacrifices precision and would struggle to meet the <1% FPR requirement while catching 95% of frauds. Option D is wrong because the team explicitly has limited resources to collect additional data, and even with more fraud examples, the extreme imbalance and cost asymmetry would still require cost-sensitive techniques.

41
MCQeasy

A team is deploying a deep learning model for real-time image classification on edge devices with limited computational resources. Which technique would best help reduce model size and inference time without significant accuracy loss?

A.Data augmentation
B.Model pruning and quantization
C.Transfer learning
D.Ensemble learning
AnswerB

Pruning removes redundant weights and neurons, while quantization reduces numeric precision (for example FP32 to INT8). Combined, they shrink model size and cut inference latency substantially on resource-constrained edge hardware, with accuracy loss typically recoverable through fine-tuning.

Why this answer

Model pruning and quantization directly reduce the number of parameters and the precision of weights (e.g., from 32-bit floats to 8-bit integers), which shrinks the model size and speeds up inference on edge devices. This technique is specifically designed to minimize computational load while preserving accuracy, making it ideal for resource-constrained environments like real-time image classification on edge hardware.

Exam trap

Candidates often mistakenly believe that transfer learning alone reduces model size, but it only reuses weights—the architecture remains unchanged. For resource-constrained edge devices, pruning and quantization are the direct methods for compression and speed optimization.

How to eliminate wrong answers

Option A is wrong because data augmentation increases the diversity of training data to improve generalization, but it does not reduce model size or inference time; in fact, it may increase training time. Option C is wrong because transfer learning reuses a pre-trained model to accelerate training on a new task, but it does not inherently reduce model size or inference speed—the model remains large unless combined with pruning or quantization. Option D is wrong because ensemble learning combines multiple models to improve accuracy, but it multiplies the computational cost and memory footprint, which is counterproductive for edge devices with limited resources.

42
MCQeasy

A retail analytics team has a labeled dataset of 50,000 customer transactions where each record is tagged as either 'fraudulent' or 'legitimate.' They need a supervised learning approach that outputs a probability between 0 and 1 for the fraudulent class so it can be compared against a business threshold. Which algorithm is most appropriate for this task?

A.K-means clustering
B.Apriori association rule mining
C.Principal component analysis
D.Logistic regression
AnswerD

Logistic regression is a supervised binary classifier that applies a sigmoid function to a linear combination of features, producing a probability between 0 and 1 for the positive class. This exactly matches the requirement to compare the fraud probability against a threshold. It also trains efficiently on 50,000 labeled records and yields interpretable coefficients, which is valuable for explaining fraud decisions to stakeholders.

Why this answer

The task is supervised binary classification with a need for a 0-to-1 probability, which points to logistic regression because its sigmoid output is directly comparable to a decision threshold. The other techniques either lack supervision (K-means, Apriori) or do not model the target label (PCA). Logistic regression also scales well to 50,000 records and provides coefficients that help explain which transaction features drive fraud risk.

Exam trap

The trap here is assuming any algorithm that groups or summarizes data can classify labeled records, when supervised classification requires a model that learns from the target label and emits a class probability.

43
MCQmedium

A team is training a convolutional neural network (CNN) for medical image diagnosis. They have a limited dataset of 500 labeled images. Which strategy is most effective to improve model generalization?

A.Increasing network depth
B.Data augmentation
C.Using a larger batch size
D.Reducing the number of filters
AnswerB

Data augmentation synthetically expands the 500 labelled images via rotations, flips and crops, directly addressing the limited-dataset constraint. This reduces overfitting and improves generalisation without requiring new labelled data, unlike transfer learning which needs a pretrained model or regularisation which only penalises complexity.

Why this answer

With only 500 labeled medical images, the primary challenge is overfitting due to limited data. Data augmentation (e.g., random rotations, flips, zooms) artificially expands the training set by creating varied but realistic transformations, which forces the CNN to learn invariant features and significantly improves generalization to unseen data.

Exam trap

The AI0-001 exam often tests the misconception that increasing model complexity (depth or filters) always improves performance, but with limited data, the correct strategy is to use regularization techniques like data augmentation to combat overfitting.

How to eliminate wrong answers

Option A is wrong because increasing network depth adds more parameters, which exacerbates overfitting on a small dataset and requires more data to train effectively. Option C is wrong because using a larger batch size provides a noisier gradient estimate and can lead to sharper minima, often reducing generalization, especially with limited data. Option D is wrong because reducing the number of filters lowers the model's capacity, which may cause underfitting and fail to capture the complex patterns needed for medical image diagnosis.

44
MCQeasy

A junior data scientist is training a supervised classification model to predict whether a loan applicant will default. The dataset has 40,000 labeled historical records with a clear binary outcome column. The team needs a model that outputs a probability between 0 and 1 for the default class. Which algorithm is the most appropriate choice for this task?

A.Logistic regression
B.Principal component analysis
C.Apriori association rule mining
D.K-means clustering
AnswerA

Logistic regression applies a sigmoid function to a linear combination of features, directly producing a probability between 0 and 1 for a binary outcome. With 40,000 labeled records and a binary target, it fits the supervised classification scenario, trains quickly, and yields interpretable coefficients that can support lending decisions and regulatory review.

Why this answer

Logistic regression is designed for binary classification and directly estimates the probability that an observation belongs to the positive class through the sigmoid link function. With abundant labeled data and a clear binary target, it satisfies the supervised learning requirement and delivers interpretable, well-calibrated outputs suitable for credit risk decisions.

Exam trap

The trap here is assuming any algorithm that groups or transforms data can serve as a classifier, when unsupervised methods such as clustering and dimensionality reduction never use the target label.

45
MCQmedium

A machine learning engineer is building a spam filter. The dataset contains 10,000 emails, of which 1,000 are spam. The engineer decides to use a Random Forest classifier. Which preprocessing step is most critical to ensure the model generalizes well to new, unseen emails?

A.Apply Principal Component Analysis (PCA) to reduce dimensionality
B.Normalize the numerical features to have zero mean and unit variance
C.Split the data into training and testing sets before any other preprocessing
D.Encode all features using one-hot encoding
AnswerC

Holding out a test set before any preprocessing prevents data leakage, because fitting transformations such as scaling or resampling on the full dataset lets test information influence training. The split must therefore precede all other preprocessing to give an honest estimate of generalisation.

Why this answer

Splitting the data into training and testing sets before any other preprocessing prevents data leakage. If preprocessing like normalization or PCA is applied to the entire dataset first, the test set information influences the training process, leading to overly optimistic performance estimates and poor generalization to new, unseen emails.

Exam trap

CompTIA often tests the concept of data leakage by presenting preprocessing steps that seem harmless but actually incorporate test set information, tricking candidates into thinking scaling or dimensionality reduction is always necessary for tree-based models.

How to eliminate wrong answers

Option A is wrong because PCA is an unsupervised dimensionality reduction technique that, if applied before splitting, would leak information from the test set into the training set, and Random Forest is robust to high-dimensional sparse data, making PCA unnecessary for generalization. Option B is wrong because Random Forest is a tree-based ensemble method that is invariant to monotonic transformations and does not require feature scaling; normalizing before splitting would also risk data leakage if done on the full dataset. Option D is wrong because one-hot encoding is only relevant for categorical features, and applying it before splitting could introduce data leakage if the encoding uses levels present only in the test set; moreover, not all features in an email dataset are categorical, and Random Forest can handle label encoding without one-hot encoding.

46
Multi-Selecthard

Which TWO are valid techniques to reduce overfitting in a deep neural network? (Choose TWO.)

Select 2 answers
A.Increase batch size
B.Increase learning rate
C.L2 regularization
D.Gradient clipping
E.Dropout
AnswersC, E

L2 regularization adds a penalty proportional to the squared magnitude of weights to the loss function, shrinking weights toward zero. This constrains model complexity, reducing variance so the network generalises better rather than memorising training samples, which directly counteracts overfitting.

Why this answer

L2 regularization (option C) is a valid technique to reduce overfitting by adding a penalty term proportional to the square of the weight magnitudes to the loss function. This discourages the network from learning overly complex patterns, effectively shrinking weights and improving generalization. Dropout (option E) randomly drops a fraction of neurons during training, which prevents co-adaptation of features and forces the network to learn more robust representations, also reducing overfitting.

Exam trap

CompTIA often tests the distinction between techniques that improve training stability (like gradient clipping or adjusting batch size/learning rate) versus those that directly regularize the model to reduce overfitting (like L2 regularization and dropout), leading candidates to confuse optimization tricks with regularization methods.

47
MCQeasy

A team is building a recommendation system using collaborative filtering. They have a sparse user-item matrix. Which technique should they use to handle the sparsity and improve recommendations?

A.Association rule mining
B.Matrix factorization
C.k-nearest neighbors
D.Content-based filtering
AnswerB

Matrix factorization decomposes the sparse user-item matrix into lower-dimensional latent factor matrices, capturing hidden relationships between users and items. This reduces dimensionality and fills implicit gaps, generating meaningful recommendations despite missing ratings that plague collaborative filtering on sparse data.

Why this answer

Matrix factorization (B) is the correct technique because it decomposes the sparse user-item matrix into lower-dimensional latent factor matrices, effectively capturing underlying patterns and filling in missing entries. This directly addresses sparsity by learning dense representations that generalize beyond observed interactions, which is a core strength in collaborative filtering for recommendation systems.

Exam trap

CompTIA often tests the misconception that k-nearest neighbors (k-NN) is the go-to for collaborative filtering, but candidates fail to recognize that k-NN's performance collapses under high sparsity, whereas matrix factorization explicitly models latent factors to overcome this.

How to eliminate wrong answers

Option A is wrong because association rule mining (e.g., Apriori algorithm) is designed for market basket analysis to find frequent itemsets and rules, not for handling sparse user-item matrices in collaborative filtering; it fails to generalize from sparse data and does not model latent factors. Option C is wrong because k-nearest neighbors (k-NN) is a memory-based collaborative filtering method that relies on direct similarity computations between users or items, which degrades severely with high sparsity due to lack of overlapping ratings, leading to poor recommendations. Option D is wrong because content-based filtering uses item features (e.g., genre, keywords) to recommend similar items, not the user-item interaction matrix; it does not address sparsity in collaborative filtering and ignores collaborative signals from other users.

48
Multi-Selecteasy

Which TWO of the following are common activation functions used in deep neural networks?

Select 2 answers
A.Linear Regression
B.Support Vector Machine
C.K-means
D.ReLU
E.Sigmoid
AnswersD, E

ReLU is the most common activation for hidden layers.

Why this answer

ReLU (Rectified Linear Unit) is a common activation function in deep neural networks because it introduces non-linearity while being computationally efficient, outputting the input directly if positive and zero otherwise. It helps mitigate the vanishing gradient problem, making it a default choice for hidden layers in many architectures.

Exam trap

CompTIA AI often tests the distinction between machine learning algorithms (like Linear Regression, SVM, K-means) and neural network components (like activation functions), so candidates mistakenly select algorithms as activation functions because they recognize them as common ML terms.

49
MCQeasy

A machine learning engineer has a dataset of 100,000 records. She splits it into 70% training, 15% validation, and 15% test sets. After training, the model achieves 95% accuracy on training and 85% on validation. What does the accuracy difference most likely indicate?

A.The validation set is too small
B.The model generalizes well
C.The model is overfitting
D.The test set should be larger
AnswerC

The model fits training data closely but generalises worse to unseen validation data, a 10-point gap indicating overfitting. It has learned noise and idiosyncrasies rather than the underlying pattern, so validation accuracy lags training accuracy.

Why this answer

The 10% gap between training accuracy (95%) and validation accuracy (85%) is a classic sign of overfitting. The model has memorized patterns specific to the training set rather than learning generalizable features, causing it to perform worse on unseen validation data. In machine learning, a significant drop in performance from training to validation indicates poor generalization, which is the hallmark of overfitting.

Exam trap

CompTIA often tests the distinction between overfitting and data split issues, trapping candidates who mistake a performance gap for an insufficient validation set rather than recognizing it as a model generalization problem.

How to eliminate wrong answers

Option A is wrong because a 15% validation set (15,000 records) is generally sufficient for reliable evaluation; the issue is not size but the performance gap. Option B is wrong because good generalization would show similar accuracy on training and validation sets, not a 10% drop. Option D is wrong because the test set size (15%) is standard and does not affect the training-to-validation accuracy discrepancy; the problem lies in model behavior, not data partitioning.

50
MCQeasy

A data analyst wants to predict housing prices based on square footage, number of bedrooms, and location. Which machine learning approach is most suitable?

A.K-means clustering
B.Decision tree regression
C.Association rule mining
D.Linear regression
AnswerD

Linear regression models a continuous numeric target, here housing price, as a weighted linear combination of input features such as square footage, bedroom count, and location. This supervised regression approach directly fits predicting a continuous price value.

Why this answer

Linear regression is the most suitable approach because the problem involves predicting a continuous numeric target (housing prices) from multiple independent variables (square footage, bedrooms, location). Linear regression models the linear relationship between the features and the target, providing interpretable coefficients and efficient training for this type of regression task.

Exam trap

The trap here is that candidates may confuse regression (predicting a continuous value) with classification or unsupervised learning, and incorrectly select decision tree regression or clustering because they see 'prediction' and assume any tree-based or grouping method works.

How to eliminate wrong answers

Option A is wrong because K-means clustering is an unsupervised learning algorithm used for grouping unlabeled data into clusters, not for predicting a continuous target variable. Option B is wrong because decision tree regression can be used for regression, but it is not the most suitable here; it tends to overfit and lacks the interpretability and simplicity of linear regression for a straightforward linear relationship. Option C is wrong because association rule mining is an unsupervised technique for discovering frequent itemsets and rules in transactional data, not for predicting numeric values.

51
MCQhard

A deep learning model for image classification is overfitting the training data. The team has already tried data augmentation and dropout. Which additional technique should they implement to reduce overfitting?

A.Batch normalization
B.Increase number of epochs
C.Gradient clipping
D.Early stopping
AnswerD

Early stopping halts training once validation loss stops improving, preventing the model from continuing to memorise training noise. Unlike augmentation and dropout, which alter inputs or activations, it directly constrains the number of optimisation steps, complementing the techniques already applied.

Why this answer

Early stopping (Option D) is the correct additional technique because it halts training when validation performance stops improving, directly preventing the model from memorizing noise in the training data. Since data augmentation and dropout are already in use, early stopping provides a complementary regularization effect by limiting the number of training iterations before overfitting occurs.

Exam trap

CompTIA often tests the distinction between techniques that address overfitting versus those that solve optimization issues, leading candidates to confuse batch normalization or gradient clipping as overfitting solutions when they are not.

How to eliminate wrong answers

Option A is wrong because batch normalization primarily accelerates training and stabilizes learning by normalizing layer inputs, but it does not directly reduce overfitting—it can even have a slight regularizing effect, but it is not a primary overfitting countermeasure. Option B is wrong because increasing the number of epochs would exacerbate overfitting by giving the model more opportunities to memorize training data, making the problem worse. Option C is wrong because gradient clipping is used to prevent exploding gradients in deep networks, especially in RNNs, and does not address overfitting from excessive model capacity or insufficient regularization.

52
MCQmedium

A team trains a model to predict whether loan applicants will default. On the holdout set the model achieves 0.86 AUC, but when audited, applicants over 60 receive systematically higher risk scores than equally qualified younger applicants. The team must reduce this disparity while preserving predictive performance. Which action should they take first?

A.Remove the age column from the training data and retrain the model.
B.Increase model complexity so it can learn finer distinctions between applicants.
C.Lower the overall decision threshold to approve more applicants.
D.Measure disparity across protected groups and inspect whether age-correlated features drive the scores.
AnswerD

Before mitigating bias, the team must quantify it and locate its source. Computing group-level metrics such as demographic parity or equal opportunity, and examining feature importance for age-correlated variables, reveals whether the disparity comes from a proxy feature, label bias, or sampling. This diagnostic step is the prerequisite for any effective, targeted fairness intervention.

Why this answer

Bias mitigation should begin with measurement, because the source of disparity determines the right remedy. Group-level fairness metrics and feature attribution reveal whether age itself, a correlated proxy, or biased labels drive the scores. Only after diagnosing the cause can the team choose an appropriate intervention, such as reweighting, proxy removal, or post-processing.

Exam trap

The trap here is assuming that deleting the protected attribute eliminates bias, when correlated proxy features often preserve the disparity and can even mask it from casual inspection.

53
MCQeasy

A company wants to deploy a machine learning model that requires continuous learning as new data arrives. The model must be able to adapt to changing patterns without retraining from scratch. Which approach should be used?

A.Transfer learning
B.Online learning
C.Batch learning
D.Unsupervised learning
AnswerB

Online learning updates model parameters incrementally as each new data point arrives, rather than retraining on the full dataset. This satisfies the stem's constraint of continuous adaptation to changing patterns without retraining from scratch, unlike batch learning.

Why this answer

Online learning (also called incremental learning) updates the model incrementally as each new data point arrives, without requiring full retraining. This makes it ideal for scenarios where data arrives continuously and patterns shift over time, as the model can adapt its parameters on the fly.

Exam trap

CompTIA often tests the distinction between training paradigms (online vs. batch) and other ML concepts like transfer learning or unsupervised learning, so candidates may confuse 'continuous learning' with 'transfer learning' or incorrectly assume that any learning method can handle streaming data.

How to eliminate wrong answers

Option A is wrong because transfer learning reuses a pre-trained model on a new but related task, but it does not inherently support continuous adaptation to streaming data—it typically requires a separate fine-tuning phase. Option C is wrong because batch learning trains the model on the entire dataset at once and requires retraining from scratch when new data arrives, making it unsuitable for continuous learning. Option D is wrong because unsupervised learning is a paradigm for finding patterns in unlabeled data, not a deployment strategy for handling streaming data or model updates.

54
MCQmedium

A healthcare organization wants to use patient data to predict disease risk. They are concerned about bias in the model. Which step is most critical during the data preparation phase to mitigate bias?

A.Applying SMOTE to oversample minority classes
B.Using a more complex algorithm
C.Removing all demographic features
D.Ensuring the training data is representative of the target population
AnswerD

Bias originates in data whose composition misrepresents the population the model will serve. Ensuring the training set reflects the target population's demographics and disease prevalence prevents systematic underrepresentation, directly satisfying the data-preparation constraint before any modelling begins.

Why this answer

Ensuring the training data is representative of the target population is the most critical step during data preparation to mitigate bias because bias often originates from skewed or incomplete data that does not reflect the real-world distribution of patient demographics, conditions, and outcomes. Without a representative dataset, any subsequent preprocessing or modeling will propagate and potentially amplify existing disparities, leading to unfair or inaccurate predictions for underrepresented groups.

Exam trap

CompTIA often tests the misconception that bias can be fixed by technical tweaks like oversampling or removing sensitive attributes, when in fact the root cause is almost always unrepresentative training data that must be addressed at the collection or sampling stage.

How to eliminate wrong answers

Option A is wrong because SMOTE (Synthetic Minority Oversampling Technique) addresses class imbalance by generating synthetic samples for minority classes, but it does not correct for broader representational bias (e.g., missing demographic subgroups) and can introduce artifacts if the minority class itself is not representative of the target population. Option B is wrong because using a more complex algorithm does not mitigate bias; in fact, complex models can overfit to spurious correlations in biased data, making bias worse rather than reducing it. Option C is wrong because removing all demographic features can mask bias but does not eliminate it—protected attributes like race or age may be correlated with other features (e.g., zip code, medical history), leading to proxy discrimination, and this approach can also remove clinically relevant information needed for accurate risk prediction.

55
Multi-Selecthard

A data science team is preparing a gradient boosting model to predict equipment failure from sensor data. They want to tune hyperparameters that primarily control model complexity and reduce overfitting. Which two hyperparameters should they focus on? (Choose two.)

Select 2 answers
A.Random seed for data shuffling
B.Output directory for saved model artifacts
C.Number of CPU threads used during training
D.Learning rate
E.Maximum tree depth
AnswersD, E

The learning rate scales each tree's contribution to the ensemble. A smaller value forces the model to build more trees and take smaller corrective steps, which typically reduces overfitting and improves generalization. Paired with an appropriate number of estimators, tuning the learning rate is a core complexity control in gradient boosting for noisy sensor data.

Why this answer

Learning rate and maximum tree depth are structural hyperparameters that govern how much each tree contributes and how complex each tree can become. Lowering them constrains the ensemble's capacity, which reduces variance and overfitting. The other choices affect runtime, reproducibility, or file storage rather than the bias-variance tradeoff.

Exam trap

The trap here is treating any configurable training setting as a hyperparameter for overfitting, when operational settings such as thread count, seed, and output paths do not change model complexity.

56
MCQmedium

A machine learning engineer is tuning a neural network for image classification. The training loss decreases steadily, but the validation loss starts increasing after 50 epochs. Which action best addresses this issue?

A.Increase the number of hidden layers
B.Add more training data
C.Apply early stopping with a patience of 10 epochs
D.Increase the batch size
AnswerC

Early stopping with patience halts training once validation loss stops improving for 10 consecutive epochs, directly countering the overfitting that begins after epoch 50. It preserves the best-performing weights rather than continuing to minimise training loss, satisfying the stem's requirement to address rising validation loss.

Why this answer

The described behavior—decreasing training loss with increasing validation loss—is a classic sign of overfitting. Early stopping with a patience of 10 epochs directly addresses this by halting training when the validation loss fails to improve for a specified number of epochs, preventing further overfitting while retaining the best model weights.

Exam trap

The AI0-001 exam often tests the distinction between underfitting and overfitting symptoms, and the trap here is that candidates may confuse a rising validation loss with a need for more data or a deeper network, when the correct action is to stop training early to combat overfitting.

How to eliminate wrong answers

Option A is wrong because increasing the number of hidden layers increases model capacity, which typically worsens overfitting by allowing the network to memorize training data more easily. Option B is wrong because adding more training data can help reduce overfitting in general, but it is not the most direct or immediate fix for the specific problem of validation loss increasing after 50 epochs; early stopping is a more targeted and efficient solution. Option D is wrong because increasing the batch size provides a more accurate gradient estimate but does not prevent overfitting; it may even lead to sharper minima and poorer generalization, making the validation loss issue worse.

57
Multi-Selecteasy

Which TWO are characteristics of supervised learning?

Select 2 answers
A.Does not require target variable
B.Requires labeled data
C.Uses reinforcement signals
D.Learns to cluster data
E.Predicts continuous or categorical output
AnswersB, E

Supervised learning trains on input-output pairs where each example carries a ground-truth target, allowing the model to minimise loss against known labels. This labelled-data requirement is the defining characteristic separating it from unsupervised and reinforcement approaches.

Why this answer

Option B is correct because supervised learning fundamentally requires a labeled dataset, where each training example is paired with a known target (ground-truth) value that the model learns to map inputs to. Option E is correct because supervised models are trained to predict an output that is either continuous (regression, e.g., predicting a price) or categorical (classification, e.g., predicting a class label). Option A is incorrect because the absence of a target variable describes unsupervised learning, not supervised learning.

Option C is incorrect because reinforcement signals (rewards/penalties from an environment) characterize reinforcement learning, a separate paradigm. Option D is incorrect because clustering is an unsupervised task that discovers structure without labeled targets.

Exam trap

The AI0-001 exam often tests the distinction between supervised and unsupervised learning by presenting 'clustering' or 'reinforcement signals' as plausible characteristics of supervised learning, trapping candidates who confuse task types.

58
MCQhard

The exhibit shows a model configuration for a classification task with 10 classes. What is wrong with this setup?

A.The loss function should be categorical crossentropy, not mean squared error
B.The metric should be precision, not accuracy
C.The activation should be sigmoid in hidden layers
D.The optimizer should be SGD, not Adam
AnswerA

Mean squared error suits regression, penalising numeric distance between continuous outputs. For 10-class classification, categorical crossentropy compares predicted probability distributions against one-hot labels, giving the correct gradient signal; MSE on softmax outputs produces weak, misleading updates.

Why this answer

In a multi-class classification task with 10 classes, the correct loss function is categorical crossentropy because it measures the dissimilarity between the true probability distribution and the predicted probability distribution. Mean squared error (MSE) is designed for regression tasks and penalizes errors in a way that is not suitable for classification probabilities, leading to poor gradient behavior and slower convergence.

Exam trap

The AI0-001 exam often tests the misconception that MSE can be used as a generic loss function for any task, but in classification, crossentropy is specifically designed to handle probability distributions and one-hot encoding.

How to eliminate wrong answers

Option B is wrong because accuracy is the standard metric for multi-class classification tasks; precision is typically used for binary classification or when focusing on specific class performance, but it is not a general replacement for accuracy. Option C is wrong because sigmoid activation in hidden layers can cause vanishing gradients and is not optimal; ReLU or its variants are preferred for hidden layers to mitigate gradient issues. Option D is wrong because Adam is a widely used optimizer that adapts learning rates and often outperforms SGD in practice; there is no inherent problem with using Adam for this setup.

59
MCQmedium

An AI engineer is training a deep neural network for image recognition. The training loss decreases steadily for the first few epochs but then plateaus and starts to oscillate. Which adjustment is most likely to improve convergence?

A.Add more layers
B.Increase the learning rate
C.Increase the batch size
D.Reduce the learning rate
AnswerD

Oscillating loss after an initial plateau indicates the optimiser is overshooting minima because each update step is too large. Lowering the learning rate shrinks those steps, allowing the network to settle into a smoother convergence path.

Why this answer

The plateau and oscillation of the training loss indicate that the optimizer is overshooting the minimum due to a learning rate that is too high. Reducing the learning rate allows the optimizer to take smaller, more precise steps, dampening oscillations and enabling convergence to a lower loss. This is a standard technique in gradient descent optimization, often implemented via learning rate schedules or adaptive methods like Adam.

Exam trap

CompTIA often tests the misconception that increasing the learning rate speeds up convergence, when in fact it causes divergence or oscillation, and that adding layers always improves performance, ignoring the risk of overfitting and optimization difficulty.

How to eliminate wrong answers

Option A is wrong because adding more layers increases model complexity, which typically exacerbates overfitting and can worsen convergence issues when the loss is already oscillating. Option B is wrong because increasing the learning rate would make the oscillations larger and more erratic, moving the optimizer further from the minimum. Option C is wrong because increasing the batch size reduces the variance of gradient estimates but does not address the fundamental issue of an overly large step size causing oscillations; it may even slow convergence by requiring more epochs to process the same data.

60
MCQhard

An AI developer observes that the training accuracy of a neural network is high, but the test accuracy is low. The model uses a ReLU activation function and Adam optimizer. Which approach is most likely to improve test accuracy?

A.Increase the learning rate
B.Add L2 regularization to the loss function
C.Switch to a stochastic gradient descent optimizer
D.Increase the number of epochs
AnswerB

L2 regularization penalises large weights, reducing overfitting so the model generalises better to unseen data, which raises test accuracy. This addresses the stem's high training accuracy versus low test accuracy gap, unlike ReLU or Adam changes.

Why this answer

The scenario describes overfitting: high training accuracy but low test accuracy. Adding L2 regularization (weight decay) penalizes large weights, which reduces model complexity and improves generalization to unseen data. This directly addresses the overfitting problem without altering the model architecture or optimization algorithm.

Exam trap

AI0-001 often tests the misconception that changing the optimizer or increasing training duration can fix overfitting, when in fact regularization techniques like L2 are the direct solution.

How to eliminate wrong answers

Option A is wrong because increasing the learning rate would likely cause unstable training and may worsen overfitting by allowing the model to fit noise more aggressively. Option C is wrong because switching to SGD does not inherently fix overfitting; while SGD can have a regularizing effect due to noise, it is not the most direct or reliable solution compared to explicit regularization. Option D is wrong because increasing the number of epochs would allow the model to continue fitting the training data even more closely, exacerbating overfitting and further reducing test accuracy.

61
Multi-Selectmedium

Which THREE techniques can help reduce overfitting in neural networks?

Select 3 answers
A.Increasing training data size
B.L2 regularization
C.Using a larger learning rate
D.Dropout
E.Increasing number of layers
AnswersA, B, D

More data helps the model generalize better.

Why this answer

Increasing the training data size helps reduce overfitting by providing the model with more examples to learn from, which reduces the variance and improves generalization. With more data, the model is less likely to memorize noise and instead learns the underlying patterns, making it more robust on unseen data.

Exam trap

CompTIA often tests the misconception that increasing model complexity (e.g., more layers or larger learning rates) can help with overfitting, when in fact these changes typically worsen it by increasing variance or destabilizing training.

62
MCQmedium

A healthcare analytics team trains a model to flag patients at risk of readmission. The dataset contains 9,500 non-readmitted patients and 500 readmitted patients. The model predicts the majority class for every patient and reports 95 percent accuracy, yet it identifies no at-risk patients. Which evaluation approach best reveals the model's failure?

A.Measure training time and inference latency across the full dataset
B.Compute the confusion matrix and examine recall and precision for the readmission class
C.Report overall accuracy on a stratified holdout set
D.Calculate the mean squared error between predicted probabilities and labels
AnswerB

The confusion matrix exposes the zero true positives directly, and recall for the readmission class is zero while precision is undefined. These class-specific metrics reveal that the model catches no at-risk patients, which accuracy hides. Recall matters most clinically because a missed readmission is a false negative with real patient harm.

Why this answer

With a 95-to-5 class split, a majority-class predictor achieves 95 percent accuracy while providing zero clinical value. The confusion matrix and minority-class recall and precision expose the absence of true positives, which is the evidence needed to justify resampling, class weighting, or threshold tuning.

Exam trap

The trap here is trusting a high accuracy number on an imbalanced dataset, when accuracy can be maximized by simply ignoring the minority class that the model exists to detect.

63
MCQmedium

A fraud-detection team at a bank trains a gradient-boosted tree model on two years of transaction data. Only 0.4% of transactions are fraudulent. The model achieves 99.7% accuracy but flags almost no fraud. Which approach best addresses the underlying problem with how the model is being trained and evaluated?

A.Remove outliers using a z-score filter and standardize all numeric features before retraining the model.
B.Increase the number of boosting rounds and lower the learning rate until training accuracy reaches 100%.
C.Optimize the model using precision-recall AUC and apply class weighting or resampling to the fraudulent class.
D.Switch the evaluation metric to root mean squared error and report it alongside accuracy for each boosting round.
AnswerC

Accuracy is misleading on a 0.4% positive class because a model that predicts 'not fraud' every time scores 99.6% accuracy while catching nothing. Precision-recall AUC focuses on the minority class, and class weighting or resampling forces the learner to pay attention to fraudulent examples. Together they fix both the training signal and the evaluation metric, which is exactly the failure observed.

Why this answer

Accuracy is a poor metric when one class is extremely rare, because a trivial majority-class predictor scores deceptively high. The real fix is twofold: train with class weighting or resampling so the learner sees the minority class, and evaluate with precision-recall AUC, which reflects performance on the fraud class rather than being swamped by true negatives.

Exam trap

The trap here is assuming that a 99.7% accuracy score means the model is performing well, when in fact it is simply predicting the majority class.

64
MCQmedium

During training of a neural network, the loss oscillates and does not converge smoothly. The learning rate is set to 0.1. What is the most likely cause and what adjustment should be made?

A.Learning rate too low; increase it
B.Batch size too small; increase it
C.Learning rate too high; decrease it
D.Too many epochs; stop early
AnswerC

A learning rate of 0.1 is large enough that each gradient step overshoots the loss minimum, producing oscillation instead of smooth convergence. Reducing the learning rate shrinks step size, allowing the optimiser to settle into the minimum rather than bouncing across it.

Why this answer

A learning rate of 0.1 is relatively high for many neural network architectures. When the learning rate is too high, the optimizer takes steps that overshoot the minimum of the loss function, causing the loss to oscillate or even diverge instead of converging smoothly. Decreasing the learning rate allows for smaller, more stable weight updates, leading to smoother convergence.

Exam trap

CompTIA often tests the misconception that a high learning rate always speeds up training; the trap here is that candidates may think increasing the learning rate will force faster convergence, when in fact it causes instability and oscillation.

How to eliminate wrong answers

Option A is wrong because increasing an already high learning rate (0.1) would exacerbate oscillations and likely cause divergence, not fix the convergence issue. Option B is wrong because a small batch size introduces noise into gradient estimates, which can cause oscillations, but the question states the loss oscillates and does not converge smoothly; while increasing batch size can reduce variance, the primary suspect with a learning rate of 0.1 is that the learning rate itself is too high. Option D is wrong because stopping early would truncate training before convergence, not address the underlying cause of oscillations; the model may still be far from optimal, and the loss pattern indicates a learning rate problem, not overfitting or excessive epochs.

65
MCQhard

A fraud detection model is trained on a dataset where only 0.1% of transactions are fraudulent. The model achieves 99.9% accuracy but fails to catch most frauds. Which metric should the team prioritize, and which technique could help?

A.Mean Squared Error; use L2 regularization
B.F1 score; use principal component analysis
C.Accuracy; collect more data
D.Precision-Recall AUC; use oversampling like SMOTE
AnswerD

With 0.1% positives, accuracy is misleading because predicting all negatives scores 99.9%. Precision-Recall AUC focuses on the minority fraud class, and SMOTE synthetically oversamples it, directly addressing the severe class imbalance that causes the model to miss most frauds.

Why this answer

The dataset is highly imbalanced (0.1% fraud), so 99.9% accuracy is misleading because a model that predicts 'not fraud' for every transaction achieves it. Precision-Recall AUC focuses on the positive class (fraud) and is robust to class imbalance, unlike accuracy or ROC-AUC. Oversampling like SMOTE generates synthetic fraud samples to balance the dataset, helping the model learn the minority class patterns.

Exam trap

Test-takers often mistakenly believe that high accuracy always indicates a good model, and that techniques like PCA or regularization can fix class imbalance. In reality, only metrics and resampling methods designed for skewed distributions are effective.

How to eliminate wrong answers

Option A is wrong because Mean Squared Error is a regression metric, not suitable for classification, and L2 regularization prevents overfitting but does not address class imbalance. Option B is wrong because while F1 score is a good metric for imbalanced data, principal component analysis (PCA) is an unsupervised dimensionality reduction technique that can discard important fraud-related features and does not solve class imbalance. Option C is wrong because accuracy is already misleadingly high due to imbalance, and simply collecting more data does not guarantee more fraud samples or fix the skew; it may even worsen the imbalance if the new data has the same distribution.

66
Multi-Selectmedium

Which THREE of the following are techniques for handling missing data in machine learning?

Select 3 answers
A.Deletion of rows with missing values
B.Autoencoder reconstruction
C.Mean imputation
D.Principal Component Analysis
E.Using a separate category for missing values
AnswersA, C, E

Listwise deletion removes incomplete records; a basic approach.

Why this answer

Deleting rows with missing values is a straightforward technique for handling missing data, often used when the missingness is random and the dataset is large enough that removing a few rows does not significantly impact model performance. This method avoids introducing bias from imputation but can lead to loss of valuable information if too many rows are removed.

Exam trap

The CompTIA AI exam often tests the distinction between techniques that directly handle missing data versus those that are preprocessing or modeling steps that assume complete data, leading candidates to mistakenly select PCA or autoencoder reconstruction as missing data methods.

67
MCQhard

A company deploys a machine learning model that makes predictions on streaming data. Over time, the data distribution shifts, causing model performance to degrade. Which monitoring strategy is most appropriate to detect this drift?

A.Compare the distribution of predictions to the training set
B.Monitor the model's training loss
C.Retrain the model daily on new data
D.Track the model's accuracy on a fixed validation set over time
AnswerD

Accuracy drop on a static validation set indicates concept drift.

Why this answer

Tracking the model's accuracy on a fixed validation set over time directly measures performance degradation caused by data drift. As the streaming data distribution shifts, the model's predictions on the static validation set will become less accurate, providing a clear signal that drift has occurred. This is a standard monitoring approach in production ML systems for detecting concept drift.

Exam trap

CompTIA often tests the distinction between monitoring for drift (which requires tracking performance on a fixed baseline) versus retraining or comparing input distributions, leading candidates to mistakenly choose Option A or C.

How to eliminate wrong answers

Option A is wrong because comparing the distribution of predictions to the training set only detects covariate shift (input distribution change) but not concept drift (change in the relationship between inputs and outputs), and it does not directly measure performance degradation. Option B is wrong because monitoring the model's training loss is irrelevant for deployed models; training loss reflects convergence during training, not performance on new streaming data. Option C is wrong because retraining the model daily on new data is a reactive mitigation strategy, not a monitoring strategy; it does not detect drift but instead attempts to fix it blindly, which can be wasteful and may not address the root cause.

68
MCQeasy

A data scientist is building a binary classification model to predict customer churn. The dataset has 10,000 samples with 80% non-churn and 20% churn. The model achieves 95% accuracy but fails to identify churners correctly. Which metric should the scientist focus on to evaluate model performance properly?

A.Precision
B.F1-score
C.Recall (TPR)
D.Specificity
AnswerC

With 80/20 class imbalance, accuracy is misleading because predicting all non-churn yields 80%. Recall measures the proportion of actual churners correctly identified, exposing the model's failure to detect the minority class the business cares about.

Why this answer

Recall (True Positive Rate) measures the proportion of actual churners correctly identified by the model. With 80% non-churn and 20% churn, a model can achieve 95% accuracy by simply predicting the majority class (non-churn) for all samples, resulting in zero true positives for churn. Recall directly exposes this failure by quantifying how many churners are captured, making it the critical metric for imbalanced classification problems.

Exam trap

CompTIA often tests the concept that accuracy is misleading in imbalanced datasets, and candidates mistakenly choose precision or F1-score because they seem more comprehensive, but the question specifically asks for the metric that reveals the model's failure to identify churners, which is recall.

How to eliminate wrong answers

Option A is wrong because precision focuses on the proportion of predicted churners that are actually churners, but a model that predicts very few or no churners can still have high precision if those few predictions are correct, masking the failure to identify churners. Option B is wrong because F1-score is the harmonic mean of precision and recall; while it balances both, it can still be misleadingly high if precision is high but recall is low, and it does not directly highlight the model's inability to detect churners as clearly as recall does. Option D is wrong because specificity (True Negative Rate) measures the proportion of actual non-churners correctly identified, which is already high in a model that predicts majority class, and it does not address the core problem of missing churners.

69
MCQhard

A retail company trains a gradient-boosted tree model to forecast weekly demand for 500 stores. After six months, forecast error rises sharply even though the model code and pipeline are unchanged. Store openings, promotions, and seasonality have shifted the underlying demand patterns. Which action best addresses the root cause?

A.Increase the number of boosting rounds in the existing model.
B.Replace the gradient-boosted model with a deeper neural network.
C.Lower the learning rate and re-evaluate on the original test set.
D.Schedule periodic retraining with recent data and monitor for data drift.
AnswerD

The scenario describes concept and data drift: the relationship between features and demand has changed due to new stores, promotions, and seasonality. Periodic retraining on recent data lets the model relearn current patterns, while drift monitoring detects when retraining is needed. This directly targets the root cause rather than masking symptoms, and it fits a production forecasting pipeline.

Why this answer

When a deployed model's error grows without code changes and the business environment has shifted, the cause is drift between training and production data. Retraining on recent data restores alignment with current patterns, and continuous drift monitoring provides early warning. This is the standard MLOps response to concept drift in forecasting systems.

Exam trap

The trap here is treating degraded accuracy as a hyperparameter problem and tuning the existing model, when the real issue is that the data distribution has moved away from what the model learned.

70
MCQmedium

A hospital wants an AI system to classify chest X-rays as normal or showing pneumonia. Radiologists have labeled 12,000 images, but only 900 show pneumonia. The team must choose a modeling approach that handles this class imbalance. Which approach is most appropriate?

A.Train a regression model to output a continuous pneumonia severity score.
B.Train with a weighted loss function that assigns higher cost to the minority class.
C.Remove 8,000 normal images so the classes are closer in size.
D.Report only accuracy because the dataset is large.
AnswerB

Class weighting in the loss function increases the penalty for misclassifying the underrepresented pneumonia cases, so the model pays more attention to them during training. This directly addresses the 900-versus-11,100 imbalance without discarding data or fabricating examples. It is a standard, well-supported technique for imbalanced medical imaging tasks and preserves the original clinical distribution.

Why this answer

Severe class imbalance causes standard training to favor the majority class because minimizing overall loss is easiest when normal images are predicted correctly. A class-weighted loss raises the cost of minority-class errors, pushing the optimizer to learn pneumonia features. It keeps all data, requires no synthetic generation, and is a widely accepted practice in medical imaging.

Exam trap

The trap here is reaching for undersampling or oversampling as the only fix, when reweighting the loss achieves the same goal without deleting data or fabricating images.

71
MCQmedium

An organization needs to classify customer emails into categories. They have labeled data for some categories but not all. Which approach should they use?

A.Unsupervised clustering then labeling
B.Supervised learning for all categories
C.Reinforcement learning
D.Semi-supervised learning
AnswerD

Semi-supervised learning combines a small labelled set with a larger unlabelled set, using the labelled examples to seed category boundaries and propagating labels through structure in the unlabelled emails. This suits partial category coverage where full supervision is unavailable.

Why this answer

Semi-supervised learning (D) is the correct approach because the organization has labeled data for some categories but not all. This technique leverages a small amount of labeled data to guide the clustering or classification of a larger pool of unlabeled data, effectively combining supervised and unsupervised methods to handle partially labeled datasets.

Exam trap

CompTIA often tests the distinction between semi-supervised and unsupervised learning, trapping candidates who assume that any use of unlabeled data automatically means unsupervised learning, ignoring the critical role of the existing labeled data.

How to eliminate wrong answers

Option A is wrong because unsupervised clustering then labeling would ignore the existing labeled data entirely, wasting valuable information and potentially producing clusters that do not align with the known categories. Option B is wrong because supervised learning requires labeled data for all categories, which the organization does not have, making it impossible to train a model for the unlabeled categories. Option C is wrong because reinforcement learning is designed for sequential decision-making with rewards and penalties, not for static classification tasks like categorizing emails.

72
MCQeasy

A junior ML engineer is asked to evaluate a binary classifier that predicts whether a bank transaction is fraudulent. The model's precision is 0.92 and recall is 0.41. The team wants to improve recall without retraining the model. Which action should the engineer take?

A.Lower the classification threshold for the fraud class.
B.Apply L2 regularization and retrain the classifier.
C.Switch the evaluation metric from F1 score to accuracy.
D.Increase the number of training epochs and re-evaluate.
AnswerA

Lowering the decision threshold makes the model more willing to label a transaction as fraudulent, which increases the number of true positives captured and therefore raises recall. Precision will typically drop as more false positives appear, but the scenario explicitly prioritizes improving recall without retraining, so adjusting the threshold is the correct, direct lever.

Why this answer

Recall measures the fraction of actual positives the model captures. When a model has high precision but low recall, it is being too conservative about predicting the positive class. Lowering the decision threshold shifts the operating point along the precision-recall curve toward higher recall, which is exactly what the scenario asks for without retraining the model.

Exam trap

The trap here is assuming that changing the evaluation metric (such as switching to accuracy) will change the model's behavior, when metrics only measure performance and do not alter predictions.

73
MCQmedium

A machine learning team is deploying a model that predicts customer churn. They notice that the model's predictions are highly sensitive to small changes in input features, leading to inconsistent outputs. Which technique should the team apply to improve model stability?

A.Increase learning rate
B.Feature scaling
C.Regularization
D.Cross-validation
AnswerC

Regularization adds a penalty on large weights to the loss function, constraining the model and reducing its sensitivity to small input perturbations. This directly addresses the instability described, yielding smoother, more consistent predictions than unregularised training, which overfits noise in the churn features.

Why this answer

Regularization (Option C) is the correct technique because it adds a penalty term to the loss function (e.g., L1 or L2 regularization), which constrains the model's weights. This reduces variance and prevents overfitting to noise in the training data, directly addressing the high sensitivity to small input changes (brittleness). By shrinking coefficients, regularization forces the model to learn more general patterns, improving stability and consistency in predictions.

Exam trap

CompTIA often tests the misconception that feature scaling alone can fix model instability, but scaling only normalizes inputs and does not penalize large weights, which is the root cause of sensitivity to small input changes.

How to eliminate wrong answers

Option A is wrong because increasing the learning rate makes gradient descent steps larger, which can cause the model to overshoot minima and increase instability, not reduce sensitivity to input changes. Option B is wrong because feature scaling normalizes input ranges (e.g., via standardization or min-max scaling) to help gradient descent converge faster, but it does not address model variance or overfitting that causes prediction instability. Option D is wrong because cross-validation is a technique for evaluating model performance and tuning hyperparameters, not a method to directly improve model stability or reduce sensitivity to input perturbations.

74
MCQeasy

A data scientist is training a binary classification model to detect fraudulent transactions. The dataset is highly imbalanced with only 1% fraud cases. Which technique is most appropriate to address the class imbalance?

A.Use a linear regression model
B.Oversample the minority class
C.Undersample the majority class
D.Increase the learning rate
AnswerB

Oversampling the minority class replicates or synthesises fraudulent examples so the training set becomes more balanced, letting the model learn fraud patterns rather than predicting the majority class. This satisfies the stem's 1% fraud constraint, unlike accuracy-based metrics that mislead on imbalanced data.

Why this answer

Oversampling the minority class (e.g., using SMOTE or random oversampling) is the most appropriate technique because it balances the dataset by generating synthetic or duplicate examples of the fraud cases, allowing the model to learn the decision boundary for the minority class without discarding valuable majority-class data. This directly addresses the class imbalance where only 1% of transactions are fraudulent, improving recall and precision for fraud detection.

Exam trap

A common misconception is that undersampling is always better because it reduces dataset size and training time, but the trap here is that undersampling discards majority-class data, which can severely degrade model performance when the imbalance is extreme (e.g., 1:99 ratio).

How to eliminate wrong answers

Option A is wrong because linear regression is a regression algorithm, not a classification model, and it cannot output binary class probabilities or handle class imbalance without modification. Option C is wrong because undersampling the majority class discards a large amount of potentially useful non-fraud data, which can lead to loss of information and poor generalization, especially when the imbalance is severe (99% majority). Option D is wrong because increasing the learning rate does not address class imbalance; it only affects the convergence speed of gradient descent and may cause the model to overshoot the optimum, not rebalance the dataset.

75
MCQmedium

Based on the exhibit, what is the most likely issue with the trained model?

A.Overfitting because training accuracy is much higher than validation accuracy
B.Data leakage artificially inflating training accuracy
C.Vanishing gradients causing no learning
D.Underfitting due to insufficient epochs
AnswerA

Overfitting occurs when a model memorises training data, producing high training accuracy but poor generalisation, evidenced by much lower validation accuracy. This directly satisfies the exhibit's constraint: the widening gap between the two accuracy curves. Regularisation, dropout, or more training data would reduce this variance.

Why this answer

The exhibit shows a significant gap between high training accuracy and lower validation accuracy, which is the classic symptom of overfitting. The model has memorized the training data rather than learning generalizable patterns, leading to poor performance on unseen validation data.

Exam trap

The AI0-001 exam often tests the distinction between overfitting and underfitting by presenting accuracy curves where candidates must recognize that high training accuracy with low validation accuracy indicates overfitting, not data leakage or gradient issues.

How to eliminate wrong answers

Option B is wrong because data leakage would cause both training and validation accuracy to be artificially high and closely aligned, not a large gap. Option C is wrong because vanishing gradients prevent the model from learning at all, resulting in both training and validation accuracy remaining low or random, not high training accuracy. Option D is wrong because underfitting due to insufficient epochs would show low accuracy on both training and validation sets, not a high training accuracy with a lower validation accuracy.

Page 1 of 2 · 95 questions totalNext →

Ready to test yourself?

Try a timed practice session using only Machine Learning and Deep Learning questions.