Courseiva

CCNA Machine Learning and Deep Learning Questions

20 of 95 questions · Page 2/2 · Machine Learning and Deep Learning · Answers revealed

76
MCQhard

A financial institution uses a deep learning model for fraud detection. The model is a feedforward neural network with three hidden layers. It was trained on a balanced dataset of 100,000 transactions. During deployment, the model achieves high accuracy on the test set but the fraud detection rate (true positive rate) is only 40% while the false positive rate is 0.1%. The business requires a true positive rate of at least 80%. Which of the following actions is most likely to achieve the required true positive rate while minimizing the increase in false positives?

A.Increase the number of hidden layers to five to capture more complex patterns
B.Use synthetic minority oversampling (SMOTE) to rebalance the training set
C.Change the threshold for classifying a transaction as fraud from the default 0.5 to a lower value
D.Add L2 regularization to reduce overfitting
AnswerC

Lowering threshold increases TPR; the optimal threshold can be chosen based on the precision-recall curve.

Why this answer

(increase hidden layers) may capture more complexity but does not directly increase TPR and could overfit. Option B (SMOTE) rebalances the training set, but the dataset is already balanced, so this is unlikely to improve TPR. Option D (L2 regularization) reduces overfitting but increases bias, which could lower TPR.

Option C (change threshold) is the most direct approach: lowering the classification threshold increases the true positive rate, and by tuning, it can achieve 80% TPR with a minimal increase in false positives.

77
MCQmedium

Based on the exhibit, what is the likely problem with the model?

A.Batch size too small
B.Overfitting
C.Learning rate too high
D.Underfitting
AnswerB

The exhibit shows training performance continuing to improve while validation performance degrades, the classic divergence indicating the model memorises training noise rather than generalising. That gap between training and validation error is the defining signature of overfitting.

Why this answer

The exhibit shows training loss decreasing to near zero while validation loss increases after a certain point, which is a classic sign of overfitting. The model is memorizing the training data rather than learning generalizable patterns, leading to poor performance on unseen data.

Exam trap

CompTIA often tests the distinction between overfitting and underfitting by showing loss curves where candidates mistakenly focus on the low training loss alone, ignoring the rising validation loss that confirms overfitting.

How to eliminate wrong answers

Option A is wrong because a batch size that is too small typically causes noisy gradient updates and slower convergence, not the divergence between training and validation loss seen here. Option C is wrong because a learning rate that is too high usually causes the loss to oscillate or diverge entirely, not a steady decrease in training loss with a rise in validation loss. Option D is wrong because underfitting would show high loss on both training and validation sets, not the low training loss and high validation loss pattern in the exhibit.

78
MCQmedium

A self-driving car company uses a reinforcement learning agent to navigate. The agent was trained in a simulated environment and achieved high rewards. When deployed in the real world, the agent fails to avoid obstacles. The team collects real-world driving data and uses it to fine-tune the model. However, fine-tuning leads to catastrophic forgetting of the simulated knowledge. Which technique should the team use to mitigate this? A. Increase the learning rate during fine-tuning. B. Use elastic weight consolidation (EWC) to regularize important weights. C. Train the model from scratch using only real-world data. D. Increase the number of layers in the network.

A.Increase the number of layers in the network.
B.Use elastic weight consolidation (EWC) to regularize important weights.
C.Train the model from scratch using only real-world data.
D.Increase the learning rate during fine-tuning.
AnswerB

Elastic weight consolidation adds a regularisation penalty that protects weights important to previously learned simulated tasks, allowing real-world fine-tuning without overwriting that knowledge. This directly counteracts catastrophic forgetting, satisfying the stem's requirement to retain simulation knowledge while adapting to real driving.

Why this answer

Elastic Weight Consolidation (EWC) is a regularization technique specifically designed to prevent catastrophic forgetting when fine-tuning a neural network on a new task. It identifies the weights that are most important for the original task (simulated driving) and penalizes large changes to those weights during fine-tuning on real-world data, thereby preserving the learned knowledge while adapting to the new domain.

Exam trap

CompTIA often tests the concept of catastrophic forgetting by presenting fine-tuning as a solution and then offering tempting but incorrect options like increasing learning rate or network depth, which candidates might mistakenly associate with improving generalization or capacity.

How to eliminate wrong answers

Option A is wrong because increasing the number of layers in the network does not address catastrophic forgetting; it adds capacity but does not constrain updates to important weights, and may even worsen overfitting. Option C is wrong because training from scratch using only real-world data discards all the valuable simulated knowledge, which is the opposite of mitigating forgetting and would likely require much more real-world data to achieve comparable performance. Option D is wrong because increasing the learning rate during fine-tuning would cause larger weight updates, accelerating the overwriting of previously learned knowledge and exacerbating catastrophic forgetting, not mitigating it.

79
MCQhard

A machine learning engineer notices that the gradient values in a deep network are becoming extremely small during backpropagation. What is this problem?

A.Dead ReLU
B.Exploding gradient
C.Covariate shift
D.Vanishing gradient
AnswerD

Repeated multiplication of small derivatives during backpropagation shrinks gradients exponentially with depth, so early layers barely update. This matches the stem's observation of extremely small gradient values, distinguishing it from exploding gradients, where values grow instead.

Why this answer

The vanishing gradient problem occurs when gradients become extremely small during backpropagation, especially in deep networks with many layers. This causes the weights in earlier layers to update very slowly or not at all, severely hindering training. The correct answer is D because the scenario directly describes the hallmark symptom of vanishing gradients.

Exam trap

The AI0-001 exam often tests the distinction between vanishing and exploding gradients by describing the symptom (small vs. large gradients) and expects candidates to recognize that vanishing gradients cause slow learning in early layers, not just any training difficulty.

How to eliminate wrong answers

Option A is wrong because Dead ReLU refers to neurons that become permanently inactive (outputting zero) due to negative inputs, not to gradients becoming small across the network. Option B is wrong because exploding gradient is the opposite problem, where gradients grow exponentially large, causing unstable updates and NaN values. Option C is wrong because covariate shift is a change in the input distribution between training and test data, addressed by batch normalization, and is unrelated to gradient magnitude during backpropagation.

80
MCQeasy

A data scientist wants to reduce the dimensionality of a dataset with 200 features before training a regression model. Which technique should they use?

A.LDA
B.t-SNE
C.Autoencoder
D.PCA
AnswerD

Principal component analysis projects the 200 correlated features onto a smaller set of orthogonal components that retain most variance, reducing dimensionality before regression. It is the standard unsupervised linear technique for this purpose, unlike feature-selection methods that simply drop columns.

Why this answer

PCA (Principal Component Analysis) is the correct technique because it is an unsupervised linear dimensionality reduction method that identifies the directions (principal components) of maximum variance in the data. For a dataset with 200 features, PCA can reduce dimensionality while preserving as much variance as possible, which is ideal before training a regression model to avoid overfitting and multicollinearity.

Exam trap

CompTIA often tests the distinction between supervised and unsupervised techniques, and the trap here is that candidates confuse LDA (supervised, classification) with PCA (unsupervised, regression-friendly) because both are linear methods for dimensionality reduction.

How to eliminate wrong answers

Option A is wrong because LDA (Linear Discriminant Analysis) is a supervised dimensionality reduction technique that requires class labels and maximizes class separability, making it unsuitable for a regression task where the target is continuous. Option B is wrong because t-SNE (t-distributed Stochastic Neighbor Embedding) is a non-linear visualization technique that does not preserve global structure or distances, and it cannot be used to transform new data for a regression model. Option C is wrong because autoencoders are neural network-based non-linear dimensionality reduction methods that require significant data and tuning, and they are not the standard first-choice technique for simple linear dimensionality reduction before regression.

81
MCQeasy

A machine learning engineer notices that a linear regression model has high bias. Which action is most likely to reduce bias?

A.Use a more complex model, such as polynomial regression
B.Reduce the number of training samples
C.Add L2 regularization
D.Apply feature scaling
AnswerA

High bias means the model underfits, so increasing capacity with polynomial regression lets it capture non-linear relationships the linear hypothesis cannot represent, reducing bias. Note this typically raises variance, trading one error source for the other.

Why this answer

High bias indicates that the model is too simple to capture the underlying patterns in the data, leading to underfitting. Using a more complex model, such as polynomial regression, increases the model's capacity to fit the training data better, directly addressing the underfitting issue. This is the standard approach to reduce bias in machine learning.

Exam trap

CompTIA often tests the bias-variance tradeoff by making candidates confuse bias-reduction techniques with variance-reduction techniques, such as regularization or reducing training data, which actually increase bias or do not affect it.

How to eliminate wrong answers

Option B is wrong because reducing the number of training samples typically increases variance and does not reduce bias; it can actually worsen underfitting by providing less data for the model to learn from. Option C is wrong because adding L2 regularization (Ridge regression) penalizes large coefficients, which increases bias by constraining the model, making it simpler and potentially worsening underfitting. Option D is wrong because feature scaling (e.g., normalization or standardization) does not change the model's complexity or bias; it only helps gradient descent converge faster and is irrelevant for bias reduction.

82
MCQmedium

A data scientist is training a binary classification model to detect fraudulent transactions. The dataset contains 99.9% legitimate transactions and 0.1% fraudulent transactions. After training a logistic regression model, the accuracy is 99.9%, but the recall for the fraud class is 0%. Which of the following is the MOST likely cause?

A.The regularization parameter is too large, causing underfitting.
B.The model is overfitting due to too many features.
C.The learning rate was too high.
D.The dataset is highly imbalanced, and the model predicts the majority class for all instances.
AnswerD

With 0.1% fraud, a model predicting legitimate for every transaction scores 99.9% accuracy yet catches no fraud, giving 0% recall. The extreme class imbalance dominates the loss function, so the model never learns the minority pattern. Resampling or class weighting is required.

Why this answer

The dataset has a severe class imbalance (99.9% legitimate, 0.1% fraudulent). A logistic regression model that predicts the majority class (legitimate) for every instance will achieve 99.9% accuracy but 0% recall for the fraud class, because it never identifies any positive fraud cases. This is the classic 'accuracy paradox' in imbalanced classification.

Exam trap

CompTIA often tests the 'accuracy paradox' where candidates mistakenly attribute high accuracy to model quality, ignoring that in imbalanced datasets, a dummy classifier predicting the majority class can achieve the same accuracy, and the trap is to overlook recall or precision for the minority class.

How to eliminate wrong answers

Option A is wrong because a large regularization parameter (e.g., high L2 penalty) causes underfitting by shrinking coefficients too much, but the model here is not underfitting—it is perfectly fitting the majority class, which is a different failure mode. Option B is wrong because overfitting due to too many features would typically cause high variance and poor generalization, not a perfect 99.9% accuracy with 0% recall on the minority class; overfitting would likely memorize some fraud examples. Option C is wrong because a learning rate that is too high would cause the model's loss to diverge or oscillate, not converge to a trivial majority-class predictor with high accuracy.

83
MCQmedium

A model trained on a dataset has high bias and low variance. What does this indicate?

A.Good fit
B.Data leakage
C.Overfitting
D.Underfitting
AnswerD

High bias means the model is too simple to capture the underlying pattern, and low variance means its predictions stay consistently wrong across samples. Together these signal underfitting, where training and validation errors both remain high.

Why this answer

High bias and low variance indicate that the model is too simple to capture the underlying patterns in the data, leading to systematic errors on both training and test sets. This is the classic signature of underfitting, where the model fails to learn the training data adequately.

Exam trap

The CompTIA AI exam often tests the bias-variance tradeoff by reversing the definitions, so candidates mistakenly associate high bias with overfitting or high variance with underfitting.

How to eliminate wrong answers

Option A is wrong because a good fit requires low bias and low variance, not high bias. Option B is wrong because data leakage typically causes overly optimistic performance metrics, not a high-bias, low-variance error pattern. Option C is wrong because overfitting is characterized by low bias and high variance, the exact opposite of the given condition.

84
MCQeasy

A dataset contains features on vastly different scales (e.g., age 0-100 vs. income 0-1,000,000). Which preprocessing step is essential before training a neural network?

A.Data augmentation
B.Dimensionality reduction
C.Feature scaling (standardization or normalization)
D.One-hot encoding
AnswerC

Features on vastly different scales cause gradient updates to be dominated by large-magnitude inputs, slowing or destabilising training. Standardisation or normalisation rescales each feature to a comparable range, ensuring the network converges efficiently and weights are not biased toward high-range variables such as income.

Why this answer

Neural networks rely on gradient-based optimization, where features with larger scales can dominate the weight updates, causing unstable convergence or slow training. Feature scaling (standardization or normalization) ensures all features contribute equally to the loss function, preventing the model from being biased toward high-magnitude features like income versus age.

Exam trap

CompTIA AI often tests the misconception that data augmentation or dimensionality reduction can substitute for feature scaling, when in fact scaling is a prerequisite for stable gradient descent in neural networks.

How to eliminate wrong answers

Option A is wrong because data augmentation is a technique to artificially increase dataset size by creating modified copies of data (e.g., rotations, flips for images), not to address scale differences among features. Option B is wrong because dimensionality reduction (e.g., PCA) reduces the number of features to combat the curse of dimensionality or noise, but it does not equalize the scales of existing features; scaling is still required before or after reduction. Option D is wrong because one-hot encoding is used to convert categorical variables into binary vectors, not to handle numerical features with differing magnitudes.

85
MCQmedium

A team trained a deep neural network on a limited dataset. The training loss decreases consistently, but the validation loss starts increasing after 20 epochs. What is the most likely issue and the best corrective action?

A.Vanishing gradient; use ReLU activation
B.Overfitting; apply regularization like dropout
C.Underfitting; increase model complexity
D.Data leakage; reshuffle split
AnswerB

Dropout randomly deactivates neurons during training, forcing the network to learn redundant, generalisable features rather than memorising the limited samples. This directly counteracts the diverging validation loss after epoch 20, which signals overfitting. Regularisation constrains model capacity, restoring alignment between training and validation performance.

Why this answer

The training loss decreasing while validation loss increasing after 20 epochs is the classic signature of overfitting: the model has memorized the training data but fails to generalize to unseen data. Applying regularization like dropout forces the network to learn more robust features by randomly dropping neurons during training, reducing overfitting. This is the most direct and effective corrective action for this specific symptom.

Exam trap

CompTIA often tests the distinction between overfitting and vanishing gradients by showing a loss curve that decreases initially then rises, tricking candidates into thinking the gradient is vanishing when the real issue is poor generalization.

How to eliminate wrong answers

Option A is wrong because vanishing gradient causes the training loss to stagnate or decrease very slowly from the start, not a divergence between training and validation loss after many epochs; ReLU helps mitigate vanishing gradients but does not address overfitting. Option C is wrong because underfitting would show both training and validation loss remaining high or not decreasing, and increasing model complexity would worsen overfitting, not fix it. Option D is wrong because data leakage typically causes both training and validation loss to be artificially low and correlated, not a divergence after a certain number of epochs; reshuffling the split does not address the core issue of model memorization.

86
MCQhard

An ML engineer is tuning a random forest classifier for a medical diagnosis task and observes that training accuracy is 99% while validation accuracy is 78%. She wants to reduce the gap without discarding the ensemble approach. Which change is most likely to reduce the generalization gap?

A.Add more trees to the ensemble until the out-of-bag error stops decreasing.
B.Increase the maximum tree depth so each tree can capture more interactions in the training data.
C.Reduce the maximum tree depth and increase the minimum number of samples required at each leaf.
D.Disable bootstrap sampling so every tree is trained on the full dataset instead of a random subset.
AnswerC

A 99% versus 78% split is classic overfitting: individual trees are fitting noise. Limiting depth and raising the minimum samples per leaf constrains tree complexity, forcing each tree to learn broader, more generalizable splits. Combined with the ensemble averaging of a random forest, this typically narrows the train-validation gap while preserving the benefits of bagging and feature subsampling.

Why this answer

The large train-validation gap indicates that the individual trees are too complex and memorizing the training set. Constraining tree depth and requiring more samples per leaf regularizes each estimator. Because a random forest averages many decorrelated trees, this regularization typically improves validation accuracy while keeping the ensemble's variance-reduction benefits intact.

Exam trap

The trap here is assuming that adding more trees or more depth will always improve a random forest, when the gap points to per-tree overfitting that only complexity limits can fix.

87
Multi-Selectmedium

A deep learning engineer is training a convolutional neural network for image classification. The model is overfitting the training data. Which three techniques can help reduce overfitting? (Choose three.)

Select 3 answers
A.Add dropout layers
B.Apply L2 regularization
C.Use data augmentation
D.Use a smaller learning rate
E.Increase the number of convolutional layers
AnswersA, B, C

Dropout randomly drops units during training, reducing co-adaptation.

Why this answer

Dropout layers randomly deactivate a fraction of neurons during training, which prevents co-adaptation and forces the network to learn more robust features. This reduces overfitting by acting as a form of ensemble learning without increasing model complexity.

Exam trap

In CompTIA AI+ exams, a common trap is assuming that reducing the learning rate or increasing model depth can mitigate overfitting, when in fact these adjustments either have no effect on overfitting or worsen it.

88
MCQeasy

A junior data scientist is asked to explain the difference between supervised and unsupervised learning to a product manager. She wants to give a single concrete example that clearly illustrates unsupervised learning. Which example should she choose?

A.Grouping retail customers into segments based on purchasing behavior using k-means clustering.
B.Predicting whether a loan applicant will default using historical labeled outcomes.
C.Training an agent to play a board game by rewarding winning moves over many episodes.
D.Estimating a house's sale price from its square footage and location features.
AnswerA

K-means clustering groups unlabeled data points based on similarity in feature space, with no target variable provided. Customer segmentation is the canonical unsupervised example because the algorithm discovers structure in purchasing behavior rather than predicting a known label. This directly illustrates the defining characteristic of unsupervised learning: learning patterns from data without ground-truth outputs.

Why this answer

Unsupervised learning finds structure in data that has no labels. Clustering customers by purchasing behavior is a textbook example because k-means discovers groups from feature similarity alone. The other choices involve a known target (default status, sale price) or a reward signal (game playing), which place them in supervised or reinforcement learning rather than unsupervised learning.

Exam trap

The trap here is conflating any 'data-driven' task with unsupervised learning, when the presence or absence of labeled targets is what actually distinguishes the paradigms.

89
MCQhard

A financial firm trained a gradient boosting model on two years of loan data. It reported strong AUC during development, but after six months in production, approval rates for a newly launched loan product diverge sharply from expectations. The data science lead suspects the model is stale. Which approach best addresses this deployment issue?

A.Retrain the model on the original two-year dataset with a lower learning rate
B.Switch the evaluation metric from AUC to accuracy without changing the data
C.Increase the number of boosting rounds to improve fit on the existing data
D.Monitor prediction distributions and retrain on recent labeled data when drift is detected
AnswerD

This combines drift detection with periodic retraining on fresh, labeled outcomes, which is the standard remedy for data and concept drift. Monitoring feature and prediction distributions catches when the new loan product shifts the input space, and retraining on recent labels updates the model's learned relationships. It directly addresses staleness while preserving the ability to validate the refreshed model before redeployment in a regulated lending context.

Why this answer

The symptoms describe drift: a new product changes the applicant distribution and recent economic conditions change the relationship between features and default. The durable fix is to monitor for drift and retrain on recent labeled data, then revalidate before deployment. Changing hyperparameters, adding boosting rounds, or swapping metrics leaves the model trained on outdated patterns and cannot restore alignment with current production behavior.

Exam trap

The trap here is treating a production performance drop as a tuning problem, when the real cause is that the training distribution no longer matches the live applicant population.

90
MCQhard

A team is building a model to predict stock prices based on time series data. They need to capture long-term dependencies and avoid vanishing gradients. Which architecture is best suited?

A.Standard RNN
B.LSTM
C.Autoencoder
D.CNN
AnswerB

LSTM networks use gating mechanisms and a cell state that preserve information across many time steps, directly addressing vanishing gradients. This lets the model capture long-term dependencies in time series data, which standard recurrent architectures cannot sustain.

Why this answer

LSTM (Long Short-Term Memory) networks are specifically designed to capture long-term dependencies in sequential data through their gating mechanisms (input, forget, and output gates), which regulate the flow of information and mitigate the vanishing gradient problem that plagues standard RNNs. This makes them ideal for time series forecasting tasks like stock price prediction, where historical context over many time steps is critical.

Exam trap

CompTIA often tests the misconception that any recurrent architecture (like a standard RNN) can handle long sequences, when in fact only gated variants like LSTM or GRU are designed to overcome vanishing gradients in practice.

How to eliminate wrong answers

Option A (Standard RNN) is wrong because it suffers from the vanishing gradient problem during backpropagation through time, making it unable to effectively learn long-term dependencies in sequences longer than about 10–20 steps. Option C (Autoencoder) is wrong because it is an unsupervised learning architecture designed for dimensionality reduction or feature learning (e.g., anomaly detection), not for sequential prediction or capturing temporal dependencies. Option D (CNN) is wrong because while CNNs can be used for time series via 1D convolutions, they lack inherent memory mechanisms and are not optimized for capturing long-range temporal dependencies without extensive stacking or dilation, and they do not directly address vanishing gradients in the same way as LSTMs.

91
MCQeasy

A team wants to predict monthly sales using historical data. Which algorithm is most appropriate?

A.Linear regression
B.K-means
C.Decision tree
D.Logistic regression
AnswerA

Monthly sales is a continuous numeric target, so regression rather than classification applies. Linear regression models the relationship between historical input variables and the sales figure directly, producing a predicted value. It satisfies the forecasting constraint with a simple, interpretable model suited to this trend-based prediction task.

Why this answer

Linear regression is the most appropriate algorithm because the goal is to predict a continuous numerical value (monthly sales) based on historical data. It models the relationship between input features and the target variable by fitting a linear equation, making it ideal for regression tasks where the output is a real number.

Exam trap

CompTIA often tests the distinction between regression and classification algorithms, and the trap here is that candidates may confuse 'regression' in logistic regression with continuous prediction, not realizing it is actually a classification algorithm.

How to eliminate wrong answers

Option B (K-means) is wrong because it is an unsupervised clustering algorithm used to group unlabeled data into clusters, not for predicting continuous values. Option C (Decision tree) is wrong because while it can handle regression, it is more prone to overfitting and less optimal for simple linear relationships compared to linear regression; it is not the most appropriate choice for this straightforward prediction task. Option D (Logistic regression) is wrong because it is designed for binary classification problems, predicting probabilities of discrete outcomes, not continuous numerical values like sales.

92
Multi-Selecthard

Which THREE of the following are best practices for preventing overfitting in deep learning models?

Select 3 answers
A.L2 regularization
B.Increasing the number of layers
C.Dropout
D.Using a larger batch size
E.Data augmentation
AnswersA, C, E

L2 adds penalty on weights, keeping them small and reducing overfitting.

Why this answer

L2 regularization (also known as weight decay) adds a penalty term proportional to the square of the weight magnitudes to the loss function. This discourages the model from learning overly complex patterns by forcing weights to remain small, which reduces variance and helps prevent overfitting. It is a standard technique in deep learning frameworks like TensorFlow and PyTorch, where it is implemented via the `kernel_regularizer` or `weight_decay` parameter.

Exam trap

The AI0-001 exam often tests the misconception that increasing model complexity (e.g., more layers) or adjusting batch size are regularization techniques, when in fact they either worsen overfitting or serve different purposes like optimization speed.

93
Multi-Selecteasy

A data scientist is tuning hyperparameters for a support vector machine (SVM) with an RBF kernel. Which two hyperparameters most significantly affect model performance? (Select TWO.)

Select 2 answers
A.gamma (kernel coefficient)
B.learning rate
C.epsilon (for epsilon-SVR)
D.degree (for polynomial kernel)
E.C (regularization parameter)
AnswersA, E

gamma determines the radius of influence of support vectors.

Why this answer

Gamma defines the influence of a single training example, with low values meaning a far reach and high values meaning a close reach. C controls the trade-off between achieving a low error on the training data and minimizing the margin, directly impacting overfitting. Together, they are the two most critical hyperparameters for an SVM with an RBF kernel.

Exam trap

Candidates often mistake kernel-specific hyperparameters (e.g., degree for polynomial, gamma for RBF) for general SVM parameters, selecting options like degree or epsilon without realizing they do not apply to the RBF kernel.

94
MCQmedium

A retail company uses a gradient boosting model to predict customer lifetime value (CLV). The model currently uses 50 features including purchase history, demographics, and web behavior. The model's RMSE on the test set is 120. The data science team wants to improve the model's accuracy without increasing training time significantly. They have access to additional data: customer support interaction logs (text), social media sentiment (text), and third-party credit scores (numeric). They also have the ability to perform feature engineering, hyperparameter tuning, and ensemble methods. Which approach is most likely to yield the best improvement in predictive performance with minimal increase in training time?

A.Add the customer support text as a feature using TF-IDF vectors
B.Use an ensemble of gradient boosting and random forest models
C.Perform hyperparameter tuning using grid search
D.Engineer new features such as average purchase value and recency
AnswerD

Feature engineering can capture patterns without adding new data sources or significant time.

Why this answer

Engineering domain-relevant features like average purchase value and recency directly captures the underlying behavioral patterns that drive customer lifetime value, often providing a higher signal-to-noise ratio than adding raw text or third-party data. This approach leverages existing data without significantly increasing the feature dimensionality or training time, unlike adding TF-IDF vectors which would dramatically expand the feature space and slow training.

Exam trap

CompTIA often tests the misconception that adding more data (especially text) or complex ensemble methods always improves model accuracy, while the correct approach is to engineer features that capture domain-specific patterns with minimal computational overhead.

How to eliminate wrong answers

Option A is wrong because adding customer support text as TF-IDF vectors would introduce thousands of sparse features, significantly increasing training time and risking overfitting without guaranteed improvement in RMSE. Option B is wrong because ensembling gradient boosting with random forest typically increases training time substantially (both models must be trained) and may not outperform a well-tuned single gradient boosting model on structured data. Option C is wrong because hyperparameter tuning using grid search is computationally expensive, often requiring many model fits, and would increase training time more than feature engineering without leveraging the new data sources.

95
MCQhard

A machine learning engineer is training a deep neural network for image classification. The training loss decreases steadily, but the validation loss starts to increase after 20 epochs. The engineer wants to implement a technique that dynamically adjusts the learning rate during training to improve convergence and generalization. Which method should the engineer use?

A.Stochastic Gradient Descent (SGD) with a fixed learning rate
B.Batch Normalization
C.Dropout regularization
D.Learning Rate Scheduler with exponential decay
AnswerD

An exponential decay learning rate scheduler reduces the learning rate over time, which can help the model settle into a deeper minimum and improve generalization. As training progresses, smaller learning rates allow finer adjustments to weights, potentially mitigating the increasing validation loss by preventing overshooting and encouraging convergence to a smoother minimum.

Why this answer

A learning rate scheduler with exponential decay dynamically reduces the learning rate as training progresses, which helps the model converge more smoothly and can improve generalization. This directly addresses the need to adjust the learning rate during training to combat the rising validation loss, unlike fixed learning rates or other regularization techniques that do not modify the learning rate.

Exam trap

The trap here is confusing regularization techniques like dropout with optimization techniques that adjust the learning rate, even though both can help with overfitting.

← PreviousPage 2 of 2 · 95 questions total

Ready to test yourself?

Try a timed practice session using only Machine Learning and Deep Learning questions.