Courseiva

CCNA Ai Concepts Foundations Questions

25 of 100 questions · Page 2/2 · Ai Concepts Foundations topic · Answers revealed

76
MCQhard

An AI team notices that their model's performance degrades over time because the statistical relationship between input features and the target variable changes. This issue is called:

A.Data drift
B.Overfitting
C.Concept drift
D.Model drift
AnswerC

Concept drift is the change in the statistical relationship between input features and the target variable over time, degrading model performance. This matches the stem exactly, unlike data drift, which shifts input distributions while the underlying mapping stays constant.

Why this answer

Concept drift occurs when the statistical relationship between input features and the target variable changes over time, causing model performance to degrade. This is distinct from data drift, which involves changes in the input data distribution alone. In the AI0-001 context, concept drift directly addresses the shift in the underlying mapping from features to labels.

Exam trap

CompTIA often tests the distinction between data drift and concept drift, where candidates mistakenly choose data drift because they focus on the input features changing, rather than the relationship between features and the target.

How to eliminate wrong answers

Option A is wrong because data drift refers to changes in the distribution of input features, not the relationship between features and the target. Option B is wrong because overfitting is a model that memorizes training data noise and fails to generalize, not a temporal degradation due to shifting relationships. Option D is wrong because 'model drift' is not a standard term in machine learning; the correct term for the described phenomenon is concept drift.

77
Multi-Selecteasy

Which TWO of the following are common techniques to reduce overfitting in a neural network?

Select 2 answers
A.Increasing the number of hidden layers
B.Using a larger learning rate
C.L2 regularization
D.Training for more epochs
E.Dropout
AnswersC, E

Correct; L2 regularization adds a penalty on squared weights.

Why this answer

L2 regularization (option C) reduces overfitting by adding a penalty term proportional to the squared magnitude of the weights to the loss function. This forces the network to keep weights small, preventing it from fitting noise in the training data and improving generalization.

Exam trap

CompTIA often tests the misconception that adding more layers or training longer always improves accuracy, when in fact these actions typically increase overfitting without proper regularization or validation monitoring.

78
Multi-Selecthard

Which THREE of the following are key considerations when deploying an AI model in a production environment?

Select 3 answers
A.Basing acceptance solely on training accuracy
B.Maximizing model complexity to achieve the best accuracy
C.Monitoring model performance for data drift
D.Ensuring inference latency meets service-level agreements
E.Providing explainability for model decisions
AnswersC, D, E

Correct; models degrade over time if data changes.

Why this answer

Data drift refers to the change in the statistical properties of the input data over time, which can degrade model accuracy. Continuous monitoring for data drift is essential in production to detect when the model's assumptions about the data distribution are no longer valid, triggering retraining or alerts.

Exam trap

CompTIA often tests the misconception that high training accuracy is the primary goal for production deployment, when in reality operational concerns like latency, explainability, and drift monitoring are prioritized over raw accuracy.

79
MCQmedium

A hospital uses an AI system to predict patient deterioration from vital signs. The system currently uses a logistic regression model trained on data from the past year. Recently, the hospital adopted a new patient monitoring device that provides more accurate readings. The model's performance has dropped significantly. The data science team has access to the new device's data for the past month and wants to improve the model with minimal disruption. The team also wants to ensure the model remains interpretable for regulatory compliance. Which approach should they take?

A.Retrain the logistic regression model on a combined dataset of old and new device data
B.Continue using the current model and manually adjust predictions based on device differences
C.Build an ensemble of logistic regression and a neural network using new data only
D.Replace the logistic regression model with a gradient boosting model using only new device data
AnswerA

Retraining the logistic regression on combined old and new device data adapts the model to the new readings while retaining its interpretable, coefficient-based form, satisfying the regulatory requirement. More complex models would improve accuracy but sacrifice the interpretability compliance demands.

Why this answer

Retraining the logistic regression model on a combined dataset of old and new device data is the best approach because it leverages all available data to adapt the model to the new device's measurement distribution while preserving the model's inherent interpretability. Logistic regression is a linear model that remains fully transparent for regulatory compliance, and combining both datasets helps the model learn the systematic shift in vital sign readings without discarding valuable historical patterns. This minimizes disruption by avoiding a complete overhaul and directly addresses the performance drop caused by the change in input data distribution.

Exam trap

CompTIA often tests the trade-off between model performance and interpretability, and the trap here is that candidates may prioritize performance gains from complex models (like gradient boosting or neural networks) without recognizing that regulatory compliance mandates interpretability, making logistic regression the only viable choice despite its simplicity.

How to eliminate wrong answers

Option B is wrong because manually adjusting predictions based on device differences is ad-hoc, non-scalable, and introduces subjective bias, which undermines both reliability and regulatory compliance. Option C is wrong because building an ensemble with a neural network reduces interpretability, violating the regulatory requirement, and using only new data ignores valuable historical patterns, leading to overfitting and poor generalization. Option D is wrong because replacing logistic regression with a gradient boosting model sacrifices interpretability, as gradient boosting is a black-box model, and training only on one month of new data risks overfitting and fails to capture long-term trends.

80
MCQmedium

A team deploying an AI model for real-time fraud detection notices that inference latency is too high. The model is a deep neural network with 50 layers, deployed on a cloud GPU. Which of the following is the BEST approach to reduce latency while maintaining acceptable accuracy?

A.Deploy the model on a more powerful GPU.
B.Reduce the batch size for inference.
C.Replace the DNN with a logistic regression model.
D.Apply knowledge distillation to create a smaller model.
AnswerD

Knowledge distillation trains a smaller student network to mimic the 50-layer teacher, cutting inference computation and latency while retaining most accuracy. This satisfies the stem's constraint by reducing depth and parameters rather than merely quantising or batching the existing model.

Why this answer

Knowledge distillation trains a smaller 'student' model to mimic the behavior of a larger 'teacher' model, significantly reducing the number of parameters and layers while preserving most of the original accuracy. This directly addresses the high inference latency caused by the 50-layer DNN by producing a compact model that runs faster on the same GPU hardware.

Exam trap

CompTIA often tests the misconception that simply upgrading hardware or reducing batch size is the best latency fix, when in fact architectural compression techniques like knowledge distillation are the most effective for deep models with strict latency budgets.

How to eliminate wrong answers

Option A is wrong because upgrading to a more powerful GPU only provides a linear speedup and does not address the fundamental architectural overhead of a 50-layer network; it also increases cost without guaranteeing latency targets. Option B is wrong because reducing batch size actually increases the number of inference passes per transaction, which can increase per-request latency due to underutilized GPU parallelism and higher overhead from frequent kernel launches. Option C is wrong because replacing the DNN with logistic regression would cause a catastrophic drop in accuracy for complex fraud patterns, as logistic regression cannot model non-linear interactions and high-dimensional feature spaces that the DNN captures.

81
MCQhard

A self-driving car company is developing an object detection system using a convolutional neural network (CNN). The system needs to detect pedestrians and vehicles in real-time with high accuracy. Which technique can reduce inference time while maintaining accuracy?

A.Apply model pruning and quantization
B.Use a pre-trained model and fine-tune it
C.Add more convolutional layers
D.Increase number of filters in each layer
AnswerA

Pruning removes redundant weights and quantisation reduces numerical precision of remaining parameters, shrinking model size and memory bandwidth demands. Both cut inference latency on embedded automotive hardware while preserving detection accuracy, satisfying the real-time constraint that a full-precision CNN cannot meet.

Why this answer

Model pruning removes redundant or less important weights from the CNN, reducing computational load, while quantization converts floating-point weights to lower-precision integers (e.g., INT8). Together, they shrink model size and speed up inference without significantly degrading accuracy, making them ideal for real-time object detection in resource-constrained environments like autonomous vehicles.

Exam trap

CompTIA often tests the misconception that adding more layers or filters always improves performance, when in fact it increases latency and resource usage, while pruning and quantization are the standard techniques for reducing inference time without sacrificing accuracy.

How to eliminate wrong answers

Option B is wrong because fine-tuning a pre-trained model improves accuracy for a specific task but does not inherently reduce inference time; it may even increase it if the model remains large. Option C is wrong because adding more convolutional layers increases the network depth and computational cost, which slows inference and can cause overfitting without careful regularization. Option D is wrong because increasing the number of filters in each layer expands the feature map channels, raising the number of parameters and FLOPs, which directly increases inference time.

82
MCQhard

A data scientist splits a dataset into training (80%) and test (20%). After training, the model achieves 95% accuracy on training and 60% on test. Which step should the data scientist take first?

A.Collect more data
B.Use cross-validation
C.Apply regularization
D.Increase model complexity
AnswerC

Regularization penalizes large weights, reducing overfitting.

Why this answer

The model shows high training accuracy (95%) but significantly lower test accuracy (60%), which is a classic sign of overfitting. Regularization (Option C) directly addresses overfitting by adding a penalty term to the loss function (e.g., L1 or L2 regularization), discouraging the model from learning overly complex patterns that do not generalize. This is the first step because it targets the core issue without requiring additional data or increasing complexity.

Exam trap

CompTIA often tests the misconception that overfitting is always solved by more data or cross-validation, but the immediate corrective action is to apply regularization to penalize model complexity.

How to eliminate wrong answers

Option A is wrong because collecting more data can help reduce overfitting, but it is not the first step; regularization is a simpler, more immediate fix that does not depend on data availability. Option B is wrong because cross-validation is a technique for model evaluation and hyperparameter tuning, not a direct remedy for overfitting; it would help assess the severity but does not solve the underlying problem. Option D is wrong because increasing model complexity would worsen overfitting, as it allows the model to fit noise even more closely, further reducing test accuracy.

83
MCQhard

A team is training a deep learning model for natural language processing using a large corpus. They notice the model has a very high number of parameters and training is slow. Which technique can reduce the number of parameters without significant performance loss?

A.Apply embedding compression
B.Add more dropout layers
C.Use a larger batch size
D.Increase learning rate
AnswerA

Embedding compression reduces the dimensionality or parameter count of embedding layers through techniques such as low-rank factorisation or weight sharing, directly cutting total parameters. This addresses the slow training caused by the oversized embedding table while retaining semantic representation quality.

Why this answer

Embedding compression reduces the dimensionality of the embedding layer, which often contains the majority of the model's parameters in NLP tasks. By using techniques like low-rank factorization or pruning, the model retains most of its representational power while significantly decreasing the parameter count and training time.

Exam trap

The trap here is that candidates confuse regularization techniques (like dropout) or training speed optimizations (batch size, learning rate) with actual parameter reduction, which only embedding compression directly achieves.

How to eliminate wrong answers

Option B is wrong because adding more dropout layers does not reduce the number of parameters; it only randomly drops neurons during training to prevent overfitting, leaving the parameter count unchanged. Option C is wrong because using a larger batch size improves training speed through better hardware utilization but does not reduce the number of parameters. Option D is wrong because increasing the learning rate can speed up convergence but does not affect the parameter count and may cause training instability or divergence.

84
MCQeasy

A company wants to use AI to analyze customer reviews and determine sentiment (positive, negative, neutral). Which AI subfield is most directly applicable?

A.Reinforcement learning
B.Computer vision
C.Natural language processing
D.Robotics
AnswerC

Natural language processing handles unstructured text, enabling tokenisation, feature extraction and classification of review content into positive, negative or neutral sentiment. It directly addresses the textual analysis the scenario requires, unlike vision or speech subfields.

Why this answer

Natural language processing (NLP) is the AI subfield that enables machines to understand, interpret, and generate human language. Analyzing customer reviews for sentiment requires processing text, extracting meaning, and classifying it as positive, negative, or neutral, which is a core NLP task called sentiment analysis.

Exam trap

The trap here is that candidates often confuse natural language processing with computer vision or reinforcement learning because they see 'AI' broadly, but the specific task of analyzing text directly maps to NLP, not the other subfields.

How to eliminate wrong answers

Option A is wrong because reinforcement learning is a training paradigm where an agent learns by interacting with an environment and receiving rewards or penalties; it is not designed for text classification or sentiment analysis. Option B is wrong because computer vision focuses on interpreting visual data such as images and videos, not textual content like customer reviews. Option D is wrong because robotics deals with the design and control of physical machines to perform tasks in the real world, which is unrelated to analyzing text-based sentiment.

85
MCQeasy

A marketing team uses a recommendation system to suggest products to customers. The system currently uses collaborative filtering. Which scenario would most likely cause the cold-start problem?

A.A new product is added to the catalog with no purchase history.
B.The system switches from collaborative filtering to content-based filtering.
C.The website interface is redesigned, affecting user navigation.
D.A seasonal product experiences a sudden spike in sales.
AnswerA

Collaborative filtering derives recommendations from user-item interaction patterns, so a product with no purchase history has no latent factor to learn from. This absence of interaction data is precisely the cold-start constraint, making the new catalogue item the scenario that triggers it.

Why this answer

The cold-start problem occurs when a recommendation system lacks sufficient data to make accurate predictions. In collaborative filtering, recommendations rely on historical user-item interactions (e.g., purchase history). A new product with no purchase history has no interaction data, so the system cannot find similar users or items to generate recommendations, directly causing the cold-start problem.

Exam trap

CompTIA often tests the cold-start problem by making candidates confuse it with performance issues or UI changes, but the trap here is that the cold-start problem is specifically about insufficient interaction data for new users or items, not about algorithm switches or interface redesigns.

How to eliminate wrong answers

Option B is wrong because switching from collaborative filtering to content-based filtering does not inherently cause the cold-start problem; content-based filtering uses item features (e.g., product attributes) to make recommendations, which can still work for new items if features are available. Option C is wrong because a website interface redesign affects user navigation but does not impact the underlying recommendation algorithm's data availability or the cold-start problem. Option D is wrong because a sudden spike in sales for a seasonal product provides abundant interaction data, which actually helps collaborative filtering make better recommendations, not cause a cold-start.

86
MCQeasy

A machine learning engineer wants to evaluate a binary classifier. Which metric is MOST appropriate when the positive class is rare (e.g., 1% of total data)?

A.True negative rate
B.F1-score
C.Mean squared error
D.Accuracy
AnswerB

With a 1% positive rate, accuracy is misleading because predicting all negatives scores 99%. F1-score is the harmonic mean of precision and recall, so it penalises both missed positives and false alarms, giving a meaningful measure of minority-class performance.

Why this answer

When the positive class is rare (e.g., 1% of total data), accuracy is misleading because a classifier that always predicts the negative class would achieve 99% accuracy. The F1-score is the harmonic mean of precision and recall, making it robust to class imbalance by focusing on the positive class performance. It is the most appropriate metric for evaluating binary classifiers on imbalanced datasets.

Exam trap

CompTIA often tests the misconception that accuracy is always the best metric, but in imbalanced datasets it is misleading, and candidates must recognize that F1-score (or precision-recall curves) is the correct choice for rare positive classes.

How to eliminate wrong answers

Option A is wrong because the true negative rate (specificity) measures the proportion of actual negatives correctly identified, which is not sensitive to the rare positive class and can be high even if the classifier misses all positives. Option C is wrong because mean squared error (MSE) is a regression metric that measures average squared differences between predicted and actual values, not suitable for binary classification outcomes. Option D is wrong because accuracy ( (TP+TN)/(TP+TN+FP+FN) ) is dominated by the majority class in imbalanced datasets, giving a falsely high score even when the classifier fails to detect the rare positive class.

87
Multi-Selectmedium

Which TWO techniques are commonly used to handle missing data in a dataset?

Select 2 answers
A.Feature scaling
B.One-hot encoding
C.Remove rows with missing values
D.Impute with mean or median
E.Principal component analysis (PCA)
AnswersC, D

Dropping rows containing missing values is a straightforward deletion technique that keeps remaining data untouched. It works when missingness is random and sparse, but discards information and can bias results if the missing records differ systematically from the retained ones.

Why this answer

Removing rows with missing values is a straightforward technique to handle missing data, especially when the missingness is random and the dataset is large enough that dropping a few rows does not significantly reduce the sample size or introduce bias. Option D is correct because imputing missing values with the mean or median is a common statistical method that preserves the dataset size and is simple to implement, though it can reduce variance and may distort relationships if the data is not missing completely at random.

Exam trap

CompTIA often tests the distinction between data preprocessing techniques that handle missing values versus those that transform or reduce features, so candidates may confuse feature scaling or PCA with missing data handling because they are all part of data preparation.

88
Multi-Selecthard

Which THREE factors are common causes of bias in AI systems?

Select 3 answers
A.Cross-validation
B.Lack of diversity in the development team
C.Unrepresentative training sample
D.Biased historical data used for training
E.High regularization
AnswersB, C, D

Homogeneous teams may overlook biased assumptions.

Why this answer

A lack of diversity in the development team leads to homogeneity of thought, which can cause blind spots in identifying potential biases in data, features, or model behavior. When the team does not represent the full spectrum of end users, the AI system may inadvertently encode assumptions that disadvantage underrepresented groups, resulting in biased outcomes.

Exam trap

CompTIA often tests the distinction between statistical bias (e.g., from regularization or validation techniques) and harmful societal bias that leads to unfair outcomes, so candidates mistakenly select options like cross-validation or high regularization as causes of bias.

89
Multi-Selectmedium

A company is implementing an AI solution for fraud detection. The dataset is highly imbalanced (only 1% fraudulent transactions). Which THREE techniques are most appropriate to address class imbalance? (Select three.)

Select 3 answers
A.Apply cost-sensitive learning by assigning a higher misclassification cost to the minority class.
B.Reduce the number of features using principal component analysis (PCA).
C.Use accuracy as the primary evaluation metric.
D.Evaluate model performance using precision-recall curves and F1 score.
E.Use synthetic oversampling (SMOTE) to create additional minority class samples.
AnswersA, D, E

Cost-sensitive methods penalize minority class errors more heavily.

Why this answer

Cost-sensitive learning directly addresses class imbalance by assigning a higher misclassification cost to the minority class (fraudulent transactions). This forces the model to penalize false negatives more heavily, thereby improving recall for the minority class without altering the dataset distribution.

Exam trap

CompTIA often tests the misconception that accuracy is a valid metric for imbalanced datasets, but the trap here is that candidates overlook how a high accuracy can mask poor minority class performance, leading them to select option C instead of focusing on precision-recall curves and F1 score.

90
MCQmedium

Refer to the exhibit. The training log shows loss and accuracy for a binary classification model. What is the most likely issue with this model?

A.Overfitting
B.Insufficient epochs
C.Underfitting
D.Data leakage
AnswerA

Overfitting is indicated when training loss keeps falling and training accuracy approaches 100% while validation loss rises and validation accuracy stagnates or degrades. That divergence between training and validation curves is the classic signature, matching the exhibit's logged behaviour.

Why this answer

The training log shows that training loss continues to decrease while validation loss increases after a certain point, and training accuracy approaches 100% while validation accuracy plateaus or drops. This divergence is the classic signature of overfitting, where the model memorizes noise in the training data rather than learning generalizable patterns.

Exam trap

The key trap is that candidates may only look at the final accuracy numbers without comparing training and validation curves, missing the divergence that indicates overfitting.

How to eliminate wrong answers

Option B is wrong because insufficient epochs would show both training and validation loss still decreasing at the end of training, not a divergence. Option C is wrong because underfitting would show high loss and low accuracy on both training and validation sets, not the high training accuracy seen here. Option D is wrong because data leakage typically causes unusually high performance on both sets that does not degrade, not a gap between training and validation metrics.

91
MCQmedium

A team is training a neural network for image classification. They observe that training loss decreases steadily but validation loss starts increasing after 20 epochs. What is the most likely issue?

A.Underfitting
B.Vanishing gradients
C.Data leakage
D.Overfitting
AnswerD

Overfitting occurs when the model memorises training data, so training loss keeps falling while validation loss rises after epoch 20. The divergence between decreasing training loss and increasing validation loss is the defining signature, indicating the network no longer generalises to unseen images.

Why this answer

The training loss decreasing while validation loss increases after 20 epochs is the classic signature of overfitting. The model is memorizing the training data (including noise) rather than learning generalizable patterns, causing it to perform poorly on unseen validation data.

Exam trap

The AI0-001 exam often tests the distinction between overfitting and underfitting by showing a loss curve that decreases then increases, which candidates may misinterpret as a learning rate issue or vanishing gradient problem.

How to eliminate wrong answers

Option A is wrong because underfitting would show both training and validation loss remaining high or not decreasing, not a divergence. Option B is wrong because vanishing gradients cause the network to stop learning early (loss plateaus), not a late-stage validation loss increase. Option C is wrong because data leakage typically causes both training and validation metrics to be artificially high or inconsistent from the start, not a clear divergence after many epochs.

92
Multi-Selecteasy

Which THREE are common machine learning algorithms used for regression?

Select 3 answers
A.Logistic regression
B.K-means
C.Linear regression
D.Decision tree
E.K-nearest neighbors
AnswersC, D, E

Linear regression directly models a continuous target by fitting a linear relationship between input features and output, satisfying the regression requirement. Unlike classification algorithms that predict discrete class labels, it minimises squared error to produce numeric predictions, making it a foundational supervised regression technique.

Why this answer

Linear regression (C) is the canonical regression algorithm, modeling a continuous target as a linear combination of input features by minimizing squared error. Decision tree (D) regressors recursively split data on feature thresholds to predict continuous values, e.g., via variance reduction or mean-squared-error minimization at each node. K-nearest neighbors (E) performs regression by averaging the target values of the k closest training points under a chosen distance metric, making it a standard non-parametric regressor.

Logistic regression (A) is excluded because it is a classification algorithm that outputs class probabilities via the sigmoid/softmax function, not a continuous value. K-means (B) is excluded because it is an unsupervised clustering algorithm that partitions data into groups, not a supervised regression method.

Exam trap

CompTIA often tests the distinction between regression and classification algorithms, and the trap here is that candidates mistakenly associate 'logistic regression' with regression tasks due to its name, when it is actually a classification algorithm.

93
MCQmedium

A logistics company is building a model to estimate delivery times. The team has a dataset with 120,000 labeled historical deliveries, but the labels for arrival times are noisy because some drivers manually entered them hours later. The team wants to improve label quality without discarding the dataset. Which approach best addresses the noisy-label problem?

A.Remove all records whose arrival times fall outside two standard deviations from the mean.
B.Train a model on the noisy labels, then use its predictions to relabel the most confident examples and retrain.
C.Model the label noise explicitly, for example with a noise-robust loss or a probabilistic noise model, and train with the noisy labels.
D.Apply a clustering algorithm to the delivery records and relabel each cluster with its centroid value.
AnswerC

Noise-robust losses and probabilistic noise models are designed to learn from corrupted labels by down-weighting or modeling the error process. They allow the team to retain all 120,000 records while reducing the influence of late manual entries. This directly targets the stated problem of noisy arrival-time labels without discarding data, making it the most appropriate and technically grounded solution.

Why this answer

Noise-robust training methods explicitly account for corrupted labels, allowing the team to use the full dataset while reducing the impact of late manual entries. Unlike trimming or self-relabeling, this approach models the noise process and preserves legitimate variation, which is essential for accurate delivery-time estimation.

Exam trap

The trap here is treating noisy labels as outliers to delete, when label noise can be distributed throughout the dataset and requires a modeling or robust-loss strategy rather than simple filtering.

94
MCQhard

A research team is training a deep neural network for image classification. The training loss decreases rapidly for the first few epochs but then plateaus, while validation loss starts to increase after epoch 10. Which action would best address this issue?

A.Reduce the batch size to introduce more noise during training.
B.Increase the learning rate to help the model escape the plateau.
C.Implement early stopping based on validation loss to prevent further overfitting.
D.Add more convolutional layers to increase model capacity.
AnswerC

Early stopping halts training when validation loss begins rising at epoch 10, while training loss keeps falling — the signature of overfitting. Restoring the best-performing weights prevents the network memorising training data, directly addressing the diverging loss curves described.

Why this answer

The training loss decreasing rapidly then plateauing while validation loss increases after epoch 10 is a classic sign of overfitting. Early stopping monitors validation loss and halts training when it begins to rise, preventing the model from memorizing noise in the training data. This directly addresses the overfitting issue without requiring architectural or hyperparameter changes that could destabilize training.

Exam trap

CompTIA often tests the misconception that plateauing training loss always requires adjusting learning rate or batch size, when in fact the simultaneous rise in validation loss is the definitive indicator of overfitting that early stopping is designed to solve.

How to eliminate wrong answers

Option A is wrong because reducing batch size increases gradient noise, which can actually worsen overfitting by preventing the model from converging to a stable minimum and may amplify validation loss increases. Option B is wrong because increasing the learning rate when validation loss is already rising risks overshooting the optimal weights, causing divergence or even higher validation loss. Option D is wrong because adding more convolutional layers increases model capacity, which exacerbates overfitting by giving the model more parameters to memorize training data rather than generalizing.

95
MCQhard

A financial institution uses a machine learning model to approve loan applications. The model was trained on historical data that inadvertently encoded a bias against applicants from certain zip codes, leading to discriminatory lending practices. A recent audit reveals that the model's decisions are unfair, and regulators require the bank to remediate the bias without significantly reducing overall approval accuracy. The data science team has access to the training data, the model, and a set of fairness metrics. They also have a small, unbiased validation set. Which course of action should the team take to satisfy regulatory requirements?

A.Remove the zip code feature from the model and retrain
B.Implement adversarial debiasing using the unbiased validation set to enforce fairness constraints
C.Increase the weight of samples from disadvantaged zip codes in the training data
D.Retrain the model using only the unbiased validation set
AnswerB

Adversarial debiasing directly optimizes for fairness and accuracy.

Why this answer

Adversarial debiasing directly addresses the bias encoded in the model by training a predictor and an adversary simultaneously. The adversary tries to predict the protected attribute (e.g., zip code) from the model's predictions, while the predictor is penalized for allowing such inference, enforcing fairness constraints. Using the unbiased validation set ensures the debiasing process is guided by ground truth labels that are free from historical bias, allowing the model to retain high accuracy while reducing discrimination.

Exam trap

CompTIA often tests the misconception that removing a sensitive feature (like zip code) is sufficient to eliminate bias, but the trap is that models can learn proxy features, so a more sophisticated debiasing technique like adversarial debiasing is required.

How to eliminate wrong answers

Option A is wrong because simply removing the zip code feature does not eliminate bias; the model can still learn proxy features (e.g., income, loan amount) that correlate with zip code, leading to continued discriminatory outcomes. Option C is wrong because increasing sample weights for disadvantaged zip codes may overcorrect and reduce overall accuracy, and it does not directly enforce a fairness constraint; it can also introduce new biases if the weighting is not carefully tuned. Option D is wrong because retraining on only the small unbiased validation set would likely lead to severe overfitting and poor generalization, as the dataset is too small to capture the full distribution of loan applications, significantly reducing approval accuracy.

96
MCQhard

A team is building a natural language processing (NLP) model to analyze customer feedback. They have a large corpus of unlabeled text data and want to generate word embeddings that capture semantic meaning. Which approach should they use?

A.One-hot encoding
B.TF-IDF vectorization
C.Word2Vec
D.Bag-of-words model
AnswerC

Word2Vec learns dense vector representations from unlabelled text by predicting a word from its neighbours (skip-gram) or vice versa (CBOW), so semantic relationships emerge from co-occurrence statistics. This directly satisfies the stem's requirement for embeddings from a large unlabelled corpus, unlike supervised approaches needing labelled data.

Why this answer

Word2Vec is the correct approach because it learns dense, distributed word embeddings from large unlabeled corpora by training a shallow neural network to predict words in context (CBOW) or context from words (Skip-gram). This captures semantic relationships such as analogy and similarity, which is essential for analyzing customer feedback without labeled data.

Exam trap

CompTIA often tests the distinction between frequency-based vectorization (TF-IDF, bag-of-words) and prediction-based embedding methods (Word2Vec, GloVe), trapping candidates who think TF-IDF captures semantic meaning when it only captures term importance in a document.

How to eliminate wrong answers

Option A is wrong because one-hot encoding produces sparse, high-dimensional vectors with no semantic meaning—each word is represented as a binary vector with a single 1, and all vectors are orthogonal, so no similarity or relationship between words is captured. Option B is wrong because TF-IDF vectorization relies on term frequency and inverse document frequency to produce weighted sparse vectors, which reflect word importance in a document but do not capture semantic meaning or word relationships; it is a bag-of-words variant that ignores word order and context. Option D is wrong because the bag-of-words model creates sparse vectors based on word counts, losing all word order and context, and cannot generate embeddings that capture semantic similarity or analogy.

97
MCQhard

A media company uses a reinforcement learning agent to schedule promotional banners on its homepage. The agent receives a reward when users click a banner, and it has learned to show the same sensational headline repeatedly because it historically generated high clicks. The editorial team is concerned that this harms long-term user trust. Which modification best aligns the agent's objective with long-term user satisfaction?

A.Increase the discount factor so future rewards are weighted more heavily.
B.Reduce the size of the experience replay buffer so the agent forgets older click patterns.
C.Switch from an epsilon-greedy policy to a softmax exploration policy.
D.Add a reward penalty for showing the same banner repeatedly and include a signal for user retention or satisfaction.
AnswerD

Redefining the reward to penalize repetition and include retention or satisfaction directly changes what the agent optimizes. This aligns the objective with long-term trust rather than short-term clicks. By incorporating both a repetition penalty and a satisfaction signal, the agent learns to balance engagement with sustainable user experience, which is exactly what the editorial team needs.

Why this answer

The agent's behavior stems from a reward function that values clicks without regard for long-term consequences. Adding a repetition penalty and a retention or satisfaction signal reshapes the objective so the agent learns to avoid overexposing sensational content and to prioritize sustainable user engagement.

Exam trap

The trap here is assuming that changing exploration or discount parameters will fix a misaligned objective, when the core issue is that the reward function itself rewards repetitive clickbait.

98
MCQmedium

An AI model is trained to predict loan default. The training data contains 95% non-default and 5% default. Which metric is most appropriate to evaluate model performance given the imbalanced dataset?

A.Mean squared error
B.F1-score
C.Accuracy
D.R-squared
AnswerB

With only 5% defaults, accuracy misleads by favouring the majority class. F1-score is the harmonic mean of precision and recall, so it penalises both false positives and false negatives, giving a balanced view of performance on the minority default class.

Why this answer

The F1-score is the harmonic mean of precision and recall, making it robust to class imbalance. In this dataset with 95% non-default and 5% default, accuracy would be misleadingly high (95%) even if the model never predicts default, while F1-score penalizes poor recall of the minority class.

Exam trap

CompTIA often tests the misconception that accuracy is always the best metric, leading candidates to overlook its failure in imbalanced scenarios where a trivial classifier can achieve high accuracy.

How to eliminate wrong answers

Option A is wrong because Mean Squared Error (MSE) is a regression metric that measures average squared differences between predicted and actual values, not suitable for binary classification tasks like loan default prediction. Option C is wrong because accuracy is misleading on imbalanced datasets; a model predicting all non-default would achieve 95% accuracy but fail to identify any actual defaults. Option D is wrong because R-squared is a regression metric that indicates the proportion of variance explained by the model, inappropriate for evaluating classification performance on imbalanced data.

99
MCQhard

A data scientist notices the model overfits. Which change to the exhibit's configuration would most likely reduce overfitting?

A.Remove dropout layers
B.Increase learning rate to 0.01
C.Add L2 regularization to dense layers
D.Increase units in the first dense layer to 512
AnswerC

Adding L2 regularization to dense layers penalises large weights in the loss function, reducing the network's capacity to fit training noise. This directly counteracts the overfitting observed in the exhibit's configuration without altering the architecture.

Why this answer

Adding L2 regularization to dense layers penalizes large weights by adding a squared magnitude term to the loss function, which forces the model to learn simpler patterns and reduces overfitting. This directly addresses the core issue of the model memorizing noise in the training data.

Exam trap

CompTIA often tests the misconception that increasing model capacity (more units or layers) or removing regularization always improves performance, when in fact these changes exacerbate overfitting; candidates must recognize that regularization techniques like L2 are specifically designed to penalize complexity and reduce overfitting.

How to eliminate wrong answers

Option A is wrong because removing dropout layers would actually increase overfitting, as dropout is a regularization technique that randomly drops neurons during training to prevent co-adaptation. Option B is wrong because increasing the learning rate to 0.01 (a relatively high value) can cause the optimizer to overshoot minima and lead to unstable training, but it does not directly reduce overfitting; in fact, a too-high learning rate may prevent convergence altogether. Option D is wrong because increasing units in the first dense layer to 512 adds more parameters to the model, which increases capacity and typically worsens overfitting rather than reducing it.

100
Multi-Selecthard

A natural language processing team is building a sentiment analysis model for customer reviews. They want to ensure the model generalizes well to new, unseen reviews and does not simply memorize the training data. Which TWO techniques are most appropriate to achieve this goal? (Choose two.)

Select 2 answers
A.Remove all stop words from the reviews before training.
B.Use a larger batch size during training.
C.Apply L2 regularization to the model's weights.
D.Increase the number of training epochs until training loss approaches zero.
E.Implement early stopping based on validation loss.
AnswersC, E

L2 regularization adds a penalty proportional to the square of the weights to the loss function, discouraging large weights and reducing model complexity. This helps prevent overfitting by forcing the model to learn smoother, more generalizable patterns rather than memorizing training examples. For sentiment analysis, it is a standard and effective technique to improve performance on unseen reviews.

Why this answer

L2 regularization penalizes large weights, encouraging simpler models that generalize better, while early stopping halts training when validation performance degrades, preventing overfitting. Both directly target the goal of avoiding memorization of training data. Increasing epochs, larger batch sizes, and stop word removal do not reliably improve generalization and may even hurt performance.

Exam trap

The trap here is assuming that any technique that speeds up training or reduces data size will also improve generalization, when in fact only methods that constrain model complexity or monitor validation performance directly address overfitting.

← PreviousPage 2 of 2 · 100 questions total

Ready to test yourself?

Try a timed practice session using only Ai Concepts Foundations questions.