You must diagnose model behavior from loss and accuracy curves and choose the right fix—regularization, early stopping, dropout, or more data. The single most important skill is correctly distinguishing overfitting from underfitting before selecting any remedy.
Start practicing
Machine Learning and Deep Learning — choose a session length
Free · No account required
Domain overview
This domain covers core machine learning and deep learning concepts on the AI0-001 exam: supervised and unsupervised learning, model training and evaluation, overfitting and regularization, neural network architecture, and hyperparameter tuning. Questions present realistic engineering scenarios—retail CLV prediction, image classification, dataset splitting—and ask you to diagnose problems and select appropriate remedies.
Exam objectives
Diagnosing overfitting via training versus validation loss curves and applying dropout, early stopping, and regularization
Interpreting train/validation/test accuracy gaps to identify variance, bias, or data leakage
Selecting evaluation metrics such as RMSE for regression and accuracy for classification tasks
Understanding gradient boosting, CNNs, and neural network hyperparameter tuning workflows
Confusing overfitting with underfitting: rising validation loss while training loss falls means overfitting, not a need for more epochs.
Assuming a large train-validation accuracy gap is acceptable; it signals variance and requires regularization or more data.
Treating test-set accuracy as a tuning signal; the test set must stay untouched until final evaluation.
Click any question to see the full explanation and answer options, or start a focused practice session above.
A data scientist is building a classification model to detect fraudulent transactions. The dataset is highly imbalanced with only 1% fraudulent cases. Which approach should the scientist use to evaluate model performance most effectively?
2A machine learning team is deploying a model that predicts customer churn. They notice that the model's predictions are highly sensitive to small changes in input features, leading to inconsistent outputs. Which technique should the team apply to improve model stability?
3A deep learning model for image classification is overfitting the training data. The team has already tried data augmentation and dropout. Which additional technique should they implement to reduce overfitting?
4A company wants to deploy a machine learning model that requires continuous learning as new data arrives. The model must be able to adapt to changing patterns without retraining from scratch. Which approach should be used?
5A team is building a recommendation system using collaborative filtering. They have a sparse user-item matrix. Which technique should they use to handle the sparsity and improve recommendations?
6Which TWO techniques are commonly used to handle missing data in a machine learning dataset? (Choose TWO.)
7Which THREE are common activation functions used in neural networks? (Choose THREE.)
8Which TWO are valid techniques to reduce overfitting in a deep neural network? (Choose TWO.)
9A machine learning engineer is building a spam filter. The dataset contains 10,000 emails, of which 1,000 are spam. The engineer decides to use a Random Forest classifier. Which preprocessing step is most critical to ensure the model generalizes well to new, unseen emails?
10A retail company uses a gradient boosting model to predict customer lifetime value (CLV). The model currently uses 50 features including purchase history, demographics, and web behavior. The model's RMSE on the test set is 120. The data science team wants to improve the model's accuracy without increasing training time significantly. They have access to additional data: customer support interaction logs (text), social media sentiment (text), and third-party credit scores (numeric). They also have the ability to perform feature engineering, hyperparameter tuning, and ensemble methods. Which approach is most likely to yield the best improvement in predictive performance with minimal increase in training time?
11A data scientist is training a binary classification model to detect fraudulent transactions. The dataset is highly imbalanced with only 1% fraud cases. Which technique is most appropriate to address the class imbalance?
12A machine learning engineer is tuning a neural network for image classification. The training loss decreases steadily, but the validation loss starts increasing after 50 epochs. Which action best addresses this issue?
13A company deploys a deep learning model for real-time object detection in autonomous vehicles. The model was trained on high-end GPUs but needs to run on edge devices with limited computational resources. Which technique is most effective for reducing model size and inference latency while maintaining acceptable accuracy?
14A data analyst wants to predict housing prices based on square footage, number of bedrooms, and location. Which machine learning approach is most suitable?
15A team is training a convolutional neural network (CNN) for medical image diagnosis. They have a limited dataset of 500 labeled images. Which strategy is most effective to improve model generalization?
16An AI developer observes that the training accuracy of a neural network is high, but the test accuracy is low. The model uses a ReLU activation function and Adam optimizer. Which approach is most likely to improve test accuracy?
17A machine learning engineer needs to choose an algorithm for grouping customers into segments based on purchasing behavior without any labels. Which algorithm should the engineer use?
18While training a deep neural network, the loss function fails to converge and oscillates wildly. Which adjustment is most likely to stabilize training?
19A data scientist is training a random forest model on a large dataset and notices that the model is overfitting. Which hyperparameter adjustment is most likely to reduce overfitting?
20A deep learning engineer is training a convolutional neural network for image classification. The model is overfitting the training data. Which three techniques can help reduce overfitting? (Choose three.)
21A data scientist is building a binary classification model to predict customer churn. The dataset has 10,000 samples with 80% non-churn and 20% churn. The model achieves 95% accuracy but fails to identify churners correctly. Which metric should the scientist focus on to evaluate model performance properly?
22A team is implementing a machine learning pipeline to classify images for a defect detection system. They are considering using a pre-trained convolutional neural network (CNN) and fine-tuning it on their small dataset. What is the primary advantage of transfer learning in this scenario?
23A company uses linear regression to predict sales based on advertising spend. The model's residuals show a pattern of increasing variance as spend increases. Which assumption of linear regression is violated?
24An AI engineer is training a deep neural network for image recognition. The training loss decreases steadily for the first few epochs but then plateaus and starts to oscillate. Which adjustment is most likely to improve convergence?
25A healthcare organization wants to use patient data to predict disease risk. They are concerned about bias in the model. Which step is most critical during the data preparation phase to mitigate bias?
26A deep learning model for natural language processing uses a recurrent neural network (RNN) to process long sequences. The gradients vanish after many time steps. Which architectural change is most effective to mitigate this problem?
27An organization has a dataset with categorical features having high cardinality (e.g., ZIP codes). They plan to use a tree-based model. Which encoding method is most appropriate?
28A company deploys a machine learning model that makes predictions on streaming data. Over time, the data distribution shifts, causing model performance to degrade. Which monitoring strategy is most appropriate to detect this drift?
29A data scientist is tuning hyperparameters for a support vector machine (SVM) with an RBF kernel. Which two hyperparameters most significantly affect model performance? (Select TWO.)
30A team is designing a deep learning pipeline for a computer vision task. They want to reduce overfitting. Which two techniques are specifically effective for this purpose? (Select TWO.)
31Refer to the exhibit. The training log shows losses and accuracies over 5 epochs. What is the most likely problem?
32A data scientist needs to predict whether a customer will churn based on historical data containing features like account age, monthly charges, and support tickets. The target variable is binary (churn or not). Which type of machine learning algorithm should be used?
33A team trained a deep neural network on a limited dataset. The training loss decreases consistently, but the validation loss starts increasing after 20 epochs. What is the most likely issue and the best corrective action?
34A company is building a computer vision system to detect defects in manufactured parts. They have 10,000 labeled images per class (defective and non-defective). They want to achieve high accuracy with limited computational resources. Which deep learning architecture and approach is most appropriate?
35A machine learning engineer has a dataset of 100,000 records. She splits it into 70% training, 15% validation, and 15% test sets. After training, the model achieves 95% accuracy on training and 85% on validation. What does the accuracy difference most likely indicate?
36A deep learning model for sentiment analysis uses a softmax output layer. The hidden layers currently use tanh activation. Which activation function should replace tanh to mitigate vanishing gradients in deeper networks?
37A fraud detection model is trained on a dataset where only 0.1% of transactions are fraudulent. The model achieves 99.9% accuracy but fails to catch most frauds. Which metric should the team prioritize, and which technique could help?
38A dataset contains features on vastly different scales (e.g., age 0-100 vs. income 0-1,000,000). Which preprocessing step is essential before training a neural network?
39During training of a neural network, the loss oscillates and does not converge smoothly. The learning rate is set to 0.1. What is the most likely cause and what adjustment should be made?
40A team is building a model to predict stock prices based on time series data. They need to capture long-term dependencies and avoid vanishing gradients. Which architecture is best suited?
41Which TWO are characteristics of supervised learning?
42Which THREE techniques can help reduce overfitting in neural networks?
43Which TWO are key differences between Convolutional Neural Networks (CNN) and Recurrent Neural Networks (RNN)?
44Refer to the exhibit. What is the most likely issue and what action should be taken?
45Refer to the exhibit. What is the recall of the model?
46A data scientist trains a linear regression model on housing prices. The training error is low, but test error is high. What is the most likely issue?
47A company uses a neural network for fraud detection. The dataset has 99% legitimate, 1% fraudulent. The model achieves 99% accuracy but fails to detect most frauds. Which metric should they focus on?
48A data scientist wants to reduce the dimensionality of a dataset with 200 features before training a regression model. Which technique should they use?
49An organization needs to classify customer emails into categories. They have labeled data for some categories but not all. Which approach should they use?
50A machine learning engineer notices that the gradient values in a deep network are becoming extremely small during backpropagation. What is this problem?
51A team wants to predict monthly sales using historical data. Which algorithm is most appropriate?
52A model trained on a dataset has high bias and low variance. What does this indicate?
53Which TWO are evaluation metrics for classification problems? (Choose two.)
54Based on the exhibit, what is the likely problem with the model?
55The exhibit shows a model configuration for a classification task with 10 classes. What is wrong with this setup?
56A data scientist is training a binary classification model to detect fraudulent transactions. The dataset is highly imbalanced with 99% legitimate and 1% fraudulent. Which evaluation metric should be prioritized to assess model performance?
57A team is deploying a deep learning model for real-time image classification on edge devices with limited computational resources. Which technique would best help reduce model size and inference time without significant accuracy loss?
58A machine learning engineer notices that a linear regression model has high bias. Which action is most likely to reduce bias?
59A team is developing a recommendation system for an e-commerce platform. They want to use collaborative filtering but are concerned about cold-start problems for new users. Which approach would best mitigate the cold-start problem?
60A data scientist is training a deep neural network for sentiment analysis. The training loss decreases steadily but the validation loss starts to increase after 10 epochs. What is the most likely cause and best corrective action?
61An organization wants to automate the detection of defective products on an assembly line using computer vision. They have a limited number of labeled images for defective items. Which approach would be most effective?
62A machine learning engineer is troubleshooting a recurrent neural network that fails to learn long-range dependencies in sequential data. The gradients are computed using backpropagation through time. Which phenomenon is most likely occurring, and what architectural change would best address it?
63Which TWO of the following are common activation functions used in deep neural networks?
64Which THREE of the following are techniques for handling missing data in machine learning?
65Which THREE of the following are best practices for preventing overfitting in deep learning models?
66A hospital wants to deploy a machine learning model to predict patient readmission risk within 30 days. They have a dataset with 10,000 records, 70 features including demographics, lab results, and past admissions. The target variable is binary (readmitted or not). The data scientist trains a logistic regression model and achieves an AUC of 0.85 on the test set. However, the hospital's clinicians require interpretability of predictions to trust the model. Which action should the data scientist take to ensure the model meets the interpretability requirement while maintaining performance?
67An e-commerce company uses a gradient boosting model to forecast daily sales. Recently, the model's predictions have become less accurate, showing a significant drop in R-squared on validation data. The data scientist checks for data drift but finds no significant changes in feature distributions. The model was trained on data from the past 24 months and is retrained monthly. Upon inspecting the feature importance, the data scientist notices that the top feature 'promotion_flag' has decreased in importance over time. What is the most likely cause of the performance degradation, and what should be done?
68A financial institution uses a deep learning model for fraud detection. The model is a feedforward neural network with three hidden layers. It was trained on a balanced dataset of 100,000 transactions. During deployment, the model achieves high accuracy on the test set but the fraud detection rate (true positive rate) is only 40% while the false positive rate is 0.1%. The business requires a true positive rate of at least 80%. Which of the following actions is most likely to achieve the required true positive rate while minimizing the increase in false positives?
69A data scientist is training a binary classification model to detect fraudulent transactions. The dataset contains 99.9% legitimate transactions and 0.1% fraudulent transactions. After training a logistic regression model, the accuracy is 99.9%, but the recall for the fraud class is 0%. Which of the following is the MOST likely cause?
70A machine learning engineer is preparing to train a deep neural network for image classification. To avoid overfitting, which TWO techniques should the engineer apply? (Select TWO.)
71A self-driving car company uses a reinforcement learning agent to navigate. The agent was trained in a simulated environment and achieved high rewards. When deployed in the real world, the agent fails to avoid obstacles. The team collects real-world driving data and uses it to fine-tune the model. However, fine-tuning leads to catastrophic forgetting of the simulated knowledge. Which technique should the team use to mitigate this? A. Increase the learning rate during fine-tuning. B. Use elastic weight consolidation (EWC) to regularize important weights. C. Train the model from scratch using only real-world data. D. Increase the number of layers in the network.
72A media company uses a natural language processing (NLP) model to classify news articles into topics. The model was trained on articles from 2015-2018. In 2023, the model's F1 score drops significantly. The data scientists find that the word embeddings no longer capture the meaning of some terms (e.g., 'covid', 'metaverse'). The model uses static word embeddings (Word2Vec) trained on the original corpus. Which solution BEST addresses the observed degradation? A. Replace static embeddings with contextual embeddings from a transformer model like BERT, then fine-tune the classifier. B. Retrain the static Word2Vec embeddings on a larger corpus from 2023. C. Apply data augmentation to the original training data by replacing words with synonyms. D. Increase the dimensionality of the static embeddings.
73Based on the exhibit, what is the most likely issue with the trained model?
74A financial institution is developing a fraud detection model using historical transaction data. The dataset contains over 10 million records, but only 0.01% of transactions are fraudulent. The current model uses a neural network trained with standard cross-entropy loss, and the team applies random undersampling of the majority class to create a balanced training set. However, the model still produces a high number of false positives (legitimate transactions flagged as fraud) and misses approximately 30% of actual fraud cases. The business requires that at least 95% of frauds be caught, and the false positive rate must be below 1% to avoid overwhelming fraud analysts. The team has limited resources to collect additional data and cannot change the model architecture significantly. Which approach should the team take to best meet the business requirements?
75A fraud-detection team at a bank trains a gradient-boosted tree model on two years of transaction data. Only 0.4% of transactions are fraudulent. The model achieves 99.7% accuracy but flags almost no fraud. Which approach best addresses the underlying problem with how the model is being trained and evaluated?
76A junior data scientist is training a supervised classification model to predict whether a loan applicant will default. The dataset has 40,000 labeled historical records with a clear binary outcome column. The team needs a model that outputs a probability between 0 and 1 for the default class. Which algorithm is the most appropriate choice for this task?
77A junior data scientist is asked to explain the difference between supervised and unsupervised learning to a product manager. She wants to give a single concrete example that clearly illustrates unsupervised learning. Which example should she choose?
78A machine learning engineer trains a decision tree to predict customer churn. The tree achieves 99 percent accuracy on the training set but only 68 percent on a held-out test set. The engineer wants to reduce this gap. Which single action is most likely to improve test performance?
79An ML engineer is tuning a random forest classifier for a medical diagnosis task and observes that training accuracy is 99% while validation accuracy is 78%. She wants to reduce the gap without discarding the ensemble approach. Which change is most likely to reduce the generalization gap?
80A machine learning engineer is training a deep neural network for image classification. The training loss decreases steadily, but the validation loss starts to increase after 20 epochs. The engineer wants to implement a technique that dynamically adjusts the learning rate during training to improve convergence and generalization. Which method should the engineer use?
81A junior ML engineer is asked to evaluate a binary classifier that predicts whether a bank transaction is fraudulent. The model's precision is 0.92 and recall is 0.41. The team wants to improve recall without retraining the model. Which action should the engineer take?
82A healthcare analytics team trains a model to flag patients at risk of readmission. The dataset contains 9,500 non-readmitted patients and 500 readmitted patients. The model predicts the majority class for every patient and reports 95 percent accuracy, yet it identifies no at-risk patients. Which evaluation approach best reveals the model's failure?
83A team is deploying a sentiment classifier and notices that the model outputs probabilities such as 0.83 for the positive class, but the actual positive rate among examples scored near 0.83 is only about 0.55. Stakeholders need the scores to reflect true likelihoods. Which action should the team take?
84A hospital wants an AI system to classify chest X-rays as normal or showing pneumonia. Radiologists have labeled 12,000 images, but only 900 show pneumonia. The team must choose a modeling approach that handles this class imbalance. Which approach is most appropriate?
85A data science team is preparing a gradient boosting model to predict equipment failure from sensor data. They want to tune hyperparameters that primarily control model complexity and reduce overfitting. Which two hyperparameters should they focus on? (Choose two.)
86A retail analytics team has a labeled dataset of 50,000 customer transactions where each record is tagged as either 'fraudulent' or 'legitimate.' They need a supervised learning approach that outputs a probability between 0 and 1 for the fraudulent class so it can be compared against a business threshold. Which algorithm is most appropriate for this task?
87A retail company trains a gradient-boosted tree model to forecast weekly demand for 500 stores. After six months, forecast error rises sharply even though the model code and pipeline are unchanged. Store openings, promotions, and seasonality have shifted the underlying demand patterns. Which action best addresses the root cause?
88A team trains a recurrent neural network to translate sentences averaging 60 words. During evaluation they notice that translations of the final words in long sentences are frequently wrong, while the opening words are translated accurately. Which architectural change best addresses this behavior?
89A data science team is preparing a dataset of customer support tickets to train a supervised model that routes each ticket to the correct department. They have 40,000 tickets labeled with one of eight departments. Which TWO preprocessing steps are most appropriate before training? (Choose two.)
90A machine learning engineer is evaluating a classifier on a dataset with 1,000 examples where only 30 are positive. The model predicts the negative class for almost every example. The team reports 97% accuracy and claims success. Which metric should the engineer introduce to reveal the model's poor performance on the positive class?
91A team trains a decision tree on a customer churn dataset with 40 features. The unpruned tree reaches 100% accuracy on the training set but only 68% on a held-out validation set. The team wants to reduce this gap without changing the algorithm. Which action is most appropriate?
92A team trains a model to predict whether loan applicants will default. On the holdout set the model achieves 0.86 AUC, but when audited, applicants over 60 receive systematically higher risk scores than equally qualified younger applicants. The team must reduce this disparity while preserving predictive performance. Which action should they take first?
93A financial firm trained a gradient boosting model on two years of loan data. It reported strong AUC during development, but after six months in production, approval rates for a newly launched loan product diverge sharply from expectations. The data science lead suspects the model is stale. Which approach best addresses this deployment issue?
94A data scientist is building a model to predict the likelihood of a patient having a rare disease. The dataset is highly imbalanced, with only 2% of patients having the disease. The data scientist trains a logistic regression model and achieves 98% accuracy, but the model predicts 'no disease' for all patients. Which evaluation metric should the data scientist use to better assess the model's performance?
95An AI team is preparing a support-vector machine to classify handwritten digits. Before training, they want to apply preprocessing steps that help the linear kernel separate the classes more effectively and improve generalization. Which two steps are most appropriate? (Choose two.)
Deep-dive questions
The most-searched questions in this domain — detailed explanations, worked examples, full answer breakdowns.
You must diagnose model behavior from loss and accuracy curves and choose the right fix—regularization, early stopping, dropout, or more data. The single most important skill is correctly distinguishing overfitting from underfitting before selecting any remedy.
The Courseiva AI0-001 question bank contains 95 questions in the Machine Learning and Deep Learning domain, covering the 10% of the exam attributed to this domain in the official CompTIA blueprint. Click any question to see the full explanation and answer breakdown.
Start with a 10-question focused session to identify your baseline accuracy in this domain. Read every explanation — even for questions you answer correctly — to understand the reasoning. Once you score consistently above 80%, move to a 20–30 question session to confirm depth before moving to the next domain.
Yes — the session launcher on this page draws questions exclusively from the Machine Learning and Deep Learning domain. Choose 10, 20, 30, or 50 questions for a focused session, or click individual questions to review them one by one.
Save your results, see per-domain analytics, and get readiness scores — free, for every certification.
Sign Up FreeFree forever · Every certification included