AI0-001 · domain
Machine Learning and Deep Learning
This domain covers core machine learning and deep learning concepts on the AI0-001 exam: supervised and unsupervised learning, model training and evaluation, overfitting and regularization, neural network architecture, and hyperparameter tuning. Questions present realistic engineering scenarios—retail CLV prediction, image classification, dataset splitting—and ask you to diagnose problems and select appropriate remedies.
Focused practice
Practice Machine Learning and Deep Learning questions
Scored sessions drawing only from this domain — pick a length below.
Start 20-question practice test →What this domain covers
What to know about Machine Learning and Deep Learning
You must diagnose model behavior from loss and accuracy curves and choose the right fix—regularization, early stopping, dropout, or more data. The single most important skill is correctly distinguishing overfitting from underfitting before selecting any remedy.
Diagnosing overfitting via training versus validation loss curves and applying dropout, early stopping, and regularization
Interpreting train/validation/test accuracy gaps to identify variance, bias, or data leakage
Selecting evaluation metrics such as RMSE for regression and accuracy for classification tasks
Understanding gradient boosting, CNNs, and neural network hyperparameter tuning workflows
Watch out for
Common Machine Learning and Deep Learning exam traps
- ▸Confusing overfitting with underfitting: rising validation loss while training loss falls means overfitting, not a need for more epochs.
- ▸Assuming a large train-validation accuracy gap is acceptable; it signals variance and requires regularization or more data.
- ▸Treating test-set accuracy as a tuning signal; the test set must stay untouched until final evaluation.
Question index
All Machine Learning and Deep Learning questions (95)
Click any question to see the full explanation, or start a practice session above.
A team is designing a deep learning pipeline for a computer vision task. They want to reduce overfitting. Which two techniques are specifically effective for this purpose? (Select TWO.)
Medium2Refer to the exhibit. What is the recall of the model?
Easy3A machine learning engineer trains a decision tree to predict customer churn. The tree achieves 99 percent accuracy on the training set but only 68 percent on a held-out test set. The engineer wants to reduce this gap. Which single action is most likely to improve test performance?
Medium4Which TWO techniques are commonly used to handle missing data in a machine learning dataset? (Choose TWO.)
Medium5A data scientist is building a classification model to detect fraudulent transactions. The dataset is highly imbalanced with only 1% fraudulent cases. Which approach should the scientist use to evaluate model performance most effectively?
Easy6An e-commerce company uses a gradient boosting model to forecast daily sales. Recently, the model's predictions have become less accurate, showing a significant drop in R-squared on validation data. The data scientist checks for data drift but finds no significant changes in feature distributions. The model was trained on data from the past 24 months and is retrained monthly. Upon inspecting the feature importance, the data scientist notices that the top feature 'promotion_flag' has decreased in importance over time. What is the most likely cause of the performance degradation, and what should be done?
Medium7A data science team is preparing a dataset of customer support tickets to train a supervised model that routes each ticket to the correct department. They have 40,000 tickets labeled with one of eight departments. Which TWO preprocessing steps are most appropriate before training? (Choose two.)
Medium8A data scientist is building a model to predict the likelihood of a patient having a rare disease. The dataset is highly imbalanced, with only 2% of patients having the disease. The data scientist trains a logistic regression model and achieves 98% accuracy, but the model predicts 'no disease' for all patients. Which evaluation metric should the data scientist use to better assess the model's performance?
Hard9A deep learning model for sentiment analysis uses a softmax output layer. The hidden layers currently use tanh activation. Which activation function should replace tanh to mitigate vanishing gradients in deeper networks?
Medium10A data scientist is training a random forest model on a large dataset and notices that the model is overfitting. Which hyperparameter adjustment is most likely to reduce overfitting?
Hard11A company is building a computer vision system to detect defects in manufactured parts. They have 10,000 labeled images per class (defective and non-defective). They want to achieve high accuracy with limited computational resources. Which deep learning architecture and approach is most appropriate?
Hard12A team is developing a recommendation system for an e-commerce platform. They want to use collaborative filtering but are concerned about cold-start problems for new users. Which approach would best mitigate the cold-start problem?
Medium13A team trains a recurrent neural network to translate sentences averaging 60 words. During evaluation they notice that translations of the final words in long sentences are frequently wrong, while the opening words are translated accurately. Which architectural change best addresses this behavior?
Hard14A media company uses a natural language processing (NLP) model to classify news articles into topics. The model was trained on articles from 2015-2018. In 2023, the model's F1 score drops significantly. The data scientists find that the word embeddings no longer capture the meaning of some terms (e.g., 'covid', 'metaverse'). The model uses static word embeddings (Word2Vec) trained on the original corpus. Which solution BEST addresses the observed degradation? A. Replace static embeddings with contextual embeddings from a transformer model like BERT, then fine-tune the classifier. B. Retrain the static Word2Vec embeddings on a larger corpus from 2023. C. Apply data augmentation to the original training data by replacing words with synonyms. D. Increase the dimensionality of the static embeddings.
Hard15A machine learning engineer is preparing to train a deep neural network for image classification. To avoid overfitting, which TWO techniques should the engineer apply? (Select TWO.)
Easy16A machine learning engineer is troubleshooting a recurrent neural network that fails to learn long-range dependencies in sequential data. The gradients are computed using backpropagation through time. Which phenomenon is most likely occurring, and what architectural change would best address it?
Hard17A data scientist trains a linear regression model on housing prices. The training error is low, but test error is high. What is the most likely issue?
Easy18A company uses linear regression to predict sales based on advertising spend. The model's residuals show a pattern of increasing variance as spend increases. Which assumption of linear regression is violated?
Easy19An AI team is preparing a support-vector machine to classify handwritten digits. Before training, they want to apply preprocessing steps that help the linear kernel separate the classes more effectively and improve generalization. Which two steps are most appropriate? (Choose two.)
Medium20A company uses a neural network for fraud detection. The dataset has 99% legitimate, 1% fraudulent. The model achieves 99% accuracy but fails to detect most frauds. Which metric should they focus on?
Hard21While training a deep neural network, the loss function fails to converge and oscillates wildly. Which adjustment is most likely to stabilize training?
Medium22An organization has a dataset with categorical features having high cardinality (e.g., ZIP codes). They plan to use a tree-based model. Which encoding method is most appropriate?
Hard23A data scientist is training a binary classification model to detect fraudulent transactions. The dataset is highly imbalanced with 99% legitimate and 1% fraudulent. Which evaluation metric should be prioritized to assess model performance?
Easy24A team is implementing a machine learning pipeline to classify images for a defect detection system. They are considering using a pre-trained convolutional neural network (CNN) and fine-tuning it on their small dataset. What is the primary advantage of transfer learning in this scenario?
Easy25An organization wants to automate the detection of defective products on an assembly line using computer vision. They have a limited number of labeled images for defective items. Which approach would be most effective?
Medium26A hospital wants to deploy a machine learning model to predict patient readmission risk within 30 days. They have a dataset with 10,000 records, 70 features including demographics, lab results, and past admissions. The target variable is binary (readmitted or not). The data scientist trains a logistic regression model and achieves an AUC of 0.85 on the test set. However, the hospital's clinicians require interpretability of predictions to trust the model. Which action should the data scientist take to ensure the model meets the interpretability requirement while maintaining performance?
Easy27A data scientist needs to predict whether a customer will churn based on historical data containing features like account age, monthly charges, and support tickets. The target variable is binary (churn or not). Which type of machine learning algorithm should be used?
Easy28A deep learning model for natural language processing uses a recurrent neural network (RNN) to process long sequences. The gradients vanish after many time steps. Which architectural change is most effective to mitigate this problem?
Hard29A data scientist is training a deep neural network for sentiment analysis. The training loss decreases steadily but the validation loss starts to increase after 10 epochs. What is the most likely cause and best corrective action?
Medium30A team trains a decision tree on a customer churn dataset with 40 features. The unpruned tree reaches 100% accuracy on the training set but only 68% on a held-out validation set. The team wants to reduce this gap without changing the algorithm. Which action is most appropriate?
Medium31Refer to the exhibit. The training log shows losses and accuracies over 5 epochs. What is the most likely problem?
Easy32Refer to the exhibit. What is the most likely issue and what action should be taken?
Medium33A team is deploying a sentiment classifier and notices that the model outputs probabilities such as 0.83 for the positive class, but the actual positive rate among examples scored near 0.83 is only about 0.55. Stakeholders need the scores to reflect true likelihoods. Which action should the team take?
Medium34Which TWO are key differences between Convolutional Neural Networks (CNN) and Recurrent Neural Networks (RNN)?
Hard35A company deploys a deep learning model for real-time object detection in autonomous vehicles. The model was trained on high-end GPUs but needs to run on edge devices with limited computational resources. Which technique is most effective for reducing model size and inference latency while maintaining acceptable accuracy?
Hard36Which THREE are common activation functions used in neural networks? (Choose THREE.)
Medium37Which TWO are evaluation metrics for classification problems? (Choose two.)
Easy38A machine learning engineer needs to choose an algorithm for grouping customers into segments based on purchasing behavior without any labels. Which algorithm should the engineer use?
Easy39A machine learning engineer is evaluating a classifier on a dataset with 1,000 examples where only 30 are positive. The model predicts the negative class for almost every example. The team reports 97% accuracy and claims success. Which metric should the engineer introduce to reveal the model's poor performance on the positive class?
Medium40A financial institution is developing a fraud detection model using historical transaction data. The dataset contains over 10 million records, but only 0.01% of transactions are fraudulent. The current model uses a neural network trained with standard cross-entropy loss, and the team applies random undersampling of the majority class to create a balanced training set. However, the model still produces a high number of false positives (legitimate transactions flagged as fraud) and misses approximately 30% of actual fraud cases. The business requires that at least 95% of frauds be caught, and the false positive rate must be below 1% to avoid overwhelming fraud analysts. The team has limited resources to collect additional data and cannot change the model architecture significantly. Which approach should the team take to best meet the business requirements?
Hard41A team is deploying a deep learning model for real-time image classification on edge devices with limited computational resources. Which technique would best help reduce model size and inference time without significant accuracy loss?
Easy42A retail analytics team has a labeled dataset of 50,000 customer transactions where each record is tagged as either 'fraudulent' or 'legitimate.' They need a supervised learning approach that outputs a probability between 0 and 1 for the fraudulent class so it can be compared against a business threshold. Which algorithm is most appropriate for this task?
Easy43A team is training a convolutional neural network (CNN) for medical image diagnosis. They have a limited dataset of 500 labeled images. Which strategy is most effective to improve model generalization?
Medium44A junior data scientist is training a supervised classification model to predict whether a loan applicant will default. The dataset has 40,000 labeled historical records with a clear binary outcome column. The team needs a model that outputs a probability between 0 and 1 for the default class. Which algorithm is the most appropriate choice for this task?
Easy45A machine learning engineer is building a spam filter. The dataset contains 10,000 emails, of which 1,000 are spam. The engineer decides to use a Random Forest classifier. Which preprocessing step is most critical to ensure the model generalizes well to new, unseen emails?
Medium46Which TWO are valid techniques to reduce overfitting in a deep neural network? (Choose TWO.)
Hard47A team is building a recommendation system using collaborative filtering. They have a sparse user-item matrix. Which technique should they use to handle the sparsity and improve recommendations?
Easy48Which TWO of the following are common activation functions used in deep neural networks?
Easy49A machine learning engineer has a dataset of 100,000 records. She splits it into 70% training, 15% validation, and 15% test sets. After training, the model achieves 95% accuracy on training and 85% on validation. What does the accuracy difference most likely indicate?
Easy50A data analyst wants to predict housing prices based on square footage, number of bedrooms, and location. Which machine learning approach is most suitable?
Easy51A deep learning model for image classification is overfitting the training data. The team has already tried data augmentation and dropout. Which additional technique should they implement to reduce overfitting?
Hard52A team trains a model to predict whether loan applicants will default. On the holdout set the model achieves 0.86 AUC, but when audited, applicants over 60 receive systematically higher risk scores than equally qualified younger applicants. The team must reduce this disparity while preserving predictive performance. Which action should they take first?
Medium53A company wants to deploy a machine learning model that requires continuous learning as new data arrives. The model must be able to adapt to changing patterns without retraining from scratch. Which approach should be used?
Easy54A healthcare organization wants to use patient data to predict disease risk. They are concerned about bias in the model. Which step is most critical during the data preparation phase to mitigate bias?
Medium55A data science team is preparing a gradient boosting model to predict equipment failure from sensor data. They want to tune hyperparameters that primarily control model complexity and reduce overfitting. Which two hyperparameters should they focus on? (Choose two.)
Hard56A machine learning engineer is tuning a neural network for image classification. The training loss decreases steadily, but the validation loss starts increasing after 50 epochs. Which action best addresses this issue?
Medium57Which TWO are characteristics of supervised learning?
Easy58The exhibit shows a model configuration for a classification task with 10 classes. What is wrong with this setup?
Hard59An AI engineer is training a deep neural network for image recognition. The training loss decreases steadily for the first few epochs but then plateaus and starts to oscillate. Which adjustment is most likely to improve convergence?
Medium60An AI developer observes that the training accuracy of a neural network is high, but the test accuracy is low. The model uses a ReLU activation function and Adam optimizer. Which approach is most likely to improve test accuracy?
Hard61Which THREE techniques can help reduce overfitting in neural networks?
Medium62A healthcare analytics team trains a model to flag patients at risk of readmission. The dataset contains 9,500 non-readmitted patients and 500 readmitted patients. The model predicts the majority class for every patient and reports 95 percent accuracy, yet it identifies no at-risk patients. Which evaluation approach best reveals the model's failure?
Medium63A fraud-detection team at a bank trains a gradient-boosted tree model on two years of transaction data. Only 0.4% of transactions are fraudulent. The model achieves 99.7% accuracy but flags almost no fraud. Which approach best addresses the underlying problem with how the model is being trained and evaluated?
Medium64During training of a neural network, the loss oscillates and does not converge smoothly. The learning rate is set to 0.1. What is the most likely cause and what adjustment should be made?
Medium65A fraud detection model is trained on a dataset where only 0.1% of transactions are fraudulent. The model achieves 99.9% accuracy but fails to catch most frauds. Which metric should the team prioritize, and which technique could help?
Hard66Which THREE of the following are techniques for handling missing data in machine learning?
Medium67A company deploys a machine learning model that makes predictions on streaming data. Over time, the data distribution shifts, causing model performance to degrade. Which monitoring strategy is most appropriate to detect this drift?
Hard68A data scientist is building a binary classification model to predict customer churn. The dataset has 10,000 samples with 80% non-churn and 20% churn. The model achieves 95% accuracy but fails to identify churners correctly. Which metric should the scientist focus on to evaluate model performance properly?
Easy69A retail company trains a gradient-boosted tree model to forecast weekly demand for 500 stores. After six months, forecast error rises sharply even though the model code and pipeline are unchanged. Store openings, promotions, and seasonality have shifted the underlying demand patterns. Which action best addresses the root cause?
Hard70A hospital wants an AI system to classify chest X-rays as normal or showing pneumonia. Radiologists have labeled 12,000 images, but only 900 show pneumonia. The team must choose a modeling approach that handles this class imbalance. Which approach is most appropriate?
Medium71An organization needs to classify customer emails into categories. They have labeled data for some categories but not all. Which approach should they use?
Medium72A junior ML engineer is asked to evaluate a binary classifier that predicts whether a bank transaction is fraudulent. The model's precision is 0.92 and recall is 0.41. The team wants to improve recall without retraining the model. Which action should the engineer take?
Easy73A machine learning team is deploying a model that predicts customer churn. They notice that the model's predictions are highly sensitive to small changes in input features, leading to inconsistent outputs. Which technique should the team apply to improve model stability?
Medium74A data scientist is training a binary classification model to detect fraudulent transactions. The dataset is highly imbalanced with only 1% fraud cases. Which technique is most appropriate to address the class imbalance?
Easy75Based on the exhibit, what is the most likely issue with the trained model?
Medium76A financial institution uses a deep learning model for fraud detection. The model is a feedforward neural network with three hidden layers. It was trained on a balanced dataset of 100,000 transactions. During deployment, the model achieves high accuracy on the test set but the fraud detection rate (true positive rate) is only 40% while the false positive rate is 0.1%. The business requires a true positive rate of at least 80%. Which of the following actions is most likely to achieve the required true positive rate while minimizing the increase in false positives?
Hard77Based on the exhibit, what is the likely problem with the model?
Medium78A self-driving car company uses a reinforcement learning agent to navigate. The agent was trained in a simulated environment and achieved high rewards. When deployed in the real world, the agent fails to avoid obstacles. The team collects real-world driving data and uses it to fine-tune the model. However, fine-tuning leads to catastrophic forgetting of the simulated knowledge. Which technique should the team use to mitigate this? A. Increase the learning rate during fine-tuning. B. Use elastic weight consolidation (EWC) to regularize important weights. C. Train the model from scratch using only real-world data. D. Increase the number of layers in the network.
Medium79A machine learning engineer notices that the gradient values in a deep network are becoming extremely small during backpropagation. What is this problem?
Hard80A data scientist wants to reduce the dimensionality of a dataset with 200 features before training a regression model. Which technique should they use?
Easy81A machine learning engineer notices that a linear regression model has high bias. Which action is most likely to reduce bias?
Easy82A data scientist is training a binary classification model to detect fraudulent transactions. The dataset contains 99.9% legitimate transactions and 0.1% fraudulent transactions. After training a logistic regression model, the accuracy is 99.9%, but the recall for the fraud class is 0%. Which of the following is the MOST likely cause?
Medium83A model trained on a dataset has high bias and low variance. What does this indicate?
Medium84A dataset contains features on vastly different scales (e.g., age 0-100 vs. income 0-1,000,000). Which preprocessing step is essential before training a neural network?
Easy85A team trained a deep neural network on a limited dataset. The training loss decreases consistently, but the validation loss starts increasing after 20 epochs. What is the most likely issue and the best corrective action?
Medium86An ML engineer is tuning a random forest classifier for a medical diagnosis task and observes that training accuracy is 99% while validation accuracy is 78%. She wants to reduce the gap without discarding the ensemble approach. Which change is most likely to reduce the generalization gap?
Hard87A deep learning engineer is training a convolutional neural network for image classification. The model is overfitting the training data. Which three techniques can help reduce overfitting? (Choose three.)
Medium88A junior data scientist is asked to explain the difference between supervised and unsupervised learning to a product manager. She wants to give a single concrete example that clearly illustrates unsupervised learning. Which example should she choose?
Easy89A financial firm trained a gradient boosting model on two years of loan data. It reported strong AUC during development, but after six months in production, approval rates for a newly launched loan product diverge sharply from expectations. The data science lead suspects the model is stale. Which approach best addresses this deployment issue?
Hard90A team is building a model to predict stock prices based on time series data. They need to capture long-term dependencies and avoid vanishing gradients. Which architecture is best suited?
Hard91A team wants to predict monthly sales using historical data. Which algorithm is most appropriate?
Easy92Which THREE of the following are best practices for preventing overfitting in deep learning models?
Hard93A data scientist is tuning hyperparameters for a support vector machine (SVM) with an RBF kernel. Which two hyperparameters most significantly affect model performance? (Select TWO.)
Easy94A retail company uses a gradient boosting model to predict customer lifetime value (CLV). The model currently uses 50 features including purchase history, demographics, and web behavior. The model's RMSE on the test set is 120. The data science team wants to improve the model's accuracy without increasing training time significantly. They have access to additional data: customer support interaction logs (text), social media sentiment (text), and third-party credit scores (numeric). They also have the ability to perform feature engineering, hyperparameter tuning, and ensemble methods. Which approach is most likely to yield the best improvement in predictive performance with minimal increase in training time?
Medium95A machine learning engineer is training a deep neural network for image classification. The training loss decreases steadily, but the validation loss starts to increase after 20 epochs. The engineer wants to implement a technique that dynamically adjusts the learning rate during training to improve convergence and generalization. Which method should the engineer use?
HardOther domains
All AI0-001 exam domains
Frequently asked questions
- What does the Machine Learning and Deep Learning domain cover on the AI0-001 exam?
- You must diagnose model behavior from loss and accuracy curves and choose the right fix—regularization, early stopping, dropout, or more data. The single most important skill is correctly distinguishing overfitting from underfitting before selecting any remedy.
- How many questions are in this domain?
- This page lists all 95 Machine Learning and Deep Learning questions in the AI0-001 question bank. The actual exam draws from this domain proportionally to its weighting in the official exam blueprint.
- What is the best way to practise this domain?
- Start with a short focused session (10 questions) to identify gaps, then work through explanations. Repeat with a longer session once the weak areas feel solid.
- Can I practise only Machine Learning and Deep Learning questions?
- Yes — the session launcher on this page filters questions to this domain only. Choose any session length for inline explanations and scoring.