AI0-001 · domain
AI Concepts and Foundations
Domain 1 covers foundational AI concepts: machine learning paradigms, neural network training dynamics, data handling, and model evaluation. Questions are scenario-based, asking you to diagnose training problems, select adaptation strategies for limited data, and choose architectures or techniques that fit compute and accuracy constraints. Expect applied reasoning over definitions.
Focused practice
Practice AI Concepts and Foundations questions
Scored sessions drawing only from this domain — pick a length below.
Start 20-question practice test →What this domain covers
What to know about AI Concepts and Foundations
You must diagnose training and deployment scenarios and pick the right technique: tune hyperparameters to fix convergence, use transfer learning or fine-tuning for small datasets, and choose metrics that reflect class imbalance. Getting the data-and-metric fit right matters most.
Supervised, unsupervised, and reinforcement learning paradigm selection for a given business problem
Effect of hyperparameters like learning rate and batch size on training convergence and stability
Transfer learning, fine-tuning, and feature extraction when labeled data is scarce
Evaluation metrics such as precision, recall, F1, and their use in imbalanced recommendation tasks
Watch out for
Common AI Concepts and Foundations exam traps
- ▸Assuming a lower learning rate always improves results; too small a rate slows convergence rather than fixing slow training loss.
- ▸Confusing fine-tuning with training from scratch; with only hundreds of labeled examples, full retraining overfits and wastes compute.
- ▸Optimizing overall accuracy on imbalanced data, which hides poor recall on niche or minority classes the business cares about.
Question index
All AI Concepts and Foundations questions (100)
Click any question to see the full explanation, or start a practice session above.
A data scientist is training a binary classification model to detect fraudulent transactions. The dataset has 99% legitimate transactions and 1% fraudulent. The model achieves 99% accuracy but fails to catch most fraud. Which metric should the team prioritize to evaluate model performance?
Easy2When evaluating a binary classification model, which two metrics are most appropriate for imbalanced datasets? (Choose two.)
Medium3A self-driving car company is testing a perception model that detects pedestrians. The model achieves 99% accuracy on the test set but fails to detect pedestrians wearing dark clothing at night. The company wants to improve the model's robustness. Which action should the team take to best address this specific weakness?
Hard4A marketing team wants to segment customers into groups based on purchasing behavior without predefined categories. Which algorithm should they use?
Easy5An AI team is deploying a predictive maintenance model for industrial equipment. The model predicts failure within a 30-day window. The cost of a false positive is 10% of the cost of a false negative. Which evaluation metric should the team prioritize?
Hard6A retail company deploys a machine learning model to predict customer churn. The model outputs a probability between 0 and 1, and churn is predicted if probability > 0.5. After deployment, the model has a high false positive rate (many non-churning customers labeled as churn), which leads to unnecessary retention offers and increased costs. The data science team confirms the model was trained on historical data with a balanced class distribution. The business team wants to reduce false positives while maintaining a reasonable true positive rate. However, they cannot retrain the model because the original training data is no longer available. What is the best course of action to reduce false positives?
Hard7A company deploys a chatbot using a large language model (LLM). After launch, users report that the chatbot sometimes generates plausible but false information. This phenomenon is known as:
Medium8A data analyst needs to select two appropriate unsupervised learning techniques for clustering unlabeled data. (Choose two.)
Easy9An AI governance committee is reviewing a resume-screening model. The model's accuracy is high overall, but its false negative rate is much higher for applicants from one demographic group than for others. The committee wants to address this disparity. Which action best targets the problem?
Hard10A company uses an AI model to screen job applications. The model is trained on historical hiring data that reflects past biases. After deployment, the model disproportionately rejects candidates from certain demographics. Which concept does this best illustrate?
Medium11A self-driving car uses an AI model that learns by trial and error, receiving rewards for correct actions and penalties for mistakes. This type of learning is:
Medium12Which TWO of the following are appropriate uses of unsupervised learning?
Medium13A data scientist trains a deep neural network for image classification. The training loss decreases but validation loss starts increasing after 50 epochs. What should the data scientist do to improve generalization?
Hard14Which TWO of the following are techniques used for reducing overfitting in neural networks? (Choose two.)
Hard15An organization is developing an AI system to approve loan applications. They want to ensure the model does not discriminate based on race or gender. Which technique BEST addresses this concern?
Hard16A cybersecurity firm is developing an AI system to detect zero-day malware using behavior analysis. The team collects a dataset of 1,000 malware samples and 10,000 benign files from corporate endpoints. The model is a random forest classifier. After deployment, the false positive rate is 5%, which is acceptable, but the detection rate for new malware variants drops to 30%. The security analyst suspects the model is overfitting to the specific malware families in the training set. Which improvement should the team implement first?
Hard17An e-commerce company deploys a recommendation system using collaborative filtering. After launch, the system shows high accuracy for popular items but fails to recommend niche products to users who would likely buy them. Which technique should the team implement to improve recommendations for long-tail items?
Hard18A data scientist is training a neural network to classify images of animals. The training accuracy is 99%, but validation accuracy is only 65%. Which technique should the data scientist use to address this issue?
Medium19A startup is building a chatbot to handle customer inquiries. They want the chatbot to understand context and provide accurate responses without requiring extensive labeled data. Which AI approach is most suitable?
Medium20An organization wants to classify support tickets into categories (billing, technical, etc.). Which type of machine learning is most suitable?
Easy21A hospital wants to deploy an AI system that analyzes chest X-rays to detect pneumonia. The radiology team insists that the system provide a confidence score alongside each diagnosis so they can decide whether to trust the output. Which AI concept are they primarily concerned with?
Medium22A hospital uses an AI system to prioritize patient triage based on vital signs and medical history. During a trial, the system consistently assigns lower urgency to elderly patients with chronic conditions, even when their symptoms suggest high risk. Which approach best addresses this bias?
Medium23A team is designing an AI system for autonomous driving. They need to decide between an end-to-end deep learning approach versus a modular pipeline (perception, planning, control). Which is a key advantage of the modular approach?
Hard24A data scientist is preparing a dataset for supervised learning. Which TWO steps are essential?
Easy25Refer to the exhibit. An AI auditor reviews the fairness configuration. What is the purpose of this policy?
Medium26An AI system is being developed to diagnose diseases from medical images. The model achieves 99% accuracy on the test set, but when deployed in a different hospital, performance drops significantly. Which of the following is the MOST likely cause?
Hard27An AI system for autonomous vehicles uses reinforcement learning (RL) to navigate. The reward function encourages reaching the destination quickly but penalizes collisions heavily. The agent learns to drive aggressively, causing minor accidents. Which modification to the reward function would best align the agent's behavior with desired safe driving?
Hard28A team is deploying an AI model for credit approval. Which TWO ethical considerations must be addressed?
Medium29A retail company wants to build a model to predict customer churn based on purchase history and demographics. The dataset includes categorical features like region and gender, and numerical features like total spend. What is the best initial step before training the model?
Easy30A startup is designing an AI assistant that must handle a wide range of requests, including summarizing text, answering questions, and translating languages. The team is selecting a foundation model approach. Which two characteristics are typical of foundation models? (Choose two.)
Medium31A company develops an AI model that recommends job candidates. The model inadvertently discriminates against a protected group. Which approach is most effective for mitigating this bias?
Hard32A data scientist is preparing a dataset for a classification task. The dataset contains 10,000 rows and 50 features, but many features have missing values. Which approach should the scientist take first to address the missing data?
Easy33A hospital deploys an AI system to detect pneumonia from chest X-rays. The model achieves 95% accuracy on the test set but later is found to be less accurate for patients under 18. The development team suspects bias. Which step should be taken first to investigate?
Medium34A research team is developing an AI system to predict patient outcomes from electronic health records. The team must ensure the system adheres to ethical AI principles. Which TWO practices best align with the principle of transparency and explainability? (Choose two.)
Medium35Which TWO statements correctly describe the difference between supervised and unsupervised learning?
Medium36An AI model for detecting fraudulent transactions has high precision but low recall. Which business impact is most likely?
Medium37A small e-commerce company wants to implement a chatbot to handle customer inquiries about order status and returns. The company has limited historical chat data and wants a solution that can be deployed quickly without extensive training. Which type of AI solution is most appropriate?
Easy38A support team wants an AI assistant that can answer employee questions by retrieving passages from the company's internal policy documents and generating a response grounded in those passages. The documents change weekly. Which approach should the team implement?
Medium39A manufacturing company uses a computer vision AI to inspect products on an assembly line for defects. The AI model was trained on images from a single camera angle under bright, uniform lighting. Recently, the company moved the inspection station to a different part of the factory where lighting is dimmer and varies due to nearby windows. The model now misclassifies many non-defective products as defective, causing false alarms and production delays. The team has limited labeled data from the new environment. Which action should the team take to restore inspection accuracy while minimizing downtime?
Medium40An AI model achieves high accuracy on training data but performs poorly on new test data. The data scientist suspects the model has memorized noise. Which technique directly adds a penalty term to the loss function to address this?
Hard41A financial institution uses a regression model to predict credit risk. The model has a high R-squared on training data but low R-squared on test data. Which of the following is the most likely cause?
Medium42A hospital's AI governance committee is reviewing a sepsis-prediction model before deployment. The model was trained on five years of historical ICU data in which patients who received early antibiotics had better outcomes, and the model learned to recommend antibiotics for nearly every patient with any fever. The committee wants to determine whether the model has learned a spurious correlation rather than a true clinical signal. Which action best evaluates this concern?
Medium43A hospital wants an AI system to review chest X-ray images and flag those that may show pneumonia, but a radiologist will make the final diagnosis. The IT team must classify this system for documentation. Which category of AI best describes this deployment?
Easy44A company uses a pre-trained language model for a legal document classification task. They have limited labeled data (500 documents). Which strategy is MOST effective for adapting the model to this domain?
Medium45A data scientist trains a linear regression model to predict house prices. The model has high bias and low variance. Which action would most likely reduce bias?
Medium46A data science team is preparing a dataset of loan applications. Each row contains income, credit score, employment length, and a loan amount. Before training a model, the team wants to reduce the influence of income, which is measured in dollars and ranges into the hundreds of thousands, compared with credit score, which ranges from 300 to 850. Which technique should the team apply?
Medium47An organization is deploying a deep learning model in production. Which THREE components are essential for maintaining model performance over time?
Hard48A financial services firm is deploying a credit-scoring model that uses alternative data such as utility payments and rental history. The compliance team is concerned about fairness and transparency. Which TWO practices best support responsible AI deployment in this scenario? (Choose two.)
Hard49A healthcare provider wants to use AI to predict patient readmission risk. They have structured data (age, diagnosis, lab results) and unstructured clinical notes. Which approach is most appropriate?
Easy50Refer to the exhibit. The data scientist notices that the model achieves 98% accuracy on the training set but only 72% on the test set. Which change to the model parameters is most likely to reduce this gap?
Easy51Refer to the exhibit. A data scientist observes the training output. Which issue is most likely?
Medium52Based on the exhibit, what is the most likely issue with the model training?
Medium53A machine learning team notices that their model's performance degrades when deployed to a new geographic region. The data distribution in the new region differs from the training data. Which concept best describes this issue?
Medium54A startup is building a chatbot for customer service. They have 500 recorded conversations and want to use a pre-trained language model to generate responses. However, they have limited computational resources and need the chatbot to respond in real-time. They are considering fine-tuning a large model like GPT-3 or using a smaller model like DistilBERT. The conversation data contains industry-specific jargon. Which approach should they take?
Easy55A financial services firm is deploying a credit-scoring model that uses alternative data such as utility payments and rental history. The model shows high accuracy but the firm is concerned about regulatory compliance and explainability. The firm must provide adverse action notices to applicants who are denied credit. Which approach best satisfies the need for explainability while maintaining model performance?
Hard56An AI engineer is tuning a deep learning model and observes that the training loss decreases very slowly. The learning rate is set to 0.001. Which adjustment is most likely to speed up convergence?
Medium57A software company wants to add a feature that automatically transcribes customer support phone calls into text for analysis. Which type of AI technology is best suited for this task?
Easy58Which TWO of the following are common activation functions used in neural networks? (Choose two.)
Easy59A company wants to use AI to automatically categorize customer support tickets into topics like 'billing', 'technical', 'account'. They have 10,000 labeled examples. Which algorithm is most suitable for this task?
Easy60Which THREE of the following are types of machine learning paradigms? (Choose three.)
Medium61An AI model is being developed for medical diagnosis from X-ray images. The dataset contains only frontal chest X-rays. The model achieves high accuracy on test set but fails on lateral views. What is the most likely cause?
Medium62In the AI lifecycle, which phase involves splitting data into training, validation, and test sets?
Easy63A government agency is deploying an AI model to screen loan applications. The model uses features like income, credit score, employment history, and zip code. During fairness auditing, the model is found to deny a disproportionately high number of applicants from a particular demographic group, even when controlling for legitimate financial factors. The agency wants to mitigate this bias without significantly reducing overall accuracy. Which approach should the data scientist prioritize?
Medium64A small e-commerce startup has only 800 labeled customer-support tickets and needs to classify new tickets into categories such as billing, shipping, and returns. The team has no budget for large-scale annotation and wants to leverage a model already trained on millions of general text documents. Which approach best fits this constraint?
Easy65A company built a speech-to-text model using a recurrent neural network (RNN). During deployment, the model performs poorly on accented speech. Which action would most effectively improve model robustness?
Medium66A company wants to create an AI system that can identify objects in images. They have a large dataset of labeled images. Which type of neural network architecture is most suitable?
Medium67A team is deploying a deep learning model that uses a convolutional neural network (CNN) for image recognition. The model achieves high accuracy but is very slow to infer on edge devices. Which THREE optimization techniques should the team consider to speed up inference without significant accuracy loss? (Select three.)
Hard68Which TWO of the following are key characteristics of unsupervised learning?
Hard69A financial analyst is using a linear regression model to predict housing prices based on square footage. The model's predictions are consistently off by a large margin for both very small and very large houses, while performing well for average-sized houses. Which phenomenon is most likely occurring?
Medium70A company implements a chatbot using a rule-based system. Users complain the chatbot cannot handle new queries. Which AI approach should be considered to improve flexibility?
Easy71A data scientist wants to group customers into segments based on purchasing behavior without predefined labels. Which type of machine learning is most appropriate?
Easy72A company deploys an AI model to predict equipment failure. The model performs well on historical data but fails to generalize to new data from a different factory. Which concept best describes this issue?
Easy73Refer to the exhibit. A team deploys a sentiment analysis model with this policy. After one month, the monitoring system triggers an alert for feature drift. Which action should the team take first?
Hard74A hospital's AI governance committee is reviewing a diagnostic model that performs well on the general population but poorly on a rare disease subgroup. The committee wants to determine whether the model's poor performance on this subgroup is due to a data problem or a model problem. Which action should the committee take FIRST to make this determination?
Medium75A chatbot developer uses a transformer-based model for customer service. Users complain that the chatbot sometimes gives offensive responses. Which technique should be applied first to mitigate this issue?
Easy76An AI team notices that their model's performance degrades over time because the statistical relationship between input features and the target variable changes. This issue is called:
Hard77Which TWO of the following are common techniques to reduce overfitting in a neural network?
Easy78Which THREE of the following are key considerations when deploying an AI model in a production environment?
Hard79A hospital uses an AI system to predict patient deterioration from vital signs. The system currently uses a logistic regression model trained on data from the past year. Recently, the hospital adopted a new patient monitoring device that provides more accurate readings. The model's performance has dropped significantly. The data science team has access to the new device's data for the past month and wants to improve the model with minimal disruption. The team also wants to ensure the model remains interpretable for regulatory compliance. Which approach should they take?
Medium80A team deploying an AI model for real-time fraud detection notices that inference latency is too high. The model is a deep neural network with 50 layers, deployed on a cloud GPU. Which of the following is the BEST approach to reduce latency while maintaining acceptable accuracy?
Medium81A self-driving car company is developing an object detection system using a convolutional neural network (CNN). The system needs to detect pedestrians and vehicles in real-time with high accuracy. Which technique can reduce inference time while maintaining accuracy?
Hard82A data scientist splits a dataset into training (80%) and test (20%). After training, the model achieves 95% accuracy on training and 60% on test. Which step should the data scientist take first?
Hard83A team is training a deep learning model for natural language processing using a large corpus. They notice the model has a very high number of parameters and training is slow. Which technique can reduce the number of parameters without significant performance loss?
Hard84A company wants to use AI to analyze customer reviews and determine sentiment (positive, negative, neutral). Which AI subfield is most directly applicable?
Easy85A marketing team uses a recommendation system to suggest products to customers. The system currently uses collaborative filtering. Which scenario would most likely cause the cold-start problem?
Easy86A machine learning engineer wants to evaluate a binary classifier. Which metric is MOST appropriate when the positive class is rare (e.g., 1% of total data)?
Easy87Which TWO techniques are commonly used to handle missing data in a dataset?
Medium88Which THREE factors are common causes of bias in AI systems?
Hard89A company is implementing an AI solution for fraud detection. The dataset is highly imbalanced (only 1% fraudulent transactions). Which THREE techniques are most appropriate to address class imbalance? (Select three.)
Medium90Refer to the exhibit. The training log shows loss and accuracy for a binary classification model. What is the most likely issue with this model?
Medium91A team is training a neural network for image classification. They observe that training loss decreases steadily but validation loss starts increasing after 20 epochs. What is the most likely issue?
Medium92Which THREE are common machine learning algorithms used for regression?
Easy93A logistics company is building a model to estimate delivery times. The team has a dataset with 120,000 labeled historical deliveries, but the labels for arrival times are noisy because some drivers manually entered them hours later. The team wants to improve label quality without discarding the dataset. Which approach best addresses the noisy-label problem?
Medium94A research team is training a deep neural network for image classification. The training loss decreases rapidly for the first few epochs but then plateaus, while validation loss starts to increase after epoch 10. Which action would best address this issue?
Hard95A financial institution uses a machine learning model to approve loan applications. The model was trained on historical data that inadvertently encoded a bias against applicants from certain zip codes, leading to discriminatory lending practices. A recent audit reveals that the model's decisions are unfair, and regulators require the bank to remediate the bias without significantly reducing overall approval accuracy. The data science team has access to the training data, the model, and a set of fairness metrics. They also have a small, unbiased validation set. Which course of action should the team take to satisfy regulatory requirements?
Hard96A team is building a natural language processing (NLP) model to analyze customer feedback. They have a large corpus of unlabeled text data and want to generate word embeddings that capture semantic meaning. Which approach should they use?
Hard97A media company uses a reinforcement learning agent to schedule promotional banners on its homepage. The agent receives a reward when users click a banner, and it has learned to show the same sensational headline repeatedly because it historically generated high clicks. The editorial team is concerned that this harms long-term user trust. Which modification best aligns the agent's objective with long-term user satisfaction?
Hard98An AI model is trained to predict loan default. The training data contains 95% non-default and 5% default. Which metric is most appropriate to evaluate model performance given the imbalanced dataset?
Medium99A data scientist notices the model overfits. Which change to the exhibit's configuration would most likely reduce overfitting?
Hard100A natural language processing team is building a sentiment analysis model for customer reviews. They want to ensure the model generalizes well to new, unseen reviews and does not simply memorize the training data. Which TWO techniques are most appropriate to achieve this goal? (Choose two.)
HardOther domains
All AI0-001 exam domains
Frequently asked questions
- What does the AI Concepts and Foundations domain cover on the AI0-001 exam?
- You must diagnose training and deployment scenarios and pick the right technique: tune hyperparameters to fix convergence, use transfer learning or fine-tuning for small datasets, and choose metrics that reflect class imbalance. Getting the data-and-metric fit right matters most.
- How many questions are in this domain?
- This page lists all 100 AI Concepts and Foundations questions in the AI0-001 question bank. The actual exam draws from this domain proportionally to its weighting in the official exam blueprint.
- What is the best way to practise this domain?
- Start with a short focused session (10 questions) to identify gaps, then work through explanations. Repeat with a longer session once the weak areas feel solid.
- Can I practise only AI Concepts and Foundations questions?
- Yes — the session launcher on this page filters questions to this domain only. Choose any session length for inline explanations and scoring.