Candidates must classify AI types, frame supervised learning problems, and design leak-free dataset splits with suitable metrics. The single most important skill is diagnosing why strong test accuracy can fail in production and choosing the right evaluation practice to prevent it.
Start practicing
AI Concepts and Techniques — choose a session length
Free · No account required
Domain overview
This domain covers foundational AI terminology and core machine learning workflow concepts for the AI0-001 exam. Expect questions on defining narrow versus general AI, identifying classification and regression problems, splitting datasets correctly, and recognizing evaluation failures such as distribution shift in deployed models.
Exam objectives
Distinguishing narrow AI, general AI, and task-specific systems such as binary classification models
Applying train/validation/test splits and cross-validation to avoid data leakage in supervised learning
Selecting appropriate metrics like accuracy, precision, recall, and F1 for imbalanced binary classification
Recognizing distribution shift, overfitting, and the gap between test performance and production performance
Assuming high test accuracy guarantees production success; distribution shift and leakage often inflate offline metrics before deployment.
Confusing general AI with narrow AI, or treating any multi-task model as artificial general intelligence when it remains domain-specific.
Tuning hyperparameters on the test set or scaling before splitting, which leaks information and produces misleading evaluation results.
Click any question to see the full explanation and answer options, or start a focused practice session above.
A company is deploying a large language model for customer support. They want to reduce the number of off-topic or nonsensical responses while maintaining creativity. Which parameter adjustment would BEST achieve this?
2A startup wants to identify unusual patterns in network traffic to detect potential security breaches. They have a large dataset of normal traffic but very few labeled attacks. Which machine learning approach is MOST suitable?
3A research team is training a deep learning model for image classification using a small dataset of 1,000 labeled images. They are concerned about overfitting. Which combination of regularisation techniques would be MOST effective?
4A developer is building a natural language processing system to classify customer reviews as positive, neutral, or negative. They have 50,000 labeled reviews. Which model architecture is MOST appropriate for this task?
5A machine learning engineer is training a logistic regression model and notices that the loss is decreasing very slowly. The learning rate is set to 0.001. What is the MOST likely cause and appropriate fix?
6A team is training a generative adversarial network (GAN) to generate realistic images of furniture. The generator loss decreases sharply while the discriminator loss increases. What is the MOST likely issue and recommended action?
7A data scientist is selecting a model for a binary classification task where interpretability is critical because of regulatory requirements. The dataset has 20 features and 10,000 samples. Which model is MOST appropriate?
8A developer is using a pre-trained BERT model for a question-answering system. They want to ensure the model can handle out-of-vocabulary words. Which component of the BERT architecture is responsible for this?
9A team is training a recurrent neural network (RNN) with LSTM units to predict stock prices. The validation loss is significantly higher than the training loss. Which action is MOST likely to reduce the gap?
10A company is deploying a chatbot using a large language model. They want to mitigate the risk of prompt injection attacks. Which TWO measures should be implemented?
11A data scientist is evaluating a binary classifier for a medical diagnosis task. The dataset is imbalanced with 5% positive cases. Which THREE metrics should the data scientist consider for a comprehensive evaluation?
12A company wants to build a system that can generate new product images for an online catalog. Which TWO generative AI approaches are most suitable?
13A company wants to build a customer service chatbot that answers questions about their internal policy documents. The documents are updated monthly, and the team cannot afford to retrain a model each time. Which approach is MOST appropriate?
14A data scientist is building a model to predict whether a credit card transaction is fraudulent, using labeled historical data. Which machine learning paradigm is being used?
15A deep learning engineer is training a transformer model and notices that validation perplexity increases after a few epochs while training perplexity continues to decrease. Which of the following is the MOST likely cause?
16A product team wants a system that can generate high-quality synthetic images of furniture in different room settings for an online catalog. The images must be photorealistic and vary in style. Which generative AI approach is BEST suited for this task?
17An AI practitioner needs to measure the performance of a binary classification model for disease detection, where the cost of false negatives is very high. Which metric should be prioritized?
18An AI engineer is designing a system to detect unusual patterns in network traffic that may indicate a security breach. The system should learn from normal traffic patterns and flag deviations. Which machine learning approach is MOST appropriate?
19A machine learning engineer is training a neural network for image classification. The training loss decreases slowly and the model accuracy improves only marginally each epoch. Which hyperparameter adjustment is MOST likely to accelerate convergence?
20An AI system that can perform any intellectual task that a human being can is referred to as:
21An AI developer is selecting a model architecture for a real-time video surveillance system that must detect objects in each frame and also track movement patterns across frames. Which TWO architectures should the developer combine? (Choose 2)
22A healthcare startup is building a diagnostic support system using a large language model. The system must provide accurate, evidence-based answers and avoid generating harmful or fabricated information. Which THREE techniques should be implemented to achieve this? (Choose 3)
23A machine learning team is splitting a dataset for a binary classification problem. They want to ensure robust evaluation and avoid data leakage. Which TWO practices should they follow? (Choose 2)
24A data scientist is building a model to predict credit default using historical loan data. The dataset contains 100,000 records with 50 features, including income, debt-to-income ratio, and loan amount. The target variable is binary (default vs. no default). The goal is to maximize interpretability while maintaining high accuracy. Which algorithm is MOST appropriate?
25A team is training a deep learning model for image classification. The training loss decreases steadily but the validation loss plateaus after 20 epochs and then starts to increase. Which action is MOST likely to improve generalization?
26Which machine learning paradigm is best suited for training a model to play a game by learning from its own actions and rewards, without labeled data?
27A company is deploying a text generation model for customer service emails. They want to ensure the model's responses are factual and based on internal knowledge bases. Which technique is most effective?
28Which neural network architecture is specifically designed to handle sequential data and mitigate the vanishing gradient problem?
29A data scientist is evaluating a binary classification model. The model achieves 95% accuracy on the test set, but the precision is 0.60 and recall is 0.55. The dataset has 90% negative class samples. Which metric should the team focus on to improve the model?
30A team is fine-tuning a BERT model for a document classification task. They notice the model achieves high F1 scores on the training set but low F1 on the validation set. Which regularization technique would be MOST effective?
31In unsupervised learning, which task involves grouping similar data points together based on feature similarities?
32A developer is using a large language model via an API. They want the model to solve a math problem step by step. Which prompt engineering technique should they use?
33A company wants to deploy an LLM-based chatbot that can handle sensitive customer information. Which THREE measures should be implemented to mitigate prompt injection attacks? (Choose 3)
34A data scientist needs to predict whether a customer will churn (yes/no) based on historical data. Which type of machine learning problem is this?
35A team trains a neural network for image classification. During training, the loss decreases on the training set but increases on the validation set after a few epochs. What is the most likely cause?
36A generative AI model is asked to 'Write a poem about AI' and returns a very short, generic response. The user wants longer, more creative outputs. Which parameter adjustment is MOST likely to help?
37Which type of neural network is BEST suited for processing sequential data such as time series or natural language?
38An AI engineer is tuning a large language model for a summarization task. The output summaries are too verbose and include irrelevant details. Which technique should be applied to encourage concise outputs?
39A model trained on customer reviews achieves 98% accuracy on the test set. However, when deployed, it performs poorly on real-world data. The data scientist suspects distribution shift. Which action is MOST important to address this?
40Which of the following is a key characteristic of Narrow AI (Weak AI)?
41A team is using a pre-trained BERT model for a sentiment analysis task on product reviews. They want to adapt it to their specific domain with limited labeled data. Which approach is MOST effective?
42An organization's AI system uses a decision tree model for loan approval. The compliance team requires explanations for each decision. Which property of decision trees makes them suitable for this requirement?
43A data scientist is building a recommendation system for an e-commerce platform. The dataset includes user purchase history, product descriptions, and user demographics. The goal is to recommend products that a user is likely to purchase. Which TWO techniques are most appropriate for this task? (Select TWO.)
44A team is deploying a sentiment analysis model that must achieve high precision and high recall. They have a labeled dataset of 10,000 samples. They want to minimize overfitting. Which THREE actions are most appropriate? (Select THREE.)
45A company wants to classify images of products into categories. They have a large dataset of labeled images. Which TWO types of neural networks are most suitable for this task? (Select TWO.)
46A data scientist is building a model to predict whether a transaction is fraudulent. The dataset has 99.9% legitimate transactions and 0.1% fraudulent ones. Which evaluation metric is MOST appropriate to assess model performance given this class imbalance?
47Which machine learning paradigm involves training an agent to make decisions by interacting with an environment and receiving rewards or penalties based on its actions?
48A developer is fine-tuning a large language model for a legal document summarization task. They notice that during training, the loss decreases rapidly in the first few epochs but then plateaus with high variance. Which hyperparameter adjustment is MOST likely to help stabilize training?
49A team is deploying a sentiment analysis model for social media posts. The model currently performs well on English text but poorly on code-switched text (e.g., Spanglish). Which approach is MOST effective for improving performance on code-switched data without starting from scratch?
50Which neural network architecture is specifically designed to process sequential data, such as time series or sentences, by maintaining a hidden state that captures information about previous inputs?
51A company wants to automatically group customer support tickets into categories (e.g., billing, technical, account) without pre-labeled data. Which machine learning approach should they use?
52A generative AI model produces images from text prompts. The outputs are often blurry and lack fine details. Which model type is MOST likely being used, and which improvement would best address this issue?
53Which of the following best describes the difference between narrow AI and general AI?
54A team is training a deep learning model for image classification. They observe that training accuracy is high but validation accuracy is low, indicating overfitting. Which TWO techniques should they apply to reduce overfitting? (Select TWO)
55An AI engineer is fine-tuning a transformer-based language model for a domain-specific task. They want to improve the model's factual accuracy and reduce hallucinations. Which THREE strategies should they consider? (Select THREE)
56A data scientist needs to select a regression model to predict house prices. The dataset contains many features, some of which are irrelevant. Which TWO algorithms are BEST suited for this scenario, and why? (Select TWO)
57A company is deploying an LLM-powered application that answers questions based on internal documents. They want to minimize prompt injection attacks where users trick the model into ignoring instructions. Which THREE measures should they implement? (Select THREE)
58A company wants to use machine learning to recommend products to customers based on their purchase history. Which TWO techniques are appropriate for this task? (Select TWO)
59A data scientist is preparing a dataset for a binary classification model. The dataset has 1000 samples, with 800 positives and 200 negatives. To evaluate the model properly, which THREE steps should they take? (Select THREE)
60A hospital's AI team is building a model that estimates a patient's 10-year risk of developing heart disease from 30 clinical and lifestyle variables. A cardiologist asks the team to explain why the model produced a high-risk score for a specific patient, because clinicians are legally required to justify their recommendations. The team needs a technique that assigns a numeric contribution to each input feature for that individual prediction. Which approach should the team use?
61A retail bank is building a churn prediction model on 12 months of customer data. The data engineering team realizes that some features, such as total transactions in the last 90 days, are recorded at the moment the extraction job runs rather than at the moment each customer's churn label was determined. The model shows suspiciously high validation accuracy. Which TWO practices should the team adopt to obtain a trustworthy estimate of model performance? (Choose two.)
62A small logistics company wants to forecast next month's shipment volume using three years of historical monthly totals. The operations manager notes that volume has grown steadily and that December is always the busiest month. The data science consultant recommends a classical time series method that explicitly separates the long-term upward movement from the repeating yearly pattern. Which technique BEST fits this requirement?
63A retail analytics team wants to group customers into segments based on purchase frequency, average order value, and recency, without having any predefined segment labels. They plan to use an algorithm that partitions customers into a fixed number of groups by minimizing within-cluster variance. Which technique should they use?
64A computer vision engineer is building a model to detect defects on a manufacturing line. Defects are rare, occurring in only 0.5% of images. The engineer trains a convolutional neural network and achieves 99.5% accuracy, but the model never predicts a defect. The engineer wants to address the underlying issue. Which approach is MOST appropriate?
65A natural language processing team is building a system to classify support tickets into categories. They have a large corpus of unlabeled ticket text and a small set of manually labeled tickets. They want to leverage both to improve classification performance. Which approach is MOST suitable?
66A financial services firm is designing an AI system to detect fraudulent transactions. The dataset is highly imbalanced, with fraud representing less than 0.1% of transactions. The team wants to build a model that reliably identifies fraud while minimizing false positives that inconvenience customers. Which TWO techniques are MOST appropriate to address the class imbalance and evaluation needs? (Choose two.)
Candidates must classify AI types, frame supervised learning problems, and design leak-free dataset splits with suitable metrics. The single most important skill is diagnosing why strong test accuracy can fail in production and choosing the right evaluation practice to prevent it.
The Courseiva AI0-001 question bank contains 66 questions in the AI Concepts and Techniques domain, covering the 3% of the exam attributed to this domain in the official CompTIA blueprint. Click any question to see the full explanation and answer breakdown.
Start with a 10-question focused session to identify your baseline accuracy in this domain. Read every explanation — even for questions you answer correctly — to understand the reasoning. Once you score consistently above 80%, move to a 20–30 question session to confirm depth before moving to the next domain.
Yes — the session launcher on this page draws questions exclusively from the AI Concepts and Techniques domain. Choose 10, 20, 30, or 50 questions for a focused session, or click individual questions to review them one by one.
Save your results, see per-domain analytics, and get readiness scores — free, for every certification.
Sign Up FreeFree forever · Every certification included