Be able to diagnose a dataset or metric and choose the right fix: median imputation for skewed data, threshold tuning to cut false negatives, and consistent scaling. The single most important thing is matching the remedy to the stated cost or distribution, not applying a default.
Start practicing
AI Models and Data Engineering — choose a session length
Free · No account required
Domain overview
This domain covers the data lifecycle behind AI systems: sourcing, cleaning, transforming, and validating data, plus selecting and evaluating models. Questions are scenario-based, asking you to pick an imputation strategy, fix class imbalance, tune a decision threshold, or interpret precision, recall, and regression metrics for a stated business cost.
Exam objectives
Choosing imputation (mean, median, mode, or model-based) based on distribution and missingness
Adjusting classification decision thresholds to trade precision against recall for business cost
Feature scaling and normalization for regression inputs with widely differing numeric ranges
Splitting data into train, validation, and test sets and detecting leakage or imbalance
Using mean imputation on skewed or non-normal data instead of median, which distorts the distribution
Lowering the threshold when false negatives are costly; you must lower it, not raise it, to catch more positives
Scaling the test set with its own statistics rather than reusing parameters fit on the training data
Click any question to see the full explanation and answer options, or start a focused practice session above.
An engineer is building a regression model to predict housing prices. The dataset includes features such as square footage, number of bedrooms, and year built. The engineer notices that the square footage values range from 500 to 10,000, while the number of bedrooms ranges from 1 to 5. Which preprocessing step is most critical before training a gradient descent-based model?
2A machine learning team is deploying a sentiment analysis model for customer reviews. The model was trained on reviews from an e-commerce site but will be used for a social media platform. The team observes a drop in accuracy. Which concept best explains this issue?
3A data engineer needs to design a data pipeline for a real-time fraud detection system. The system requires low-latency processing of streaming transactions. Which architecture is most appropriate?
4A team is training a deep learning model for image classification. The training loss decreases rapidly but validation loss starts increasing after a few epochs. Which regularization technique should be applied to mitigate this issue?
5A data analyst is cleaning a dataset and finds that 20% of the values for the 'age' column are missing. Which imputation method is most robust if the data is not normally distributed?
6Which TWO techniques are commonly used for feature selection in machine learning? (Choose 2)
7Which THREE are common data preprocessing steps in a machine learning pipeline? (Choose 3)
8Which TWO are best practices for versioning machine learning models? (Choose 2)
9A data engineer is designing a pipeline for a streaming data application that uses a machine learning model to detect anomalies in real time. Which TWO practices should the engineer implement to ensure data quality and model reliability?
10A large e-commerce company uses a recommendation system based on collaborative filtering. The system uses a matrix factorization model that is trained nightly on the entire user-item interaction history. Recently, the company launched a flash sale with thousands of new products. Users are reporting that the recommendations are not showing the new products, even for users who have purchased them during the sale. The data engineering team notices that the new products have very few interactions in the training data. The model's loss on the validation set has increased, and the recall@10 metric has dropped from 0.45 to 0.32. The team needs to improve the recommendation of new items without retraining the entire model from scratch every hour. Which approach should the team take?
11A data scientist trains a deep learning model on a large dataset. The training loss decreases steadily but the validation loss starts increasing after 20 epochs. The scientist uses early stopping with patience=5. Which of the following is the MOST likely cause and best corrective action?
12A company streams sensor data from IoT devices. The data arrives as JSON messages at high velocity. Which data pipeline architecture is BEST suited to handle this streaming data for near-real-time analytics?
13A dataset for a binary classification problem has 95% of samples in class "0" and 5% in class "1". The data scientist trains a logistic regression model and achieves 95% accuracy. Which metric should the scientist primarily use to evaluate model performance?
14An e-commerce company needs to update its recommendation model continuously as user preferences change. The model currently retrains from scratch every night, but the training time is too long. Which approach would reduce training time while keeping the model up-to-date?
15A machine learning model for credit card fraud detection is deployed. The model's precision is 0.95 and recall is 0.60. The business cost of missing a fraud is very high. Which of the following should the team prioritize to reduce the number of false negatives?
16Which THREE practices are recommended for versioning machine learning models in a production environment?
17Refer to the exhibit. A data scientist reviews the MLflow run for a Random Forest model on customer churn data. What is the most likely issue with this model?
18A financial institution is training a risk assessment model. The dataset includes customer credit scores, income, age, and past loan defaults. During feature engineering, a data engineer creates a new feature 'income_to_debt_ratio'. Which type of feature engineering technique is this?
19A machine learning team is developing a model to predict server failure from telemetry data. They use a deep neural network with 3 hidden layers. After training, the model achieves 99% accuracy on training data but only 85% on validation data. Which technique should the team apply to reduce the generalization error?
20A data engineer needs to combine two datasets, each with unique customer_id, to include all records from both datasets. Which join type should be used?
21A data engineer needs to store training data in a format that supports columnar pruning during model training. Which storage format should they use?
22During model deployment, a data engineer notices that the model's predictions are consistently lower than expected due to a shift in the distribution of one feature between training and production. Which technique should be used to detect and quantify this shift?
23A data engineer is preparing a dataset for a binary classification model. The dataset has 10,000 samples with 100 features. To improve model performance and reduce training time, the engineer decides to perform feature selection. Which two techniques are appropriate for this task? (Select TWO).
24Refer to the exhibit. A data engineer is training a binary classification neural network. The loss fluctuates and does not converge. Which hyperparameter adjustment is most likely to stabilize training?
25Refer to the exhibit. A data engineer runs a validation report on the customers table. The "income" column has 12 null values. Which imputation strategy is most appropriate for this column?
26A company is deploying an AI model to recommend products. The model's training data included historical purchases from the past two years, but the business environment has changed significantly due to a market shift. What is the most likely issue affecting model performance?
27A data engineer is building a pipeline to ingest streaming data from IoT sensors. Which data storage solution is best suited for real-time analytics on timestamped sensor readings?
28During feature engineering, a data scientist creates a new feature that is a linear combination of two existing features. What risk does this pose to the model?
29A team is using a pre-trained language model for sentiment analysis. They want to adapt it to a specific domain with limited labeled data. Which approach is most efficient?
30A data pipeline processes customer data from multiple sources. The data quality check reveals duplicate records. Which step should the pipeline include to handle this?
31An AI model is deployed to a mobile app with limited computational resources. The model is a deep neural network with high latency. Which technique is best to reduce inference time?
32A data engineer is designing a feature store for machine learning. Which THREE components are essential for a feature store? (Choose THREE.)
33A data scientist notices that a binary classification model consistently predicts the majority class. Which data engineering technique should be applied?
34A team is building a regression model to predict house prices. Which data transformation is most appropriate if the target variable exhibits right skewness?
35A model's training accuracy is 99% but validation accuracy drops to 60%. What is the most likely issue?
36An organization uses a machine learning model to approve loans. The model shows higher false positive rates for a protected group. Which data engineering step should be taken to mitigate this?
37A data pipeline ingests streaming data from IoT sensors. The current batch processing pipeline causes stale predictions. Which architecture change is most appropriate?
38A fraud detection model has high precision but low recall. The cost of false negatives is very high. Which threshold adjustment should be made?
39Which TWO data preprocessing techniques reduce the dimensionality of a dataset?
40Which THREE are common causes of data leakage in machine learning pipelines?
41A deep learning model for image classification achieves 99% training accuracy but only 85% validation accuracy. The model has millions of parameters. Which technique is most likely to reduce overfitting while maintaining high accuracy?
42A machine learning engineer is training a Support Vector Machine (SVM) with an RBF kernel on a dataset with features on different scales (e.g., age 0-100, income 0-1,000,000). The model converges slowly and yields poor accuracy. What should the engineer do first?
43A medical imaging team is developing an AI model to detect tumors from CT scans. They have 10,000 labeled scans, but the labels were created by a semi-automated process with an estimated 20% error rate (mislabeled tumor vs. no tumor). The team trains a convolutional neural network (CNN) and achieves 90% accuracy on a held-out test set that was carefully validated by an expert radiologist. However, when deployed to a new hospital's patient population, the accuracy drops to 70%. The team suspects domain shift and label noise. Which strategy is most likely to improve model robustness for the new hospital?
44An e-commerce company deploys a model to recommend products to users. The recommendation system uses collaborative filtering based on user-item interaction history. After deployment, the model shows decreasing click-through rates (CTR) over time. The data engineer notices that the model was trained on data from the past six months and is retrained daily. However, the trend suggests that user preferences are shifting more rapidly than expected. The engineer suspects that the model is suffering from distribution drift. Which approach should the engineer implement to adapt the model more quickly to changing user behavior?
45A healthcare company is developing a predictive model to identify patients at risk of readmission within 30 days. The data engineering team has built a pipeline that collects data from multiple sources, including electronic health records (EHR), lab results, and wearable device data. During initial testing, the model's performance is poor, with high false positives. Upon investigation, the team discovers that the data contains significant temporal misalignment: lab results are timestamped when ordered, not when collected; wearable data is aggregated hourly; and EHR data has inconsistent update frequencies. The data pipeline currently joins all features on the patient ID without aligning timestamps. The data volume is large, and processing time is a concern. Which action should the data engineering team take to most effectively address the issue and improve model performance?
46A retail company is building a recommendation system to suggest products to customers based on their purchase history. The data engineering team has collected data from point-of-sale systems, online browsing logs, and customer reviews. After cleaning the data, they notice that the feature set has over 500 dimensions, leading to high computational costs and potential overfitting. They need to reduce dimensionality while preserving as much variance as possible for the model. The team is considering various techniques. Which approach should they take to achieve this goal most effectively?
47A logistics company uses a machine learning model to predict delivery times based on historical data. The model was performing well, but recently it started making inaccurate predictions, especially for routes that have experienced new traffic patterns and road closures. The data engineering team receives an alert that the model's accuracy has dropped by 15% over the last week. They suspect data drift. The team has access to the original training data and a continuous stream of new data. What is the most appropriate first step for the team to take?
48A media company is building a recommendation model from user clickstream logs. The raw data arrives as millions of small JSON files in an object store, and nightly training jobs currently take over ten hours because the training cluster reads thousands of tiny files per second. The team wants to reduce training time without changing the model or the underlying data values. Which data engineering approach is most appropriate?
49A fraud detection team trains a gradient boosted tree model on transaction data. During evaluation, the team notices the model performs extremely well on the training set but poorly on a holdout set drawn from the same time period. Investigation shows that a feature named 'chargeback_flag' is populated only after a dispute is resolved, sometimes weeks after the transaction. The team wants to deploy the model to score transactions in real time. Which action best addresses the problem?
50A retail analytics team is preparing a dataset of product reviews for a sentiment classification model. The dataset contains 50,000 reviews, but only 2,000 are labeled as positive or negative. The team wants to use the unlabeled reviews to improve model performance. Which approach best leverages the unlabeled data?
51A data engineer is designing a pipeline to ingest high-velocity clickstream events from a web application into a data lake. The events must be queryable within minutes of arrival, and the schema evolves frequently as new fields are added. Which storage approach best meets these requirements?
52A retail analytics team is preparing a dataset for a demand forecasting model. The dataset contains a 'store_id' column with several thousand unique values, a 'product_category' column with about twenty values, and a 'day_of_week' column. The team wants to encode these categorical variables so a tree-based model can use them effectively without creating an enormous number of columns. Which TWO encoding approaches are most appropriate? (Choose two.)
53A machine learning engineer is preparing a dataset for a natural language processing task. The dataset contains text reviews with varying lengths, and the engineer plans to use a transformer model. Which preprocessing step is most critical to ensure the model can handle the input effectively?
54A data engineer is preparing a dataset for a machine learning model that predicts customer lifetime value. The dataset contains a 'monthly_spend' column with a highly right-skewed distribution. The team wants to apply a transformation to reduce skewness while preserving the relative order of values. Which transformation is most appropriate?
55A data scientist is building a model to predict equipment failure using sensor data. The dataset contains time-series readings from multiple sensors, and the goal is to detect anomalies that precede failures. Which TWO feature engineering techniques are most appropriate for this time-series data? (Choose two.)
56A machine learning engineer is building a model to predict whether a customer will make a purchase within the next week. The dataset contains 10,000 samples with 20 features, and the target variable is binary. The engineer wants to use a model that provides interpretable results to explain predictions to business stakeholders. Which model is most appropriate?
57A data analyst is exploring a dataset and notices that one numerical feature has a highly skewed distribution with a long right tail. The analyst wants to apply a transformation to make the distribution more symmetric for a linear model. Which transformation is most appropriate?
58A machine learning team is deploying a model that predicts loan default probabilities. The model outputs a probability score, and the team wants to convert it into a binary decision (default/no default). The costs of false positives and false negatives are not equal; a false negative (predicting no default when the customer defaults) is five times more costly than a false positive. Which approach best optimizes the decision threshold?
59A data scientist is preparing a dataset for a natural language processing task. The dataset contains a 'review_text' column with free-form customer reviews. Before feeding the text into a machine learning model, the team wants to convert the text into numerical features. Which technique is most appropriate for this purpose?
60A machine learning engineer is preparing a dataset for a model that predicts whether a customer will click on an ad. The dataset contains a feature 'time_since_last_purchase' measured in hours, which has a highly skewed distribution with a long tail. The engineer decides to apply a logarithmic transformation to this feature. Which statement BEST describes the effect of this transformation?
61A data scientist is building a model to predict the likelihood of a customer defaulting on a loan. The dataset contains a feature 'debt_to_income_ratio' that is highly skewed with a long right tail. The scientist decides to apply a logarithmic transformation to this feature. Which statement best describes the effect of this transformation?
62A data engineer is designing a data pipeline for a machine learning model that predicts equipment failure in a manufacturing plant. The pipeline ingests sensor data every second, and the model must be retrained daily with the latest data. The team needs to ensure that the training data is representative of the current operating conditions and that the model does not become stale. Which strategy is most appropriate for maintaining the training dataset?
63A data scientist is building a model to detect anomalies in server logs. The dataset contains millions of log entries, each with a timestamp and a message. The scientist wants to create features that capture the frequency of certain keywords (e.g., 'error', 'timeout') over time. Which approach is MOST appropriate for creating these features while avoiding data leakage?
64A data scientist is preparing a dataset for training a machine learning model. The dataset contains a mix of numerical and categorical features, and some features have high cardinality. The data scientist needs to apply appropriate encoding techniques to transform categorical variables into a format suitable for the model. Which TWO encoding methods are most appropriate for high-cardinality categorical features? (Choose two.)
65A data engineer is building a pipeline to ingest and process data from various sources for an AI model. The pipeline must handle both structured data from relational databases and unstructured text from documents. The engineer needs to ensure data quality and prepare the data for model training. Which TWO actions are MOST appropriate for handling missing values in the structured data? (Choose two.)
66A data scientist is working on a project to classify images of handwritten digits. The dataset consists of 60,000 training images and 10,000 test images, each 28x28 pixels in grayscale. The scientist wants to build a model that can automatically extract features and achieve high accuracy. Which type of model is most suitable for this task?
67A data scientist is preparing a dataset for a machine learning model and notices that one feature has a range from 0 to 1,000,000, while another feature ranges from 0 to 1. The model to be used is a k-nearest neighbors (KNN) classifier. Which preprocessing step is MOST important to apply before training?
Deep-dive questions
The most-searched questions in this domain — detailed explanations, worked examples, full answer breakdowns.
Be able to diagnose a dataset or metric and choose the right fix: median imputation for skewed data, threshold tuning to cut false negatives, and consistent scaling. The single most important thing is matching the remedy to the stated cost or distribution, not applying a default.
The Courseiva AI0-001 question bank contains 67 questions in the AI Models and Data Engineering domain, covering the 10% of the exam attributed to this domain in the official CompTIA blueprint. Click any question to see the full explanation and answer breakdown.
Start with a 10-question focused session to identify your baseline accuracy in this domain. Read every explanation — even for questions you answer correctly — to understand the reasoning. Once you score consistently above 80%, move to a 20–30 question session to confirm depth before moving to the next domain.
Yes — the session launcher on this page draws questions exclusively from the AI Models and Data Engineering domain. Choose 10, 20, 30, or 50 questions for a focused session, or click individual questions to review them one by one.
Save your results, see per-domain analytics, and get readiness scores — free, for every certification.
Sign Up FreeFree forever · Every certification included