Courseiva

AI0-001 · topic practice

AI Models and Data Engineering practice questions

This domain covers the data lifecycle behind AI systems: sourcing, cleaning, transforming, and validating data, plus selecting and evaluating models. Questions are scenario-based, asking you to pick an imputation strategy, fix class imbalance, tune a decision threshold, or interpret precision, recall, and regression metrics for a stated business cost.

Courseiva uses original exam-style practice questions designed for learning and revision. The goal is to understand the concepts, recognise exam patterns, and improve through explanations — not memorise copied exam dumps.

Editorial oversight:Johnson Ajibi· MSc IT Security, IEEE Senior Member
20 questionsDomain: AI Models and Data Engineering

What the exam tests

What to know about AI Models and Data Engineering

Be able to diagnose a dataset or metric and choose the right fix: median imputation for skewed data, threshold tuning to cut false negatives, and consistent scaling. The single most important thing is matching the remedy to the stated cost or distribution, not applying a default.

Choosing imputation (mean, median, mode, or model-based) based on distribution and missingness

Adjusting classification decision thresholds to trade precision against recall for business cost

Feature scaling and normalization for regression inputs with widely differing numeric ranges

Splitting data into train, validation, and test sets and detecting leakage or imbalance

Watch out for

Common AI Models and Data Engineering exam traps

  • ▸Using mean imputation on skewed or non-normal data instead of median, which distorts the distribution
  • ▸Lowering the threshold when false negatives are costly; you must lower it, not raise it, to catch more positives
  • ▸Scaling the test set with its own statistics rather than reusing parameters fit on the training data

Practice set

AI Models and Data Engineering questions

20 questions · select your answer, then reveal the explanation

A data scientist is preparing a dataset for training a classification model. The dataset contains 10,000 records with a binary target variable where 9,500 belong to class A and 500 belong to class B. Which technique should the scientist use to address the class imbalance?

An organization needs to store sensitive customer data for training a machine learning model. The data must be encrypted at rest and in transit, and access must be audited. Which combination of practices should be implemented?

An engineer is training a neural network and observes the output shown. Which conclusion is most likely correct?

Exhibit

Refer to the exhibit.

```
Epoch 1/10 - loss: 0.6932 - accuracy: 0.5234 - val_loss: 0.6918 - val_accuracy: 0.5312
Epoch 2/10 - loss: 0.4231 - accuracy: 0.8047 - val_loss: 0.5234 - val_accuracy: 0.7422
Epoch 3/10 - loss: 0.3125 - accuracy: 0.8828 - val_loss: 0.6015 - val_accuracy: 0.7344
Epoch 4/10 - loss: 0.2146 - accuracy: 0.9219 - val_loss: 0.7234 - val_accuracy: 0.7188
Epoch 5/10 - loss: 0.1478 - accuracy: 0.9531 - val_loss: 0.8342 - val_accuracy: 0.7031
```

A data engineer is reviewing an S3 bucket policy for a machine learning project. The policy is intended to allow access to training data only from the corporate network (10.0.0.0/16). However, users in the corporate network report access denied. Which issue is most likely causing the problem?

Exhibit

Refer to the exhibit.

```
{
  "Version": "2012-10-17",
  "Statement": [
    {
      "Effect": "Allow",
      "Action": "s3:GetObject",
      "Resource": "arn:aws:s3:::ml-training-data/*",
      "Condition": {
        "IpAddress": {
          "aws:SourceIp": "10.0.0.0/16"
        }
      }
    }
  ]
}
```

A data scientist is training a deep learning model for image classification. The training loss decreases steadily but the validation loss starts increasing after 10 epochs. Which technique should the scientist apply to address this issue?

A financial institution is building a fraud detection system using a supervised learning model. The dataset is highly imbalanced with 99.9% legitimate transactions and 0.1% fraudulent ones. Which approach would be MOST effective to train the model to detect fraud?

A company wants to deploy an AI model for real-time inference on edge devices with limited computational resources. Which model architecture would be MOST suitable?

A team is developing a natural language processing model to classify customer feedback. The dataset contains text in multiple languages. Which THREE preprocessing steps are essential to ensure the model performs well across all languages?

A healthcare startup is developing a deep learning model to detect diabetic retinopathy from retinal images. The model is trained on a dataset of 10,000 labeled images. During initial testing, the model achieves 99% accuracy on the training set but only 85% on the test set. The startup wants to deploy the model in a clinical setting where false negatives (missing a disease) are critical. The team has access to additional unlabeled retinal images from multiple sources. Which strategy should the team use to improve the model's generalization and reduce false negatives?

A data scientist is preparing a dataset for a classification model. The dataset contains a column "Age" with 10% missing values and a column "Income" with 30% missing values. Which imputation strategy is MOST appropriate to minimize bias?

A team is building a regression model to predict house prices. The dataset includes numerical features (square footage, number of bedrooms) and categorical features (neighborhood, roof type). The categorical features have high cardinality (neighborhood has 200+ unique values). Which encoding strategy should the team use to avoid overfitting and maintain model interpretability?

A data engineer discovers that a dataset contains duplicate rows. Which data cleaning step is MOST appropriate?

A data scientist is working with a dataset that has 10,000 features but only 500 samples. The goal is to train a model for binary classification. Which feature selection technique is MOST appropriate to reduce overfitting?

A data scientist is cleaning a dataset. Which TWO actions are appropriate for handling missing data?

Which THREE data quality dimensions are critical for ensuring model reliability?

Refer to the exhibit. A stream processor ingests events. One event arrives with missing "user_id". What will happen?

Exhibit

The following is a JSON schema snippet from a data pipeline:
{
  "type": "object",
  "properties": {
    "user_id": { "type": "integer" },
    "timestamp": { "type": "string", "format": "date-time" },
    "event_type": { "type": "string" },
    "value": { "type": "number" }
  },
  "required": ["user_id", "event_type", "value"]
}

Refer to the exhibit. What is the recall of the model?

Exhibit

The following is a confusion matrix for a binary classifier:

              Predicted: Positive  Predicted: Negative
Actual Positive:     80                 20
Actual Negative:     30                 70

A data scientist is preparing a dataset for training a classification model. The dataset has a column with missing values in 5% of rows. Which action should the data engineer take to minimize bias?

A team is training a language model using a large text corpus. They want to ensure the model does not learn biased associations between gender and professions. Which data engineering technique should they apply?

A streaming data pipeline ingests sensor data from IoT devices. The data arrives at irregular intervals and contains occasional spikes. Which data transformation is most appropriate for preparing this data for a time-series model?

Free account

Track your progress over time

Create a free account to save your results and see which topics improve across sessions.

Focused AI Models and Data Engineering sessions

Start a AI Models and Data Engineering only practice session

Every question in these sessions is drawn from the AI Models and Data Engineering domain — nothing else.

Related practice questions

Related AI0-001 topic practice pages

Move into related areas when this topic feels solid.

Frequently asked questions

What does the AI0-001 exam test about AI Models and Data Engineering?
Be able to diagnose a dataset or metric and choose the right fix: median imputation for skewed data, threshold tuning to cut false negatives, and consistent scaling. The single most important thing is matching the remedy to the stated cost or distribution, not applying a default.
How should I use these practice questions?
Select your answer before revealing the explanation. Then read why each option is right or wrong — this active recall approach builds retention far faster than re-reading notes.
Can I practise just AI Models and Data Engineering questions in a focused session?
Yes — the session launcher on this page draws every question from the AI Models and Data Engineering domain. Use a 10-question session first to gauge your baseline, then move to 20 or 30 once the weak spots are clear.
Where can I practise other AI0-001 topics?
Use the topic links above to move to related areas, or go back to the AI0-001 question bank to see all topics.
Are these real exam questions or dumps?
These are original practice questions written to test the same concepts the AI0-001 exam covers. They are not copied from any real exam or dump site.