Courseiva

AI0-001 AI Models and Data Engineering Practice Question

A data engineer is building a pipeline to ingest and process data from various sources for an AI model. The pipeline must handle both structured data from relational databases and unstructured text from documents. The engineer needs to ensure data quality and prepare the data for model training. Which TWO actions are MOST appropriate for handling missing values in the structured data? (Choose two.)

⚠ Common exam trap

The trap here is assuming that removing all rows with missing values is always safe or that filling with zero is harmless, when both can introduce bias and degrade model performance.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Impute missing numerical values with the median of the available data.

For structured data with missing numerical values, median imputation is a robust simple method, while KNN imputation leverages feature correlations for more accurate estimates. Both preserve data and avoid the biases of deletion or arbitrary constant filling. These actions are appropriate for preparing data for model training, ensuring quality and completeness without introducing severe distortions.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✓

    Impute missing numerical values with the median of the available data.

    Why this is correct

    Imputing with the median is a robust method for numerical features because it is less sensitive to outliers than the mean. It preserves the central tendency and allows the model to use all samples. This is a common and effective approach when missingness is random and the proportion of missing data is not too high.

  • ✗

    Remove all rows with any missing values to ensure a complete dataset.

    Why it's wrong here

    Removing all rows with missing values can lead to significant data loss and bias if missingness is not completely random. It reduces the sample size and may disproportionately remove certain groups, harming model generalization. Unless missingness is very low and random, this is generally not recommended.

  • ✗

    Encode missing values as a separate category for numerical features.

    Why it's wrong here

    Treating missingness as a separate category is typically used for categorical features, not numerical. For numerical features, it would require converting the feature to categorical or adding a binary indicator, which may not be the best primary approach. While adding a missing indicator can be useful, encoding missing as a category for a numerical feature is not standard and can complicate modeling.

  • ✗

    Use a constant value like zero to fill missing numerical values.

    Why it's wrong here

    Filling with zero can distort the distribution and introduce bias, especially if zero is not a meaningful value for the feature. It can mislead the model into treating missingness as a specific value, which may not reflect reality. This approach is only appropriate if zero has a clear interpretation, such as absence of an event.

  • ✓

    Apply a model-based imputation method such as k-nearest neighbors (KNN) imputation.

    Why this is correct

    KNN imputation estimates missing values based on similar samples, preserving relationships between features. It can be more accurate than simple statistics when features are correlated. This method is suitable for structured data and can improve model performance by providing more realistic values, though it is computationally intensive.

About these practice questions

This AI0-001 question is part of Courseiva's 962-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official CompTIA exam blueprint

This AI0-001 practice question is part of Courseiva's free CompTIA certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the AI0-001 exam.