Courseiva
Data Analysis →mediumMultiple Choice

DA0-002 Data Analysis Practice Question

A data analyst is working with a dataset that contains a column 'region' with values such as 'North', 'South', 'East', 'West', and 'N/A'. The analyst needs to prepare this column for a machine learning model. Which of the following is the most appropriate approach to handle the 'N/A' values?

⚠ Common exam trap

The trap here is assuming that missing values must be imputed or removed, when sometimes they represent a valid category that should be preserved.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Treat 'N/A' as a separate category and encode it as its own level.

Treating 'N/A' as a separate category is the best approach because it retains the information that the region is missing or not applicable, which could be informative. Imputing with the mode or removing rows can introduce bias or lose data, and numeric encoding creates a false ordinal relationship.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    Replace 'N/A' with the mode of the 'region' column.

    Why it's wrong here

    Imputing with the mode introduces bias by assigning the most frequent region to missing entries, which may not be accurate. If 'N/A' represents truly missing data, using the mode assumes the missingness is random and that the most common region is a reasonable guess, which can distort the distribution and mislead the model.

  • ✓

    Treat 'N/A' as a separate category and encode it as its own level.

    Why this is correct

    Treating 'N/A' as a distinct category preserves the information that the region is missing or not applicable, which can be predictive. This approach avoids introducing false assumptions and allows the model to learn any signal associated with missingness. It is a common and valid strategy for categorical variables.

  • ✗

    Encode 'N/A' as a numeric value of 0 and other regions as 1-4.

    Why it's wrong here

    Assigning an arbitrary numeric value to 'N/A' implies an ordinal relationship that does not exist and can mislead the model into interpreting 'N/A' as a meaningful numeric quantity. This is particularly problematic for tree-based models that might split on the value, but also for linear models that treat it as a continuous input.

  • ✗

    Remove all rows where 'region' is 'N/A'.

    Why it's wrong here

    Dropping rows with missing values can lead to loss of valuable data and introduce bias if the missingness is not completely random. If 'N/A' is a meaningful category, removing those rows would discard potentially important information and reduce the model's ability to generalize.

About these practice questions

This DA0-002 question is part of Courseiva's 1,004-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official CompTIA exam blueprint

This DA0-002 practice question is part of Courseiva's free CompTIA certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the DA0-002 exam.