Courseiva

AI0-001 AI Models and Data Engineering Practice Question

A retail analytics team is preparing a dataset for a demand forecasting model. The dataset contains a 'store_id' column with several thousand unique values, a 'product_category' column with about twenty values, and a 'day_of_week' column. The team wants to encode these categorical variables so a tree-based model can use them effectively without creating an enormous number of columns. Which TWO encoding approaches are most appropriate? (Choose two.)

⚠ Common exam trap

The trap here is applying one encoding method uniformly across all categorical features instead of matching the technique to each feature's cardinality.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Use target encoding for 'store_id', replacing each store with a statistic such as the mean of the target computed from training folds.

High-cardinality identifiers like store_id need a compact numeric representation, and out-of-fold target encoding provides one column that captures store-level signal without exploding dimensionality. Low-cardinality nominal features like product_category and day_of_week are best handled with one-hot encoding, which avoids implying an order and keeps the feature space small. Together these choices match encoding strategy to cardinality.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✓

    Use target encoding for 'store_id', replacing each store with a statistic such as the mean of the target computed from training folds.

    Why this is correct

    Target encoding maps each high-cardinality store to a single numeric value derived from the target, so thousands of stores become one column instead of thousands of indicator columns. When computed with out-of-fold statistics it avoids leaking the current row's target into its own feature, and tree models can split on the resulting numeric value efficiently.

  • ✗

    Use one-hot encoding for 'store_id' so every store gets its own binary column.

    Why it's wrong here

    One-hot encoding a column with several thousand unique values creates thousands of sparse binary columns, which inflates memory, slows training, and forces the model to consider an impractical number of candidate splits. It can also cause overfitting on stores with few observations, and the resulting matrix is awkward for most tree implementations.

  • ✗

    Use binary encoding for 'day_of_week' so each day is represented by a binary code across multiple columns.

    Why it's wrong here

    Binary encoding is designed to compress high-cardinality features, but 'day_of_week' has only seven values, so one-hot encoding is simpler and avoids introducing an artificial numeric structure. Applying binary encoding here adds complexity without meaningful dimensionality reduction and can obscure the clear cyclic meaning of weekdays.

  • ✓

    Use one-hot encoding for 'product_category' and 'day_of_week' because they have low cardinality.

    Why this is correct

    With only about twenty categories and seven days, one-hot encoding produces a small, dense, and interpretable representation that tree models handle well. Each category becomes its own binary feature, no ordinal relationship is implied, and the total column count remains modest, making this a standard and safe choice for low-cardinality variables.

  • ✗

    Use label encoding for 'product_category' so categories become arbitrary integers based on alphabetical order.

    Why it's wrong here

    Assigning arbitrary integers to unordered categories introduces a false ordering that tree models may exploit, producing splits that treat one category as greater than another when no such relationship exists. This can degrade accuracy and make the model's behavior harder to explain, especially for a nominal feature like product category.

About these practice questions

This AI0-001 question is part of Courseiva's 962-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official CompTIA exam blueprint

This AI0-001 practice question is part of Courseiva's free CompTIA certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the AI0-001 exam.