Courseiva

AI0-001 AI Models and Data Engineering Practice Question

A retail company is building a recommendation system to suggest products to customers based on their purchase history. The data engineering team has collected data from point-of-sale systems, online browsing logs, and customer reviews. After cleaning the data, they notice that the feature set has over 500 dimensions, leading to high computational costs and potential overfitting. They need to reduce dimensionality while preserving as much variance as possible for the model. The team is considering various techniques. Which approach should they take to achieve this goal most effectively?

⚠ Common exam trap

CompTIA AI+ exam questions often test the distinction between dimensionality reduction techniques (PCA) and feature selection methods (Lasso, correlation-based selection) or visualization tools (t-SNE), expecting candidates to recognize that PCA is the only option that explicitly reduces dimensionality while preserving maximum variance in a way that is suitable for downstream modeling.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Use Principal Component Analysis (PCA) to reduce the feature space to the top 50 principal components that explain 95% of the variance.

Principal Component Analysis (PCA) is a linear dimensionality reduction technique that transforms the original high-dimensional feature space into a set of orthogonal principal components, ordered by the amount of variance they capture. By selecting the top 50 components that explain 95% of the variance, the team effectively reduces the feature set from over 500 dimensions while preserving the most informative structure in the data, directly addressing the goals of lowering computational cost and mitigating overfitting.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    Keep all features but apply L1 regularization (Lasso) in the model to automatically reduce coefficients to zero.

    Why it's wrong here

    Lasso shrinks coefficients to zero, performing feature selection, but it does not project the 500-dimensional data onto a lower-dimensional space, so the variance-preservation requirement goes unmet. It is tempting because Lasso genuinely combats overfitting in high-dimensional regression; it would suit sparse linear models, not dimensionality reduction.

  • ✗

    Apply t-Distributed Stochastic Neighbor Embedding (t-SNE) to reduce the feature space to 50 dimensions.

    Why it's wrong here

    t-SNE is a non-linear technique for visualising high-dimensional data in two or three dimensions; it does not preserve global variance, so 50 components would distort the structure the model needs. It is tempting for exploratory clustering plots, where t-SNE is the right choice.

  • ✗

    Select only features that have a high correlation with the target variable, discarding all others.

    Why it's wrong here

    Correlation filtering discards features individually, ignoring multivariate structure, so it cannot preserve overall variance across the 500 dimensions. It is tempting because correlation-based selection is a legitimate, cheap feature-selection technique; it would be correct when redundant or irrelevant predictors must be pruned before training a supervised model.

  • ✓

    Use Principal Component Analysis (PCA) to reduce the feature space to the top 50 principal components that explain 95% of the variance.

    Why this is correct

    PCA transforms the 500 correlated features into orthogonal components ranked by explained variance, so 50 components capture 95% of the variance. This satisfies the stem's dual constraint: cut dimensionality while preserving variance, unlike feature selection, which discards rather than recombines information.

About these practice questions

This AI0-001 question is part of Courseiva's 962-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

This AI0-001 practice question is part of Courseiva's free CompTIA certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the AI0-001 exam.