Courseiva

AI0-001 AI Infrastructure and Technologies Practice Question

A cybersecurity firm is building an anomaly detection system for network traffic. The dataset contains millions of connection records with dozens of features, but only 0.1% are labeled as malicious. The team needs a model that can flag suspicious connections while minimizing false positives that overwhelm analysts. Which approach is most appropriate?

⚠ Common exam trap

The trap here is assuming that class weighting or dimensionality reduction alone can overcome a 0.1% positive rate, when the scarcity of labels makes unsupervised anomaly detection the only practical path.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Train an isolation forest or autoencoder on the normal traffic to learn its structure, then flag deviations as anomalies.

Extreme class imbalance with only 0.1% malicious labels makes supervised learning unreliable. Unsupervised anomaly detection methods like isolation forest and autoencoders learn the structure of normal traffic without requiring labels, then flag deviations. Isolation forest isolates anomalies by random partitioning, while autoencoders flag high reconstruction error. Both allow threshold tuning to balance detection and false positives, which is critical for analyst workload. This makes them the most appropriate choice for the cybersecurity scenario.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    Apply principal component analysis (PCA) to reduce dimensionality, then use a threshold on reconstruction error to detect anomalies.

    Why it's wrong here

    PCA is a linear method that captures global variance, but network traffic anomalies often lie in nonlinear manifolds that PCA cannot represent. A linear projection may miss subtle malicious patterns, causing false negatives. While reconstruction error can be used, the linear assumption limits effectiveness compared to nonlinear autoencoders. PCA also does not inherently handle the extreme imbalance or provide the flexibility needed to tune false positives.

  • ✗

    Use k-means clustering with k set to the number of known attack types and treat points far from centroids as anomalies.

    Why it's wrong here

    K-means assumes spherical, similarly sized clusters and requires specifying k, which is unknown for evolving attacks. It is sensitive to outliers, which are exactly the anomalies of interest, and may absorb them into clusters. It also does not model the density of normal traffic well, leading to poor anomaly separation. This makes it unsuitable for flagging rare malicious connections with low false positives.

  • ✓

    Train an isolation forest or autoencoder on the normal traffic to learn its structure, then flag deviations as anomalies.

    Why this is correct

    With only 0.1% malicious labels, supervised classification is impractical due to extreme class imbalance. Unsupervised anomaly detection methods like isolation forest and autoencoders learn the distribution of normal traffic and flag deviations, requiring no labels. This directly suits the scenario, and by tuning the contamination parameter or reconstruction error threshold, the team can control the false positive rate to keep analysts from being overwhelmed.

  • ✗

    Train a supervised gradient boosting classifier on the labeled data and use class weights to handle imbalance.

    Why it's wrong here

    Even with class weighting, 0.1% positive labels means the model has very few malicious examples to learn from, leading to poor generalization and many false negatives. Supervised methods also cannot detect novel attack types not present in the training labels. The extreme scarcity of labels makes this approach unreliable for a robust anomaly detection system in cybersecurity.

About these practice questions

This AI0-001 question is part of Courseiva's 962-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official CompTIA exam blueprint

This AI0-001 practice question is part of Courseiva's free CompTIA certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the AI0-001 exam.