Courseiva

GCIH Malware and AI-Assisted Investigations Practice Question

An analyst is training a machine learning model to classify malware families. Which data preparation technique is most critical to prevent bias in the classification results?

⚠ Common exam trap

Candidates often confuse data balancing with feature selection or hyperparameter tuning, assuming that removing noise or adjusting model complexity solves class imbalance issues.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Ensuring a balanced distribution of samples across all malware classes.

Data balance is the foundation of effective machine learning. If a training set is heavily skewed toward one malware family, the model will develop a prediction bias that favors that family, leading to poor classification performance for novel variants. This matters because biased models provide a false sense of security, causing investigators to overlook emerging threats that do not fit the over-represented patterns learned during the training phase.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    Including only the most recent malware samples in the dataset.

    Why it's wrong here

    Using only recent samples ignores historical patterns and evolutionary trends in malware development. A model trained this way will struggle to identify older, repurposed code or common delivery mechanisms that still exist in the environment. Diversity in the dataset is required to build a robust model capable of broad threat detection.

  • ✗

    Manually labeling only a small subset of the total available samples.

    Why it's wrong here

    Small, manually labeled datasets are prone to human error and insufficient coverage. If the training data does not represent the full spectrum of malware behavior, the model will lack the generalization required to classify unseen threats accurately. Scalable, representative data labeling is essential for creating high-performance, reliable classification models.

  • ✓

    Ensuring a balanced distribution of samples across all malware classes.

    Why this is correct

    Balanced data prevents the model from favoring majority classes. By providing an equal representation of various malware families, the algorithm learns the distinct features of each, leading to higher accuracy during inference. This balanced approach is critical for maintaining high detection rates across diverse threat vectors during active incident investigations.

  • ✗

    Removing all features that contain obfuscated code or strings.

    Why it's wrong here

    Obfuscation is a primary characteristic of malware. Removing these features strips the model of the very indicators it needs to distinguish malicious from benign code. Instead of removing these features, the model should be trained to recognize the patterns associated with obfuscation, which is essential for accurate classification and analysis.

About these practice questions

This GCIH question is part of Courseiva's 322-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official GIAC exam blueprint

This GCIH practice question is part of Courseiva's free GIAC certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the GCIH exam.