Courseiva
Data Preparation →hardMultiple Choice

NCP-GENL Data Preparation Practice Question

Exhibit

{
  "action": "filter",
  "criterion": "min_lexical_diversity",
  "threshold": 0.2
}

Refer to the exhibit. What is the impact of this filter on the training corpus?

⚠ Common exam trap

Test-takers often misinterpret lexical diversity filters as removing long documents or high-frequency vocabulary, whereas they actually target repetitive, low-value boilerplate text.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

It discards documents that are overly repetitive or lack linguistic richness.

This filter removes documents with low lexical diversity, which often contain repetitive, low-value, or 'boilerplate' text. Such content provides little signal for the model to learn meaningful language patterns. By enforcing a minimum diversity threshold, you ensure the corpus consists of richer, more informative language, which typically leads to better convergence and higher quality output in the trained model.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    It limits the vocabulary size to 20% of the original content.

    Why it's wrong here

    Lexical diversity is a measure of the variety of words used in a text, not a cap on the vocabulary size. The threshold of 0.2 indicates a minimum level of language variety required for a document to be included, not a limitation on the number of unique tokens.

  • ✓

    It discards documents that are overly repetitive or lack linguistic richness.

    Why this is correct

    A low lexical diversity score indicates that a document uses a very small set of unique words relative to its total length. This is characteristic of repetitive or low-quality content. Filtering these out ensures the model learns from diverse, high-quality, and informative text, which improves overall model performance.

  • ✗

    It forces the model to use 20% more computation for tokenization.

    Why it's wrong here

    The filter operates on the text during the data preparation stage, long before tokenization or model training occur. It has absolutely no impact on the computational resources required for tokenization or model training. It is purely a data selection step to improve the quality of the training set.

  • ✗

    It reduces the training corpus size by exactly 20%.

    Why it's wrong here

    A threshold of 0.2 refers to the lexical diversity score, not the percentage of data to be removed. The actual reduction in corpus size depends entirely on how much of the dataset fails to meet this diversity threshold, which is determined by the quality of the raw data.

About these practice questions

This NCP-GENL question is part of Courseiva's 352-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official NVIDIA exam blueprint

This NCP-GENL practice question is part of Courseiva's free NVIDIA certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the NCP-GENL exam.