NCP-GENL Data Preparation Practice Question
Exhibit
{
"action": "filter",
"criterion": "min_lexical_diversity",
"threshold": 0.2
}Refer to the exhibit. What is the impact of this filter on the training corpus?
⚠ Common exam trap
Test-takers often misinterpret lexical diversity filters as removing long documents or high-frequency vocabulary, whereas they actually target repetitive, low-value boilerplate text.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
It discards documents that are overly repetitive or lack linguistic richness.
This filter removes documents with low lexical diversity, which often contain repetitive, low-value, or 'boilerplate' text. Such content provides little signal for the model to learn meaningful language patterns. By enforcing a minimum diversity threshold, you ensure the corpus consists of richer, more informative language, which typically leads to better convergence and higher quality output in the trained model.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
It limits the vocabulary size to 20% of the original content.
Why it's wrong here
Lexical diversity is a measure of the variety of words used in a text, not a cap on the vocabulary size. The threshold of 0.2 indicates a minimum level of language variety required for a document to be included, not a limitation on the number of unique tokens.
- ✓
It discards documents that are overly repetitive or lack linguistic richness.
Why this is correct
A low lexical diversity score indicates that a document uses a very small set of unique words relative to its total length. This is characteristic of repetitive or low-quality content. Filtering these out ensures the model learns from diverse, high-quality, and informative text, which improves overall model performance.
- ✗
It forces the model to use 20% more computation for tokenization.
Why it's wrong here
The filter operates on the text during the data preparation stage, long before tokenization or model training occur. It has absolutely no impact on the computational resources required for tokenization or model training. It is purely a data selection step to improve the quality of the training set.
- ✗
It reduces the training corpus size by exactly 20%.
Why it's wrong here
A threshold of 0.2 refers to the lexical diversity score, not the percentage of data to be removed. The actual reduction in corpus size depends entirely on how much of the dataset fails to meet this diversity threshold, which is determined by the quality of the raw data.
About these practice questions
This NCP-GENL question is part of Courseiva's 352-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official NVIDIA exam blueprint
This NCP-GENL practice question is part of Courseiva's free NVIDIA certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the NCP-GENL exam.