NCP-GENL Data Preparation Practice Question
Exhibit
config.yaml:
filter_policy:
min_words: 50
max_perplexity: 100
deduplication:
algorithm: minhash
threshold: 0.95Refer to the exhibit. What is the intended outcome of this data cleaning configuration for a Large Language Model pre-training corpus?
⚠ Common exam trap
Candidates often assume cleaning configurations only remove empty strings, failing to recognize the combined role of length/perplexity filters and MinHash deduplication in eliminating noise.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
It optimizes for the removal of low-quality or nonsensical text while minimizing redundancy.
This configuration aims to remove low-quality text that fails to meet minimum length requirements or exhibits high perplexity (indicating gibberish or low coherence). Simultaneously, MinHash deduplication identifies and removes near-duplicate documents exceeding a 95% similarity threshold. This cleanup process is vital for pre-training, as it filters out low-value, noisy data that could impede model convergence and general quality during the training cycle.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
It specifically targets the removal of personally identifiable information (PII).
Why it's wrong here
The configuration settings focus on text length, coherence (perplexity), and document similarity. None of these parameters are designed to detect or sanitize personally identifiable information. PII removal requires specialized regex patterns or entity recognition models, which are not represented in this configuration file or its associated policies.
- ✓
It optimizes for the removal of low-quality or nonsensical text while minimizing redundancy.
Why this is correct
The configuration uses perplexity filtering to identify incoherent content and length constraints to exclude short, low-information strings. The MinHash algorithm effectively manages the similarity threshold to eliminate near-duplicate documents. This combination ensures that the training dataset is concise, coherent, and free of redundant, low-value information inputs.
- ✗
It enforces a strict length-based chunking strategy for all documents.
Why it's wrong here
The configuration sets a minimum word count as a filtering policy, not as a chunking strategy. Filtering removes entire documents that fail to meet the criteria, whereas chunking would split long documents into smaller segments. These are distinct processes in data preparation that serve different purposes.
- ✗
It converts all text to a vector space representation before filtering.
Why it's wrong here
The policy operates on the raw text corpus (perplexity, word count, document similarity). While MinHash uses hashing, it is a technique for efficient similarity estimation of text sets, not a general vectorization process. The filtering happens prior to any full-scale embedding or vectorization performed during training.
About these practice questions
This NCP-GENL question is part of Courseiva's 352-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official NVIDIA exam blueprint
This NCP-GENL practice question is part of Courseiva's free NVIDIA certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the NCP-GENL exam.