NCP-GENL Data Preparation Practice Question
You are preparing a large corpus of customer support transcripts for continued pretraining of an NVIDIA NeMo Megatron model. The transcripts contain personally identifiable information such as names, email addresses, and account numbers, and company policy requires that this information be removed before training while preserving as much linguistic context as possible for the model to learn from. Which data preparation approach best satisfies both requirements?
⚠ Common exam trap
The trap here is treating PII removal as a binary choice between deleting records and destroying all structure, when span-level replacement preserves the linguistic signal.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Detect PII spans with a combination of regular expressions and a named-entity recognition model, then replace each detected span with a consistent placeholder token that preserves the surrounding sentence structure.
Combining regex detection for structured identifiers with named-entity recognition for names and locations covers the PII surface area, and replacing detected spans with consistent placeholders removes sensitive values while preserving the sentence context that continued pretraining depends on. Dropping records, masking all digits and capitals, or hashing every token either destroys usable language data or fails to anonymize reliably.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Hash every token in the corpus so that the original text cannot be reconstructed, and train on the hashed sequences.
Why it's wrong here
Hashing every token destroys the semantic and syntactic structure the model needs to learn language from; the resulting sequences carry no transferable linguistic signal. It also does not reliably anonymize free-form PII because hashing is deterministic and can be attacked with dictionary lookups for common names or emails. The approach fails both the privacy guarantee and the context preservation requirement.
- ✗
Replace all digits and capitalized words with a generic mask token across the entire corpus before training.
Why it's wrong here
Masking all digits and capitalized words destroys legitimate content such as product names, dates, and technical terms, producing text that no longer resembles real support language. The model would learn from a distorted distribution, and the masking is not targeted enough to guarantee that every PII span is caught, since names can appear in lowercase. Privacy improves only incidentally while data quality suffers broadly.
- ✓
Detect PII spans with a combination of regular expressions and a named-entity recognition model, then replace each detected span with a consistent placeholder token that preserves the surrounding sentence structure.
Why this is correct
Regexes reliably catch structured identifiers such as emails and account numbers, while a named-entity recognition model covers names and locations that patterns miss. Replacing spans with consistent placeholders removes the sensitive values but keeps sentence structure and surrounding context intact, which is exactly what continued pretraining needs to learn language patterns without memorizing private data.
- ✗
Drop every transcript that contains any detected PII so that no sensitive information reaches the training pipeline.
Why it's wrong here
Discarding all transcripts with PII removes a large fraction of the corpus, shrinking the training set and biasing it toward interactions that happen to lack names or contact details. Customer support text almost always contains such information, so the remaining data would be unrepresentative. This approach satisfies the privacy requirement but fails the goal of preserving linguistic context for the model to learn from.
About these practice questions
One of 352 original NCP-GENL practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official NVIDIA exam blueprint
This NCP-GENL practice question is part of Courseiva's free NVIDIA certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the NCP-GENL exam.