NCP-GENL Data Preparation Practice Question
You are preparing a large instruction-tuning dataset for an NVIDIA NeMo-based LLM. The raw data consists of user queries and assistant responses collected from a customer support system, stored as JSON lines. During preprocessing, you notice that many responses contain personally identifiable information (PII) such as names, email addresses, and phone numbers. You need to ensure the dataset is safe for training while preserving as much semantic content as possible. Which approach is most appropriate for handling PII in this dataset?
⚠ Common exam trap
The trap here is assuming that regex-based removal is sufficient, but PII can be unstructured and context-dependent, making NER more reliable.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Use a named entity recognition (NER) model to detect PII and replace each entity with a generic placeholder (e.g., [NAME]).
Replacing PII with generic placeholders using an NER model is the most balanced approach: it removes sensitive information while preserving the linguistic patterns and context needed for instruction tuning. This method aligns with NVIDIA NeMo's data preprocessing recommendations for responsible AI. It also avoids the pitfalls of discarding data or introducing unintelligible tokens.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Leave the PII intact but exclude the dataset from production use and train only in a sandbox environment.
Why it's wrong here
Keeping PII in training data poses legal and ethical risks even in a sandbox, as models can memorize and later regurgitate it. This approach does not address the core requirement of safe data preparation. De-identification is necessary regardless of the training environment.
- ✗
Hash each detected PII string with SHA-256 and substitute the hash into the text.
Why it's wrong here
Hashing PII creates nonsensical tokens that the model cannot interpret, harming training. It also does not fully anonymize if the hash space is small (e.g., phone numbers) and may still leak information. Placeholders are preferred for maintaining linguistic coherence.
- ✗
Apply a rule-based regex filter to remove any line containing PII, then train on the remaining data.
Why it's wrong here
Removing entire lines that contain PII discards valuable context and reduces dataset size unnecessarily. Regex can also miss obfuscated PII, and this approach does not preserve semantic content. It is overly aggressive and not recommended for instruction-tuning data where context matters.
- ✓
Use a named entity recognition (NER) model to detect PII and replace each entity with a generic placeholder (e.g., [NAME]).
Why this is correct
Replacing PII with placeholders preserves sentence structure and semantic meaning while removing sensitive information. This is a standard de-identification technique that maintains data utility for instruction tuning. It also handles variations better than simple regex and is compatible with NeMo preprocessing pipelines.
About these practice questions
This NCP-GENL question is part of Courseiva's 352-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official NVIDIA exam blueprint
This NCP-GENL practice question is part of Courseiva's free NVIDIA certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the NCP-GENL exam.