NCP-GENL Data Preparation Practice Question
A team is building an instruction-tuning dataset in NeMo from 40,000 internal support tickets. Each ticket contains a customer problem and a resolved answer, but the resolution text sometimes includes the customer's name, account number, and internal case IDs. The team plans to use NeMo Curator to produce training-ready JSONL. Which approach best prepares this data for instruction tuning while limiting personally identifiable information exposure?
⚠ Common exam trap
The trap here is treating PII handling as a post-training output filter, when the identifiers must be removed from the training corpus itself before the model ever sees them.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Run a PII redaction stage that detects and replaces names, account numbers, and case IDs with placeholders, then convert problem-resolution pairs into the instruction, input, and output schema before export.
Sensitive identifiers must be removed from the corpus before training, not mitigated after the fact. Running a PII detection and replacement stage in NeMo Curator, then mapping problem-resolution pairs into the instruction, input, and output schema, produces compliant JSONL that the fine-tuning job can consume directly. This ordering keeps identifiers out of both the training data and the resulting model weights.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Hash the entire ticket text with a cryptographic digest and train on the hashes, since the model can learn the mapping between hashed problems and hashed answers without ever seeing the original identifiers.
Why it's wrong here
A cryptographic hash is a one-way transformation that destroys the linguistic structure the model must learn. Training on hashes would yield a model incapable of generating coherent language because tokens no longer carry meaning. PII protection requires selective entity replacement that preserves surrounding natural language, not whole-record hashing that eliminates the instructional signal entirely.
- ✓
Run a PII redaction stage that detects and replaces names, account numbers, and case IDs with placeholders, then convert problem-resolution pairs into the instruction, input, and output schema before export.
Why this is correct
Redacting identifiers before schema conversion prevents the model from memorizing sensitive strings and keeps the pipeline deterministic. NeMo Curator supports PII detection and replacement as a distinct stage, and converting cleaned records into the instruction, input, and output fields yields the JSONL format the fine-tuning configuration expects. Doing redaction first also avoids leaking identifiers into derived fields.
- ✗
Convert the tickets into instruction, input, and output JSONL first, then fine-tune the model and rely on a post-training output filter to block any generated response that contains an account number or customer name.
Why it's wrong here
Post-training output filtering does not remove the identifiers from the training corpus, so the model can still memorize and reproduce them through paraphrasing that evades pattern matching. Sensitive data must be removed before training, not masked afterward. This approach also leaves case IDs embedded in model weights, creating an audit and compliance exposure that redaction at the source would have eliminated.
- ✗
Drop every ticket whose resolution contains any digit, since account numbers and case IDs always contain numeric characters, and train only on the remaining text-only resolutions.
Why it's wrong here
Numeric characters appear throughout legitimate technical content, including version numbers, port numbers, error codes, and command flags, so this rule would discard a large share of valid, informative answers. It is also ineffective because names are alphabetic and would survive the filter. Targeted PII detection with entity-aware replacement is required, not a blanket digit exclusion that damages data quality.
Visual reference
About these practice questions
This NCP-GENL question is part of Courseiva's 352-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official NVIDIA exam blueprint
This NCP-GENL practice question is part of Courseiva's free NVIDIA certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the NCP-GENL exam.