Courseiva
Data Preparation →hardMultiple Choice

NCP-GENL Data Preparation Practice Question

A team is building an instruction-tuning dataset in NeMo from 40,000 internal support tickets. Each ticket contains a customer problem and a resolved answer, but the resolution text sometimes includes the customer's name, account number, and internal case IDs. The team plans to use NeMo Curator to produce training-ready JSONL. Which approach best prepares this data for instruction tuning while limiting personally identifiable information exposure?

⚠ Common exam trap

The trap here is treating PII handling as a post-training output filter, when the identifiers must be removed from the training corpus itself before the model ever sees them.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Run a PII redaction stage that detects and replaces names, account numbers, and case IDs with placeholders, then convert problem-resolution pairs into the instruction, input, and output schema before export.

Sensitive identifiers must be removed from the corpus before training, not mitigated after the fact. Running a PII detection and replacement stage in NeMo Curator, then mapping problem-resolution pairs into the instruction, input, and output schema, produces compliant JSONL that the fine-tuning job can consume directly. This ordering keeps identifiers out of both the training data and the resulting model weights.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    Hash the entire ticket text with a cryptographic digest and train on the hashes, since the model can learn the mapping between hashed problems and hashed answers without ever seeing the original identifiers.

    Why it's wrong here

    A cryptographic hash is a one-way transformation that destroys the linguistic structure the model must learn. Training on hashes would yield a model incapable of generating coherent language because tokens no longer carry meaning. PII protection requires selective entity replacement that preserves surrounding natural language, not whole-record hashing that eliminates the instructional signal entirely.

  • ✓

    Run a PII redaction stage that detects and replaces names, account numbers, and case IDs with placeholders, then convert problem-resolution pairs into the instruction, input, and output schema before export.

    Why this is correct

    Redacting identifiers before schema conversion prevents the model from memorizing sensitive strings and keeps the pipeline deterministic. NeMo Curator supports PII detection and replacement as a distinct stage, and converting cleaned records into the instruction, input, and output fields yields the JSONL format the fine-tuning configuration expects. Doing redaction first also avoids leaking identifiers into derived fields.

  • ✗

    Convert the tickets into instruction, input, and output JSONL first, then fine-tune the model and rely on a post-training output filter to block any generated response that contains an account number or customer name.

    Why it's wrong here

    Post-training output filtering does not remove the identifiers from the training corpus, so the model can still memorize and reproduce them through paraphrasing that evades pattern matching. Sensitive data must be removed before training, not masked afterward. This approach also leaves case IDs embedded in model weights, creating an audit and compliance exposure that redaction at the source would have eliminated.

  • ✗

    Drop every ticket whose resolution contains any digit, since account numbers and case IDs always contain numeric characters, and train only on the remaining text-only resolutions.

    Why it's wrong here

    Numeric characters appear throughout legitimate technical content, including version numbers, port numbers, error codes, and command flags, so this rule would discard a large share of valid, informative answers. It is also ineffective because names are alphabetic and would survive the filter. Targeted PII detection with entity-aware replacement is required, not a blanket digit exclusion that damages data quality.

Visual reference

Client Recursive Resolver Root DNS (13 root servers) TLD DNS (.com, .org, …) Authoritative example.com query IP addr answer

About these practice questions

This NCP-GENL question is part of Courseiva's 352-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official NVIDIA exam blueprint

This NCP-GENL practice question is part of Courseiva's free NVIDIA certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the NCP-GENL exam.