Courseiva
Data Preparation →easyMultiple Choice

NCP-GENL Data Preparation Practice Question

You are preparing a customer-support dataset for fine-tuning an LLM with NVIDIA NeMo. The raw data includes personally identifiable information such as names, email addresses, and phone numbers. Which data preparation step must be performed before training to comply with privacy requirements?

⚠ Common exam trap

It's easy for candidates to confuse training-time hyperparameters like batch size or shuffling with data-preparation steps that actually remove sensitive content.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Apply PII detection and redaction to replace sensitive entities with placeholders.

Privacy compliance requires removing or masking personally identifiable information before training. PII detection and redaction replaces sensitive entities with placeholders, preserving the dataset's instructional value while preventing the model from memorizing real customer details. Tokenization, batch size, and shuffling do not alter the presence of PII and therefore cannot satisfy the requirement to protect sensitive data during fine-tuning.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✓

    Apply PII detection and redaction to replace sensitive entities with placeholders.

    Why this is correct

    Detecting and redacting PII replaces names, emails, and phone numbers with generic placeholders, removing sensitive information while preserving the conversational structure needed for fine-tuning. This directly addresses the privacy requirement and prevents the model from memorizing or emitting real customer data, making it the correct preparation step.

  • ✗

    Shuffle the dataset to break associations between PII and responses.

    Why it's wrong here

    Shuffling changes the order of examples but leaves the content intact, so PII is still present and learnable. The association between a customer's details and a response may still exist within each record. Shuffling is useful for training stability, not for privacy, so it does not satisfy the scenario's compliance obligation.

  • ✗

    Increase the batch size during fine-tuning to average out PII exposure.

    Why it's wrong here

    Batch size affects optimization dynamics, not data privacy. PII remains in the training records and can still be memorized and reproduced regardless of how many examples are processed together. This option does not remove or mask sensitive information and therefore fails to meet the compliance requirement in the scenario.

  • ✗

    Tokenize the dataset with a custom vocabulary that includes PII patterns.

    Why it's wrong here

    Tokenizing PII does not remove it; the model can still learn and potentially reproduce the sensitive strings. Adding PII patterns to the vocabulary makes them easier to generate, worsening the privacy risk. This step fails to anonymize the data and does not satisfy compliance requirements, so it is the wrong choice for the scenario.

About these practice questions

One of 352 original NCP-GENL practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official NVIDIA exam blueprint

This NCP-GENL practice question is part of Courseiva's free NVIDIA certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the NCP-GENL exam.