Courseiva
Data Preparation →easyMultiple Choice

NCP-GENL Data Preparation Practice Question

A team is preparing a customer-support chat dataset for instruction fine-tuning of an LLM. The raw data contains HTML tags, inconsistent date formats, and emoji. Which data preparation step should be performed first to make the text usable for tokenization?

⚠ Common exam trap

The trap here is treating tokenization as a cleaning step, when tokenizers faithfully encode whatever noise is present in the input text.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Normalize and clean the text

Cleaning and normalizing the raw chat data first removes HTML, standardizes dates, and handles emoji, producing consistent text for tokenization. This order prevents noisy artifacts from becoming tokens and ensures both training and validation splits receive the same treatment. It is the foundational step before tokenization and dataset splitting.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✓

    Normalize and clean the text

    Why this is correct

    Normalization and cleaning remove HTML tags, standardize date formats, and handle emoji so the text is consistent before tokenization. This step ensures the tokenizer sees clean input, reducing noise in the training data and improving the quality of the instruction-tuning examples. It is the logical first step in the preparation pipeline.

  • ✗

    Tokenize the text with the model's tokenizer

    Why it's wrong here

    Tokenization should happen after cleaning because HTML tags and inconsistent formatting would be converted into tokens that pollute the training signal. Tokenizing first also makes later cleaning harder, since you would need to detokenize or manipulate token IDs. The scenario asks for the step before tokenization, so this is premature.

  • ✗

    Convert the text to lowercase

    Why it's wrong here

    Lowercasing is a specific normalization choice, not the overarching first step. It can be harmful for named entities and acronyms in customer-support data. The scenario requires handling HTML, dates, and emoji, which lowercasing alone does not address, so it is too narrow to be the correct first action.

  • ✗

    Split the dataset into train and validation sets

    Why it's wrong here

    Splitting is important but should occur after cleaning and formatting, because you want both splits to benefit from the same cleaning rules and to avoid leakage of raw artifacts. Performing the split first would require cleaning each split separately, risking inconsistent preprocessing and making the pipeline harder to maintain.

About these practice questions

One of 352 original NCP-GENL practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official NVIDIA exam blueprint

This NCP-GENL practice question is part of Courseiva's free NVIDIA certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the NCP-GENL exam.