NCP-GENL Data Preparation Practice Question
A team is preparing a customer-support chat dataset for instruction fine-tuning of an LLM. The raw data contains HTML tags, inconsistent date formats, and emoji. Which data preparation step should be performed first to make the text usable for tokenization?
⚠ Common exam trap
The trap here is treating tokenization as a cleaning step, when tokenizers faithfully encode whatever noise is present in the input text.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Normalize and clean the text
Cleaning and normalizing the raw chat data first removes HTML, standardizes dates, and handles emoji, producing consistent text for tokenization. This order prevents noisy artifacts from becoming tokens and ensures both training and validation splits receive the same treatment. It is the foundational step before tokenization and dataset splitting.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✓
Normalize and clean the text
Why this is correct
Normalization and cleaning remove HTML tags, standardize date formats, and handle emoji so the text is consistent before tokenization. This step ensures the tokenizer sees clean input, reducing noise in the training data and improving the quality of the instruction-tuning examples. It is the logical first step in the preparation pipeline.
- ✗
Tokenize the text with the model's tokenizer
Why it's wrong here
Tokenization should happen after cleaning because HTML tags and inconsistent formatting would be converted into tokens that pollute the training signal. Tokenizing first also makes later cleaning harder, since you would need to detokenize or manipulate token IDs. The scenario asks for the step before tokenization, so this is premature.
- ✗
Convert the text to lowercase
Why it's wrong here
Lowercasing is a specific normalization choice, not the overarching first step. It can be harmful for named entities and acronyms in customer-support data. The scenario requires handling HTML, dates, and emoji, which lowercasing alone does not address, so it is too narrow to be the correct first action.
- ✗
Split the dataset into train and validation sets
Why it's wrong here
Splitting is important but should occur after cleaning and formatting, because you want both splits to benefit from the same cleaning rules and to avoid leakage of raw artifacts. Performing the split first would require cleaning each split separately, risking inconsistent preprocessing and making the pipeline harder to maintain.
About these practice questions
One of 352 original NCP-GENL practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official NVIDIA exam blueprint
This NCP-GENL practice question is part of Courseiva's free NVIDIA certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the NCP-GENL exam.