NCA-GENL Core Machine Learning and AI Knowledge Practice Question
A data scientist is preparing a dataset of 50,000 customer support chat transcripts to fine-tune an LLM for a helpdesk assistant. The raw text contains HTML tags, inconsistent whitespace, and occasional personal information such as email addresses. Which preprocessing step should be performed FIRST to prepare the text for tokenization?
⚠ Common exam trap
The trap here is assuming tokenization should happen first because it is the most familiar LLM step, when in fact noisy raw text must be cleaned before tokenization.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Clean the text by stripping HTML, normalizing whitespace, and masking personally identifiable information.
Raw chat transcripts often contain markup and private data that degrade fine-tuning quality and raise privacy concerns. Cleaning the text by removing HTML, normalizing whitespace, and masking emails or phone numbers before tokenization ensures the tokenizer produces meaningful tokens and that sensitive data never reaches the model. This ordering also avoids retokenization work later.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Apply tokenization using the model's tokenizer.
Why it's wrong here
Tokenization should come after cleaning because HTML tags and stray characters would be split into meaningless tokens, inflating sequence length and wasting model capacity. If tokenized first, removing HTML later would require detokenizing or retokenizing, which is error-prone. Cleaning the raw string before tokenization ensures the tokenizer sees only the intended linguistic content.
- ✗
Convert all text to lowercase and remove all punctuation.
Why it's wrong here
Lowercasing and removing punctuation is a normalization choice, not a universal first step. It can destroy useful signals like casing in product names or sentence boundaries. More importantly, it does not address HTML tags or personal information, so it leaves the primary noise and privacy issues unresolved. Cleaning should target the actual contaminants in this dataset.
- ✗
Split the transcripts into training and validation sets.
Why it's wrong here
Splitting into train and validation sets is a later step that should happen after cleaning and tokenization, or at least after cleaning. If you split before cleaning, you risk leaking cleaning statistics or applying inconsistent transformations. The immediate concern is removing HTML and masking personal data from the raw text before any model-specific processing.
- ✓
Clean the text by stripping HTML, normalizing whitespace, and masking personally identifiable information.
Why this is correct
Cleaning raw text before tokenization removes noise that would otherwise become tokens and helps protect sensitive data. Stripping HTML and normalizing whitespace standardizes the input, while masking emails prevents the model from memorizing private information. This step must precede tokenization so the tokenizer operates on the final, intended character sequence.
About these practice questions
This NCA-GENL question is part of Courseiva's 367-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official NVIDIA exam blueprint
This NCA-GENL practice question is part of Courseiva's free NVIDIA certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the NCA-GENL exam.