NCP-GENL · domain
Data Preparation
This domain covers the pipeline that turns raw, messy source data into training and retrieval corpora for LLMs on NVIDIA platforms. Questions are scenario-based: you must choose the correct preprocessing, splitting, tokenization, and safety-data technique, and explain why it matters for fine-tuning, RAG, and deployment on NVIDIA hardware.
Focused practice
Practice Data Preparation questions
Scored sessions drawing only from this domain — pick a length below.
Start 20-question practice test →What this domain covers
What to know about Data Preparation
Be able to select and justify preprocessing, splitting, chunking, tokenization, and safety-data choices for LLM training and RAG on NVIDIA stacks. The single most important thing: prevent leakage and preserve context so the model is grounded, stable, and safe.
Time-aware or chronological splitting to prevent leakage in time-series document corpora
Chunking, normalization, and metadata tagging of proprietary manuals for RAG grounding
Tokenization stability and vocabulary consistency across NVIDIA TensorRT-LLM and NeMo deployments
Adversarial and safety-protocol examples in fine-tuning datasets to harden model behavior
Watch out for
Common Data Preparation exam traps
- ▸Randomly shuffling time-ordered documents before splitting, which leaks future information into training and inflates evaluation scores
- ▸Chunking technical manuals without preserving product-configuration context, so retrieval returns ambiguous or conflicting specifications
- ▸Assuming any tokenizer works at inference, ignoring vocabulary drift that breaks token IDs and degrades NVIDIA deployment consistency
Question index
All Data Preparation questions (41)
Click any question to see the full explanation, or start a practice session above.
You are curating a 200 GB instruction-tuning corpus for an NVIDIA NeMo fine-tuning job on a Llama-based model. Post-training evaluation reveals the model regurgitates exact validation-set passages verbatim. An audit shows that near-duplicate instruction/response pairs were split randomly at the record level across train and validation partitions. Which data preparation change most directly eliminates this leakage while preserving the maximum amount of usable training data?
Medium2When preparing a proprietary technical manual dataset for a RAG pipeline, which data preprocessing step is most critical to ensure the LLM avoids hallucinations regarding specific product configurations?
Medium3You are building the data preparation stage for an NVIDIA NeMo retrieval-augmented generation pipeline that will ingest millions of internal wiki pages. The ingestion team reports that the same policy text appears in dozens of pages with minor edits, and that some pages contain copied tables from external sources. You need to produce a clean, deduplicated chunk store that supports accurate citation and avoids returning redundant passages. Which TWO actions best address these requirements? (Choose two.)
Hard4Refer to the exhibit. A team is preparing log data for a RAG-based troubleshooting assistant. Given the configuration, what is the most significant risk during the retrieval phase?
Hard5When preparing unstructured documentation for a high-performance retrieval system, which approach best balances index size and retrieval relevance?
Medium6You are preparing a dataset for continued pre-training of an LLM on internal engineering documents. The corpus contains many documents with boilerplate headers, footers, and legal disclaimers repeated across files. Which NeMo Curator approach best reduces this boilerplate while preserving unique technical content?
Hard7You are preparing a customer-support dataset for fine-tuning an LLM with NVIDIA NeMo. The raw data includes personally identifiable information such as names, email addresses, and phone numbers. Which data preparation step must be performed before training to comply with privacy requirements?
Easy8A data engineer is preparing a JSONL instruction dataset for an NVIDIA NeMo supervised fine-tuning run. Each line currently contains a free-form 'text' field with the instruction, context, and response concatenated. The training configuration expects the standard NeMo instruction-tuning schema with separate fields for the task instruction, optional context, and the expected response. What is the most appropriate data preparation step?
Easy9You are preparing a 2 TB corpus of English and German web text for continued pretraining of a NeMo-based LLM. The German portion includes many pages with unescaped HTML entities and mixed-language sentences. Which NeMo Curator stage should you apply to remove boilerplate, fix HTML artifacts, and filter low-quality documents before tokenization?
Medium10What is the primary function of data 'normalization' in the context of preparing inputs for a Transformer model?
Easy11A team is building an instruction-tuning dataset in NeMo from 40,000 internal support tickets. Each ticket contains a customer problem and a resolved answer, but the resolution text sometimes includes the customer's name, account number, and internal case IDs. The team plans to use NeMo Curator to produce training-ready JSONL. Which approach best prepares this data for instruction tuning while limiting personally identifiable information exposure?
Hard12You are curating a 2 TB corpus of NVIDIA technical documentation and Python code for continued pretraining of a NeMo-based LLM. A colleague proposes filtering out any document containing the token sequence 'CUDA' to reduce hardware-specific bias. What is the most appropriate response?
Medium13When training a model for a highly technical domain with a scarcity of high-quality data, which data augmentation strategy is most likely to preserve the model's reliability?
Hard14Which technique is most effective for mitigating data leakage during the training of an LLM on time-series-related document data?
Easy15You are preparing a large instruction-tuning dataset for an NVIDIA NeMo-based LLM. The raw data consists of user queries and assistant responses collected from a customer support system, stored as JSON lines. During preprocessing, you notice that many responses contain personally identifiable information (PII) such as names, email addresses, and phone numbers. You need to ensure the dataset is safe for training while preserving as much semantic content as possible. Which approach is most appropriate for handling PII in this dataset?
Medium16When assessing the quality of a dataset for instruction fine-tuning, which TWO metrics or methods are considered most reliable for measuring dataset diversity?
Medium17You are preparing a multilingual corpus for pre-training an LLM. The corpus contains documents in English, Spanish, and German, but the English portion is 80% of the data. You want to avoid the model becoming biased toward English. Which data preparation technique should you apply?
Medium18You are using NVIDIA NeMo Curator to filter a 600 GB web-crawl corpus before pre-training. Your team wants to remove exact duplicates and near-duplicates to reduce memorization and speed up training. Which NeMo Curator stage should you apply?
Medium19A team is building a NeMo-based LLM pipeline and must tokenize a corpus that mixes English, Japanese, and Python source code. They plan to train a custom tokenizer with NVIDIA NeMo. Which tokenizer configuration best supports all three content types without excessive sequence length?
Hard20You are fine-tuning a LLM on a large corpus of technical documentation. You notice the model struggles with domain-specific terminology that frequently appeared in the training data but was inconsistently formatted. Which data preparation technique best addresses this issue?
Medium21You are preparing a large instruction-tuning dataset for an LLM using NVIDIA NeMo. You need to ensure the dataset supports efficient training and evaluation. Which TWO steps are essential in the data preparation pipeline? (Choose two.)
Medium22You are preparing a dataset for instruction fine-tuning an LLM using NVIDIA NeMo. The dataset contains pairs of instructions and responses, but you notice that some responses are significantly longer than others, and a few are extremely short (e.g., 'Yes' or 'No'). You want to ensure the model learns to generate appropriate-length responses. Which data preparation technique is most effective?
Medium23You are preparing a dataset for pretraining an LLM using NVIDIA NeMo's Megatron-LM. The dataset consists of JSONL files where each line contains a 'text' field. You need to convert these files into the binary format required by Megatron for efficient training. Which tool or method should you use to perform this conversion?
Hard24An enterprise is preparing a massive multi-terabyte corpus of specialized PDF documents containing technical schematics and dense tables for training a domain-specific LLM using NeMo Curator. During the document extraction pipeline, you notice that tabular data layouts are scrambled into unstructured string tokens, destroying spatial relationships. Which data preparation strategy should you implement first within NeMo Curator to preserve tabular integrity before tokenization?
Medium25A team is preparing a customer-support chat dataset for instruction fine-tuning of an LLM. The raw data contains HTML tags, inconsistent date formats, and emoji. Which data preparation step should be performed first to make the text usable for tokenization?
Easy26Refer to the exhibit. You are reviewing the configuration file for a data preprocessing pipeline. Why is the 'min_words' filter set to 50 in the context of LLM training?
Medium27You are preparing a massive dataset for training a NeMo-based LLM. Which TWO data preprocessing steps are critical to prevent data leakage and ensure model quality?
Hard28Why is it important to perform 'domain-specific' data cleaning when preparing a corpus for fine-tuning a medical LLM?
Medium29Refer to the exhibit. What is the impact of this filter on the training corpus?
Hard30You are building a pretraining dataset from a large collection of source-code repositories for an NVIDIA NeMo LLM. The data includes many files with licenses, generated code, and minified JavaScript. Which NeMo Curator-based approach best improves code data quality before tokenization?
Hard31You are curating instruction-tuning data for an NVIDIA NIM-deployed LLM. The raw dataset contains many near-duplicate instruction-response pairs that differ only by punctuation and whitespace. Which data preparation step is most appropriate to remove these before fine-tuning?
Easy32You are preparing a large corpus of customer support transcripts for continued pretraining of an NVIDIA NeMo Megatron model. The transcripts contain personally identifiable information such as names, email addresses, and account numbers, and company policy requires that this information be removed before training while preserving as much linguistic context as possible for the model to learn from. Which data preparation approach best satisfies both requirements?
Medium33Refer to the exhibit. What is the intended outcome of this data cleaning configuration for a Large Language Model pre-training corpus?
Hard34Why is 'tokenization stability' a critical metric when preparing data for NVIDIA-based LLM deployment?
Medium35You are preparing a dataset of customer reviews for fine-tuning an LLM to generate concise summaries. The reviews are in multiple languages, but the target summaries must be in English. You have a limited budget for translation. Which data preparation step is most critical to ensure the fine-tuned model produces high-quality English summaries?
Medium36When fine-tuning an LLM to follow specific safety protocols, why is the inclusion of 'adversarial' examples in the training data considered a best practice?
Hard37You are preparing a dataset for supervised fine-tuning (SFT) of an NVIDIA NeMo LLM to follow instructions. Which TWO data preparation practices are essential to ensure the model learns to generalize rather than memorize? (Choose two.)
Medium38You are preparing a large instruction-tuning dataset with NVIDIA NeMo Curator. The dataset contains many near-duplicate instruction-response pairs that differ only in punctuation or minor wording. Which NeMo Curator stage should you apply to remove these near-duplicates before fine-tuning?
Medium39You are preparing a multilingual corpus for pretraining with NVIDIA NeMo. The dataset contains documents in 40 languages, but the tokenizer was trained primarily on English. Which data preparation action best ensures that non-English text is represented efficiently during tokenization?
Medium40You are preparing a dataset of support tickets for a RAG system using NVIDIA NeMo. Many tickets are short and contain little context, which hurts retrieval quality. Which data preparation technique best improves retrieval by enriching each ticket with related information before embedding?
Medium41A team is preparing a mixed-language corpus for continued pretraining of an NVIDIA NeMo Megatron model. The corpus contains English, Japanese, and Arabic documents. Tokenizer analysis shows the current English-centric BPE vocabulary produces very long token sequences for Japanese and Arabic, inflating sequence length and compute cost. The team wants to reduce sequence length for non-English text without retraining the tokenizer from scratch and without degrading English performance. Which data preparation action best achieves this?
HardOther domains
All NCP-GENL exam domains
Frequently asked questions
- What does the Data Preparation domain cover on the NCP-GENL exam?
- Be able to select and justify preprocessing, splitting, chunking, tokenization, and safety-data choices for LLM training and RAG on NVIDIA stacks. The single most important thing: prevent leakage and preserve context so the model is grounded, stable, and safe.
- How many questions are in this domain?
- This page lists all 41 Data Preparation questions in the NCP-GENL question bank. The actual exam draws from this domain proportionally to its weighting in the official exam blueprint.
- What is the best way to practise this domain?
- Start with a short focused session (10 questions) to identify gaps, then work through explanations. Repeat with a longer session once the weak areas feel solid.
- Can I practise only Data Preparation questions?
- Yes — the session launcher on this page filters questions to this domain only. Choose any session length for inline explanations and scoring.