Courseiva

NCP-GENL · domain

Data Preparation

This domain covers the pipeline that turns raw, messy source data into training and retrieval corpora for LLMs on NVIDIA platforms. Questions are scenario-based: you must choose the correct preprocessing, splitting, tokenization, and safety-data technique, and explain why it matters for fine-tuning, RAG, and deployment on NVIDIA hardware.

41 questions6 easy22 medium13 hard

Focused practice

Practice Data Preparation questions

Scored sessions drawing only from this domain — pick a length below.

Start 20-question practice test →

What this domain covers

What to know about Data Preparation

Be able to select and justify preprocessing, splitting, chunking, tokenization, and safety-data choices for LLM training and RAG on NVIDIA stacks. The single most important thing: prevent leakage and preserve context so the model is grounded, stable, and safe.

Time-aware or chronological splitting to prevent leakage in time-series document corpora

Chunking, normalization, and metadata tagging of proprietary manuals for RAG grounding

Tokenization stability and vocabulary consistency across NVIDIA TensorRT-LLM and NeMo deployments

Adversarial and safety-protocol examples in fine-tuning datasets to harden model behavior

Watch out for

Common Data Preparation exam traps

  • ▸Randomly shuffling time-ordered documents before splitting, which leaks future information into training and inflates evaluation scores
  • ▸Chunking technical manuals without preserving product-configuration context, so retrieval returns ambiguous or conflicting specifications
  • ▸Assuming any tokenizer works at inference, ignoring vocabulary drift that breaks token IDs and degrades NVIDIA deployment consistency

Question index

All Data Preparation questions (41)

Click any question to see the full explanation, or start a practice session above.

1

You are curating a 200 GB instruction-tuning corpus for an NVIDIA NeMo fine-tuning job on a Llama-based model. Post-training evaluation reveals the model regurgitates exact validation-set passages verbatim. An audit shows that near-duplicate instruction/response pairs were split randomly at the record level across train and validation partitions. Which data preparation change most directly eliminates this leakage while preserving the maximum amount of usable training data?

Medium
2

When preparing a proprietary technical manual dataset for a RAG pipeline, which data preprocessing step is most critical to ensure the LLM avoids hallucinations regarding specific product configurations?

Medium
3

You are building the data preparation stage for an NVIDIA NeMo retrieval-augmented generation pipeline that will ingest millions of internal wiki pages. The ingestion team reports that the same policy text appears in dozens of pages with minor edits, and that some pages contain copied tables from external sources. You need to produce a clean, deduplicated chunk store that supports accurate citation and avoids returning redundant passages. Which TWO actions best address these requirements? (Choose two.)

Hard
4

Refer to the exhibit. A team is preparing log data for a RAG-based troubleshooting assistant. Given the configuration, what is the most significant risk during the retrieval phase?

Hard
5

When preparing unstructured documentation for a high-performance retrieval system, which approach best balances index size and retrieval relevance?

Medium
6

You are preparing a dataset for continued pre-training of an LLM on internal engineering documents. The corpus contains many documents with boilerplate headers, footers, and legal disclaimers repeated across files. Which NeMo Curator approach best reduces this boilerplate while preserving unique technical content?

Hard
7

You are preparing a customer-support dataset for fine-tuning an LLM with NVIDIA NeMo. The raw data includes personally identifiable information such as names, email addresses, and phone numbers. Which data preparation step must be performed before training to comply with privacy requirements?

Easy
8

A data engineer is preparing a JSONL instruction dataset for an NVIDIA NeMo supervised fine-tuning run. Each line currently contains a free-form 'text' field with the instruction, context, and response concatenated. The training configuration expects the standard NeMo instruction-tuning schema with separate fields for the task instruction, optional context, and the expected response. What is the most appropriate data preparation step?

Easy
9

You are preparing a 2 TB corpus of English and German web text for continued pretraining of a NeMo-based LLM. The German portion includes many pages with unescaped HTML entities and mixed-language sentences. Which NeMo Curator stage should you apply to remove boilerplate, fix HTML artifacts, and filter low-quality documents before tokenization?

Medium
10

What is the primary function of data 'normalization' in the context of preparing inputs for a Transformer model?

Easy
11

A team is building an instruction-tuning dataset in NeMo from 40,000 internal support tickets. Each ticket contains a customer problem and a resolved answer, but the resolution text sometimes includes the customer's name, account number, and internal case IDs. The team plans to use NeMo Curator to produce training-ready JSONL. Which approach best prepares this data for instruction tuning while limiting personally identifiable information exposure?

Hard
12

You are curating a 2 TB corpus of NVIDIA technical documentation and Python code for continued pretraining of a NeMo-based LLM. A colleague proposes filtering out any document containing the token sequence 'CUDA' to reduce hardware-specific bias. What is the most appropriate response?

Medium
13

When training a model for a highly technical domain with a scarcity of high-quality data, which data augmentation strategy is most likely to preserve the model's reliability?

Hard
14

Which technique is most effective for mitigating data leakage during the training of an LLM on time-series-related document data?

Easy
15

You are preparing a large instruction-tuning dataset for an NVIDIA NeMo-based LLM. The raw data consists of user queries and assistant responses collected from a customer support system, stored as JSON lines. During preprocessing, you notice that many responses contain personally identifiable information (PII) such as names, email addresses, and phone numbers. You need to ensure the dataset is safe for training while preserving as much semantic content as possible. Which approach is most appropriate for handling PII in this dataset?

Medium
16

When assessing the quality of a dataset for instruction fine-tuning, which TWO metrics or methods are considered most reliable for measuring dataset diversity?

Medium
17

You are preparing a multilingual corpus for pre-training an LLM. The corpus contains documents in English, Spanish, and German, but the English portion is 80% of the data. You want to avoid the model becoming biased toward English. Which data preparation technique should you apply?

Medium
18

You are using NVIDIA NeMo Curator to filter a 600 GB web-crawl corpus before pre-training. Your team wants to remove exact duplicates and near-duplicates to reduce memorization and speed up training. Which NeMo Curator stage should you apply?

Medium
19

A team is building a NeMo-based LLM pipeline and must tokenize a corpus that mixes English, Japanese, and Python source code. They plan to train a custom tokenizer with NVIDIA NeMo. Which tokenizer configuration best supports all three content types without excessive sequence length?

Hard
20

You are fine-tuning a LLM on a large corpus of technical documentation. You notice the model struggles with domain-specific terminology that frequently appeared in the training data but was inconsistently formatted. Which data preparation technique best addresses this issue?

Medium
21

You are preparing a large instruction-tuning dataset for an LLM using NVIDIA NeMo. You need to ensure the dataset supports efficient training and evaluation. Which TWO steps are essential in the data preparation pipeline? (Choose two.)

Medium
22

You are preparing a dataset for instruction fine-tuning an LLM using NVIDIA NeMo. The dataset contains pairs of instructions and responses, but you notice that some responses are significantly longer than others, and a few are extremely short (e.g., 'Yes' or 'No'). You want to ensure the model learns to generate appropriate-length responses. Which data preparation technique is most effective?

Medium
23

You are preparing a dataset for pretraining an LLM using NVIDIA NeMo's Megatron-LM. The dataset consists of JSONL files where each line contains a 'text' field. You need to convert these files into the binary format required by Megatron for efficient training. Which tool or method should you use to perform this conversion?

Hard
24

An enterprise is preparing a massive multi-terabyte corpus of specialized PDF documents containing technical schematics and dense tables for training a domain-specific LLM using NeMo Curator. During the document extraction pipeline, you notice that tabular data layouts are scrambled into unstructured string tokens, destroying spatial relationships. Which data preparation strategy should you implement first within NeMo Curator to preserve tabular integrity before tokenization?

Medium
25

A team is preparing a customer-support chat dataset for instruction fine-tuning of an LLM. The raw data contains HTML tags, inconsistent date formats, and emoji. Which data preparation step should be performed first to make the text usable for tokenization?

Easy
26

Refer to the exhibit. You are reviewing the configuration file for a data preprocessing pipeline. Why is the 'min_words' filter set to 50 in the context of LLM training?

Medium
27

You are preparing a massive dataset for training a NeMo-based LLM. Which TWO data preprocessing steps are critical to prevent data leakage and ensure model quality?

Hard
28

Why is it important to perform 'domain-specific' data cleaning when preparing a corpus for fine-tuning a medical LLM?

Medium
29

Refer to the exhibit. What is the impact of this filter on the training corpus?

Hard
30

You are building a pretraining dataset from a large collection of source-code repositories for an NVIDIA NeMo LLM. The data includes many files with licenses, generated code, and minified JavaScript. Which NeMo Curator-based approach best improves code data quality before tokenization?

Hard
31

You are curating instruction-tuning data for an NVIDIA NIM-deployed LLM. The raw dataset contains many near-duplicate instruction-response pairs that differ only by punctuation and whitespace. Which data preparation step is most appropriate to remove these before fine-tuning?

Easy
32

You are preparing a large corpus of customer support transcripts for continued pretraining of an NVIDIA NeMo Megatron model. The transcripts contain personally identifiable information such as names, email addresses, and account numbers, and company policy requires that this information be removed before training while preserving as much linguistic context as possible for the model to learn from. Which data preparation approach best satisfies both requirements?

Medium
33

Refer to the exhibit. What is the intended outcome of this data cleaning configuration for a Large Language Model pre-training corpus?

Hard
34

Why is 'tokenization stability' a critical metric when preparing data for NVIDIA-based LLM deployment?

Medium
35

You are preparing a dataset of customer reviews for fine-tuning an LLM to generate concise summaries. The reviews are in multiple languages, but the target summaries must be in English. You have a limited budget for translation. Which data preparation step is most critical to ensure the fine-tuned model produces high-quality English summaries?

Medium
36

When fine-tuning an LLM to follow specific safety protocols, why is the inclusion of 'adversarial' examples in the training data considered a best practice?

Hard
37

You are preparing a dataset for supervised fine-tuning (SFT) of an NVIDIA NeMo LLM to follow instructions. Which TWO data preparation practices are essential to ensure the model learns to generalize rather than memorize? (Choose two.)

Medium
38

You are preparing a large instruction-tuning dataset with NVIDIA NeMo Curator. The dataset contains many near-duplicate instruction-response pairs that differ only in punctuation or minor wording. Which NeMo Curator stage should you apply to remove these near-duplicates before fine-tuning?

Medium
39

You are preparing a multilingual corpus for pretraining with NVIDIA NeMo. The dataset contains documents in 40 languages, but the tokenizer was trained primarily on English. Which data preparation action best ensures that non-English text is represented efficiently during tokenization?

Medium
40

You are preparing a dataset of support tickets for a RAG system using NVIDIA NeMo. Many tickets are short and contain little context, which hurts retrieval quality. Which data preparation technique best improves retrieval by enriching each ticket with related information before embedding?

Medium
41

A team is preparing a mixed-language corpus for continued pretraining of an NVIDIA NeMo Megatron model. The corpus contains English, Japanese, and Arabic documents. Tokenizer analysis shows the current English-centric BPE vocabulary produces very long token sequences for Japanese and Arabic, inflating sequence length and compute cost. The team wants to reduce sequence length for non-English text without retraining the tokenizer from scratch and without degrading English performance. Which data preparation action best achieves this?

Hard

Frequently asked questions

What does the Data Preparation domain cover on the NCP-GENL exam?
Be able to select and justify preprocessing, splitting, chunking, tokenization, and safety-data choices for LLM training and RAG on NVIDIA stacks. The single most important thing: prevent leakage and preserve context so the model is grounded, stable, and safe.
How many questions are in this domain?
This page lists all 41 Data Preparation questions in the NCP-GENL question bank. The actual exam draws from this domain proportionally to its weighting in the official exam blueprint.
What is the best way to practise this domain?
Start with a short focused session (10 questions) to identify gaps, then work through explanations. Repeat with a longer session once the weak areas feel solid.
Can I practise only Data Preparation questions?
Yes — the session launcher on this page filters questions to this domain only. Choose any session length for inline explanations and scoring.
nvidia-ncp-genl NVIDIA-NCP-GENL ncp data preparation Practice Questions