Courseiva

NCP-GENL · topic practice

Data Preparation practice questions

This domain covers the pipeline that turns raw, messy source data into training and retrieval corpora for LLMs on NVIDIA platforms. Questions are scenario-based: you must choose the correct preprocessing, splitting, tokenization, and safety-data technique, and explain why it matters for fine-tuning, RAG, and deployment on NVIDIA hardware.

Courseiva uses original exam-style practice questions designed for learning and revision. The goal is to understand the concepts, recognise exam patterns, and improve through explanations — not memorise copied exam dumps.

Editorial oversight:Johnson Ajibi· MSc IT Security, IEEE Senior Member
20 questionsDomain: Data Preparation

What the exam tests

What to know about Data Preparation

Be able to select and justify preprocessing, splitting, chunking, tokenization, and safety-data choices for LLM training and RAG on NVIDIA stacks. The single most important thing: prevent leakage and preserve context so the model is grounded, stable, and safe.

Time-aware or chronological splitting to prevent leakage in time-series document corpora

Chunking, normalization, and metadata tagging of proprietary manuals for RAG grounding

Tokenization stability and vocabulary consistency across NVIDIA TensorRT-LLM and NeMo deployments

Adversarial and safety-protocol examples in fine-tuning datasets to harden model behavior

Watch out for

Common Data Preparation exam traps

  • ▸Randomly shuffling time-ordered documents before splitting, which leaks future information into training and inflates evaluation scores
  • ▸Chunking technical manuals without preserving product-configuration context, so retrieval returns ambiguous or conflicting specifications
  • ▸Assuming any tokenizer works at inference, ignoring vocabulary drift that breaks token IDs and degrades NVIDIA deployment consistency

Practice set

Data Preparation questions

20 questions · select your answer, then reveal the explanation

When curating a high-quality instruction tuning dataset for fine-tuning a Large Language Model, which TWO factors are essential for maintaining model performance and safety?

In the context of NVIDIA NeMo, which THREE steps are foundational to the data preparation pipeline for pre-training large language models?

When preparing datasets for Retrieval-Augmented Generation (RAG), which THREE factors are essential to ensure efficient retrieval on an NVIDIA-accelerated vector database?

You are curating a pretraining corpus with NVIDIA NeMo Curator and need to remove low-quality documents before tokenization. Which two NeMo Curator heuristic filtering criteria are appropriate for this goal? (Choose two.)

Question 5mediummultiple choice
Read the full Data Preparation explanation →

When preparing a proprietary technical manual dataset for a RAG pipeline, which data preprocessing step is most critical to ensure the LLM avoids hallucinations regarding specific product configurations?

Refer to the exhibit. A team is preparing log data for a RAG-based troubleshooting assistant. Given the configuration, what is the most significant risk during the retrieval phase?

Exhibit

{
  "data_source": "logs_v1",
  "privacy_filter": "regex_mask_all",
  "embedding_model": "nv-embed-v1",
  "chunk_size": 4096,
  "overlap": 0
}

Which technique is most effective for mitigating data leakage during the training of an LLM on time-series-related document data?

Refer to the exhibit. What is the intended outcome of this data cleaning configuration for a Large Language Model pre-training corpus?

Exhibit

config.yaml:
  filter_policy:
    min_words: 50
    max_perplexity: 100
  deduplication:
    algorithm: minhash
    threshold: 0.95
Question 9mediummultiple choice
Read the full Data Preparation explanation →

Why is 'tokenization stability' a critical metric when preparing data for NVIDIA-based LLM deployment?

Question 10mediummultiple choice
Read the full Data Preparation explanation →

When preparing unstructured documentation for a high-performance retrieval system, which approach best balances index size and retrieval relevance?

When fine-tuning an LLM to follow specific safety protocols, why is the inclusion of 'adversarial' examples in the training data considered a best practice?

When assessing the quality of a dataset for instruction fine-tuning, which TWO metrics or methods are considered most reliable for measuring dataset diversity?

What is the primary function of data 'normalization' in the context of preparing inputs for a Transformer model?

Question 14mediummultiple choice
Read the full Data Preparation explanation →

Why is it important to perform 'domain-specific' data cleaning when preparing a corpus for fine-tuning a medical LLM?

Refer to the exhibit. What is the impact of this filter on the training corpus?

Exhibit

{
  "action": "filter",
  "criterion": "min_lexical_diversity",
  "threshold": 0.2
}

When training a model for a highly technical domain with a scarcity of high-quality data, which data augmentation strategy is most likely to preserve the model's reliability?

Question 17mediummultiple choice
Read the full Data Preparation explanation →

You are fine-tuning a LLM on a large corpus of technical documentation. You notice the model struggles with domain-specific terminology that frequently appeared in the training data but was inconsistently formatted. Which data preparation technique best addresses this issue?

You are preparing a massive dataset for training a NeMo-based LLM. Which TWO data preprocessing steps are critical to prevent data leakage and ensure model quality?

Question 19mediummultiple choice
Read the full Data Preparation explanation →

Refer to the exhibit. You are reviewing the configuration file for a data preprocessing pipeline. Why is the 'min_words' filter set to 50 in the context of LLM training?

Exhibit

{
  "dataset_config": {
    "path": "/data/corpus/",
    "format": "jsonl",
    "filters": {
      "min_words": 50,
      "language": "en",
      "pii_redaction": true
    },
    "tokenization": {
      "vocab_size": 32000,
      "special_tokens": ["<pad>", "<bos>", "<eos>"]
    }
  }
}
Question 20mediummultiple choice
Read the full Data Preparation explanation →

An enterprise is preparing a massive multi-terabyte corpus of specialized PDF documents containing technical schematics and dense tables for training a domain-specific LLM using NeMo Curator. During the document extraction pipeline, you notice that tabular data layouts are scrambled into unstructured string tokens, destroying spatial relationships. Which data preparation strategy should you implement first within NeMo Curator to preserve tabular integrity before tokenization?

Free account

Track your progress over time

Create a free account to save your results and see which topics improve across sessions.

Focused Data Preparation sessions

Start a Data Preparation only practice session

Every question in these sessions is drawn from the Data Preparation domain — nothing else.

Related practice questions

Related NCP-GENL topic practice pages

Move into related areas when this topic feels solid.

Frequently asked questions

What does the NCP-GENL exam test about Data Preparation?
Be able to select and justify preprocessing, splitting, chunking, tokenization, and safety-data choices for LLM training and RAG on NVIDIA stacks. The single most important thing: prevent leakage and preserve context so the model is grounded, stable, and safe.
How should I use these practice questions?
Select your answer before revealing the explanation. Then read why each option is right or wrong — this active recall approach builds retention far faster than re-reading notes.
Can I practise just Data Preparation questions in a focused session?
Yes — the session launcher on this page draws every question from the Data Preparation domain. Use a 10-question session first to gauge your baseline, then move to 20 or 30 once the weak spots are clear.
Where can I practise other NCP-GENL topics?
Use the topic links above to move to related areas, or go back to the NCP-GENL question bank to see all topics.
Are these real exam questions or dumps?
These are original practice questions written to test the same concepts the NCP-GENL exam covers. They are not copied from any real exam or dump site.