Courseiva

CCNA Data Preparation Questions

41 questions · Data Preparation · All types, answers revealed

1
MCQmedium

You are curating a 200 GB instruction-tuning corpus for an NVIDIA NeMo fine-tuning job on a Llama-based model. Post-training evaluation reveals the model regurgitates exact validation-set passages verbatim. An audit shows that near-duplicate instruction/response pairs were split randomly at the record level across train and validation partitions. Which data preparation change most directly eliminates this leakage while preserving the maximum amount of usable training data?

A.Increase the validation split ratio from 5% to 20% so that fewer training records overlap with the validation partition.
B.Perform MinHash-based near-duplicate detection across the full corpus, cluster records by similarity, and assign whole clusters to either the train or validation split using a deterministic hash of the cluster ID.
C.Apply aggressive token-level cleaning to strip boilerplate phrases and punctuation from every record before splitting the dataset randomly.
D.Deduplicate only the validation partition by removing any record whose exact text appears in the training partition, then keep the original random split.
AnswerB

Grouping near-duplicates into clusters and assigning entire clusters to one partition prevents the same or paraphrased content from appearing in both train and validation, which is exactly what caused the verbatim regurgitation. Because only duplicate clusters are collapsed rather than all similar-looking records, the maximum amount of unique training data is retained, and the deterministic hash keeps splits reproducible across reruns.

Why this answer

The leakage arises because near-duplicate instruction/response pairs were split at the record level, allowing paraphrased twins to appear in both partitions. Clustering near-duplicates and assigning entire clusters to one split closes that pathway while retaining all unique content for training. Deterministic cluster-to-split hashing also makes the partition reproducible, which matters for auditing and for comparing fine-tuning runs fairly.

Exam trap

The trap here is assuming that deduplication must be exact or that adjusting the split ratio can fix leakage caused by near-duplicate records spanning partitions.

2
MCQmedium

When preparing a proprietary technical manual dataset for a RAG pipeline, which data preprocessing step is most critical to ensure the LLM avoids hallucinations regarding specific product configurations?

A.Converting all text to lowercase to ensure uniformity in vector embeddings.
B.Performing aggressive stop-word removal to reduce the dimensionality of the vector space.
C.Implementing document-aware recursive character splitting with overlapping segments and metadata tagging.
D.Applying basic sentence tokenization based strictly on periods to create uniform chunks.
AnswerC

Maintaining document hierarchy via metadata and using recursive splitting preserves logical boundaries within technical manuals. The overlap ensures that context isn't lost at chunk edges, while metadata allows the system to filter by product version, ensuring the LLM only consumes data relevant to the specific hardware revision being queried.

Why this answer

Chunking strategies and metadata tagging ensure that context retrieval is precise. By preserving technical hierarchy and associating data with specific product versions, the LLM retrieves ground-truth documentation rather than generic information. This reduces hallucinations by constraining the search space to relevant, version-controlled text blocks, directly impacting the accuracy and reliability of downstream inference tasks in enterprise NVIDIA-based AI deployments.

Exam trap

Candidates frequently choose basic paragraph splitting over recursive character splitting with metadata, which fails to respect document hierarchy and product versioning, leading to hallucinations.

3
Multi-Selecthard

You are building the data preparation stage for an NVIDIA NeMo retrieval-augmented generation pipeline that will ingest millions of internal wiki pages. The ingestion team reports that the same policy text appears in dozens of pages with minor edits, and that some pages contain copied tables from external sources. You need to produce a clean, deduplicated chunk store that supports accurate citation and avoids returning redundant passages. Which TWO actions best address these requirements? (Choose two.)

Select 2 answers
A.Apply a fixed-size token window with no overlap to every page so that chunk boundaries are uniform across the corpus.
B.Run near-duplicate detection at the chunk level using MinHash or SimHash, and retain one canonical chunk per duplicate group while preserving a mapping from the canonical chunk to all source page identifiers.
C.Attach source metadata such as page identifier, section heading, and last-modified timestamp to each chunk, and propagate that metadata through embedding and indexing so retrieved chunks can be cited.
D.Tag chunks that contain tables copied from external sources as low-quality and exclude them from the index entirely.
E.Embed every page as a single vector and store the full page text as the retrieval unit, avoiding chunking entirely.
AnswersB, C

Chunk-level near-duplicate detection removes redundant policy passages that differ only by minor edits, which directly reduces duplicate retrieval results. Keeping a mapping from the canonical chunk back to every source page identifier preserves citation integrity, so a retrieved passage can still be attributed to all originating pages even though only one copy is embedded and indexed.

Why this answer

Redundant retrieval from repeated policy text is best solved by chunk-level near-duplicate detection that keeps one canonical copy while mapping it to all source pages, and accurate citation requires source metadata to travel with each chunk through the embedding and indexing stages. Together these actions reduce duplicate results and preserve provenance, which fixed-size chunking, page-level embedding, or blanket exclusion of external tables cannot achieve.

Exam trap

The trap here is assuming that deduplication alone is sufficient and overlooking that citation integrity depends on metadata being carried alongside the retained canonical chunk.

4
MCQhard

Refer to the exhibit. A team is preparing log data for a RAG-based troubleshooting assistant. Given the configuration, what is the most significant risk during the retrieval phase?

A.The use of 'nv-embed-v1' will cause OOM errors during the indexing phase.
B.The chunk size of 4096 tokens is too small for modern log analysis.
C.Setting the overlap to 0 will break contextual continuity between consecutive log entries.
D.The 'regex_mask_all' policy will cause the embedding model to fail during vectorization.
AnswerC

Logs are inherently sequential. A zero-overlap configuration ensures that each chunk is treated as an isolated entity, potentially splitting related log events. This prevents the LLM from seeing the full narrative of a system failure, significantly reducing the diagnostic utility of the retrieval-augmented generation output for complex errors.

Why this answer

The lack of overlap between chunks is critical. In log data, sequential events are often related; by having zero overlap, the system loses the transition between log lines that might contain a causal link. This fragmentation forces the model to interpret isolated snapshots, hindering its ability to reconstruct the sequence of errors or system states, which is vital for effective root-cause analysis.

Exam trap

Candidates often focus on chunk size or retrieval speed, missing that zero overlap in sequential data like logs destroys the context necessary for the model to understand causal relationships.

5
MCQmedium

When preparing unstructured documentation for a high-performance retrieval system, which approach best balances index size and retrieval relevance?

A.Using extremely large chunks to ensure that every document is a single vector.
B.Implementing sliding window chunking with semantic overlap based on document structure.
C.Storing every sentence as an individual chunk in the vector database.
D.Removing all overlaps to keep the index size at the absolute minimum possible.
AnswerB

Sliding window chunking with semantic overlap allows the retrieval system to maintain context across chunk boundaries. By respecting document structure, the system ensures that chunks are meaningful and logically coherent. This approach provides the best balance between retrieval granularity, context preservation, and overall index size efficiency for high-performance systems.

Why this answer

Utilizing a sliding window approach with semantic overlap ensures that retrieved chunks maintain context. By carefully selecting chunk size and overlap, engineers can optimize the index size to avoid redundant storage while ensuring that the semantic units of the text are not cut off. This balance is vital for maximizing the accuracy of RAG systems running on NVIDIA infrastructure, where memory efficiency is paramount.

Exam trap

Test-takers frequently choose fixed-size character chunking without overlap, mistakenly assuming it preserves context, when semantic overlap is essential to prevent cutting off critical context.

6
MCQhard

You are preparing a dataset for continued pre-training of an LLM on internal engineering documents. The corpus contains many documents with boilerplate headers, footers, and legal disclaimers repeated across files. Which NeMo Curator approach best reduces this boilerplate while preserving unique technical content?

A.Use the n-gram or sentence-level deduplication stage to remove repeated spans
B.Apply a heuristic filter that drops documents below a word-count threshold
C.Lowercase all text and remove punctuation
D.Split documents into fixed-size chunks and keep only the first chunk
AnswerA

NeMo Curator supports span-level deduplication, which can identify and remove repeated n-grams or sentences that appear across many documents. This directly targets boilerplate headers, footers, and disclaimers while leaving unique technical passages intact. It is the most precise way to reduce redundancy without discarding entire documents.

Why this answer

Span-level deduplication in NeMo Curator detects repeated n-grams or sentences across documents and removes them, which precisely targets boilerplate headers, footers, and legal disclaimers. Unlike document-level filtering, it preserves the unique technical content within each file, making it the best choice for a corpus where redundancy is localized rather than whole-document.

Exam trap

The trap here is assuming that document-level deduplication or filtering will remove boilerplate, when the repeated text is only a small portion of otherwise unique documents.

7
MCQeasy

You are preparing a customer-support dataset for fine-tuning an LLM with NVIDIA NeMo. The raw data includes personally identifiable information such as names, email addresses, and phone numbers. Which data preparation step must be performed before training to comply with privacy requirements?

A.Apply PII detection and redaction to replace sensitive entities with placeholders.
B.Shuffle the dataset to break associations between PII and responses.
C.Increase the batch size during fine-tuning to average out PII exposure.
D.Tokenize the dataset with a custom vocabulary that includes PII patterns.
AnswerA

Detecting and redacting PII replaces names, emails, and phone numbers with generic placeholders, removing sensitive information while preserving the conversational structure needed for fine-tuning. This directly addresses the privacy requirement and prevents the model from memorizing or emitting real customer data, making it the correct preparation step.

Why this answer

Privacy compliance requires removing or masking personally identifiable information before training. PII detection and redaction replaces sensitive entities with placeholders, preserving the dataset's instructional value while preventing the model from memorizing real customer details. Tokenization, batch size, and shuffling do not alter the presence of PII and therefore cannot satisfy the requirement to protect sensitive data during fine-tuning.

Exam trap

The trap here is confusing training-time hyperparameters like batch size or shuffling with data-preparation steps that actually remove sensitive content.

8
MCQeasy

A data engineer is preparing a JSONL instruction dataset for an NVIDIA NeMo supervised fine-tuning run. Each line currently contains a free-form 'text' field with the instruction, context, and response concatenated. The training configuration expects the standard NeMo instruction-tuning schema with separate fields for the task instruction, optional context, and the expected response. What is the most appropriate data preparation step?

A.Keep the free-form text field and modify the NeMo training configuration to treat the entire line as the response.
B.Convert the JSONL file to plain text with one example per paragraph and let the tokenizer infer the instruction and response boundaries.
C.Duplicate each record and label one copy as instruction and the other as response so the loader sees two fields per example.
D.Parse each record and rewrite it into the expected fields, such as instruction, input, and output, while preserving the original text content and escaping any embedded quotes.
AnswerD

NeMo's supervised fine-tuning data loader expects distinct fields for the instruction, optional context, and response, so parsing the concatenated text into those fields makes the dataset consumable without custom loader code. Preserving content and escaping quotes maintains data fidelity and prevents malformed JSONL lines that would break parsing during training.

Why this answer

NeMo's supervised fine-tuning loader expects separate instruction, context, and response fields, so the free-form text must be parsed into that schema with content preserved and quotes escaped. This produces a dataset the standard training configuration can consume directly, keeping the run reproducible and allowing loss to be computed on the response portion alone.

Exam trap

The trap here is assuming the training configuration can be bent to accept free-form text instead of reshaping the data to match the loader's expected schema.

9
MCQmedium

You are preparing a 2 TB corpus of English and German web text for continued pretraining of a NeMo-based LLM. The German portion includes many pages with unescaped HTML entities and mixed-language sentences. Which NeMo Curator stage should you apply to remove boilerplate, fix HTML artifacts, and filter low-quality documents before tokenization?

A.NeMo Curator's text cleaning and heuristic filtering pipeline, including HTML unescaping, language identification, and quality classifier stages.
B.NeMo's Megatron-LM pretraining script with an increased dropout rate on the embedding layer.
C.NeMo Guardrails with a custom Colang flow that blocks documents containing HTML tags.
D.NVIDIA Triton Inference Server with a Python backend that preprocesses each document at inference time.
AnswerA

NeMo Curator provides modular stages for exactly this: HTML unescaping, boilerplate removal, language identification using fastText, and quality filtering with a classifier or heuristic scores. Running these before tokenization ensures the German and English subsets are clean and consistently language-tagged, which reduces noise during continued pretraining and prevents mixed-language documents from degrading the model's language modeling.

Why this answer

NeMo Curator is the correct tool because it supplies production-grade stages for HTML unescaping, language identification, and quality filtering that operate on large-scale text before tokenization. These stages directly address the German/English mixed-language and malformed HTML issues. The other options are training, runtime safety, or serving components, none of which clean a pretraining corpus.

Exam trap

The trap here is assuming that any NVIDIA component that processes text, such as Guardrails or Triton, can substitute for a dedicated data curation pipeline when preparing a training corpus.

10
MCQeasy

What is the primary function of data 'normalization' in the context of preparing inputs for a Transformer model?

A.Increasing the complexity of the input text to help the model learn more features.
B.Standardizing the text to ensure consistent interpretation of tokens.
C.Encrypting the dataset to protect sensitive information during training.
D.Compressing the dataset to reduce storage space on the GPU disk.
AnswerB

Normalization removes variations in formatting, such as whitespace or character encoding, that do not carry semantic weight. This consistency is vital, as it ensures that the tokenizer maps the same concepts to the same token IDs, preventing unnecessary ambiguity and ensuring the model learns stable, reliable relationships between tokens.

Why this answer

Normalization, such as standardizing whitespace, handling special tokens, and ensuring consistent character encodings, ensures that the model interprets input text in a predictable manner. By removing noise that doesn't contribute to semantic meaning, the model can focus its capacity on learning complex linguistic patterns. This is a standard and essential step in any high-performance AI data pipeline using NVIDIA accelerated computing.

Exam trap

Candidates often confuse text normalization with tokenization or embedding generation, failing to realize that normalization strictly standardizes raw text characters and whitespace before token processing begins.

11
MCQhard

A team is building an instruction-tuning dataset in NeMo from 40,000 internal support tickets. Each ticket contains a customer problem and a resolved answer, but the resolution text sometimes includes the customer's name, account number, and internal case IDs. The team plans to use NeMo Curator to produce training-ready JSONL. Which approach best prepares this data for instruction tuning while limiting personally identifiable information exposure?

A.Hash the entire ticket text with a cryptographic digest and train on the hashes, since the model can learn the mapping between hashed problems and hashed answers without ever seeing the original identifiers.
B.Run a PII redaction stage that detects and replaces names, account numbers, and case IDs with placeholders, then convert problem-resolution pairs into the instruction, input, and output schema before export.
C.Convert the tickets into instruction, input, and output JSONL first, then fine-tune the model and rely on a post-training output filter to block any generated response that contains an account number or customer name.
D.Drop every ticket whose resolution contains any digit, since account numbers and case IDs always contain numeric characters, and train only on the remaining text-only resolutions.
AnswerB

Redacting identifiers before schema conversion prevents the model from memorizing sensitive strings and keeps the pipeline deterministic. NeMo Curator supports PII detection and replacement as a distinct stage, and converting cleaned records into the instruction, input, and output fields yields the JSONL format the fine-tuning configuration expects. Doing redaction first also avoids leaking identifiers into derived fields.

Why this answer

Sensitive identifiers must be removed from the corpus before training, not mitigated after the fact. Running a PII detection and replacement stage in NeMo Curator, then mapping problem-resolution pairs into the instruction, input, and output schema, produces compliant JSONL that the fine-tuning job can consume directly. This ordering keeps identifiers out of both the training data and the resulting model weights.

Exam trap

The trap here is treating PII handling as a post-training output filter, when the identifiers must be removed from the training corpus itself before the model ever sees them.

12
MCQmedium

You are curating a 2 TB corpus of NVIDIA technical documentation and Python code for continued pretraining of a NeMo-based LLM. A colleague proposes filtering out any document containing the token sequence 'CUDA' to reduce hardware-specific bias. What is the most appropriate response?

A.Accept the filter, because removing every occurrence of a hardware-specific token guarantees the model will not overfit to NVIDIA-specific APIs and will generalize better to other vendors.
B.Accept the filter but apply it only to documents where 'CUDA' appears more than ten times, since low-frequency occurrences are harmless and high-frequency ones indicate redundant marketing material.
C.Reject the filter, because removing documents based on a single high-signal domain token destroys the domain-specific signal the model needs and is better handled by deduplication and quality scoring.
D.Accept the filter, because NVIDIA NeMo requires that proprietary product names be excluded from pretraining corpora to comply with the model card's data provenance requirements.
AnswerC

The proposed filter is a crude keyword removal that discards entire documents whose core subject is the target domain. In NeMo Curator pipelines, quality is improved through exact and fuzzy deduplication, heuristic quality filters, and classifier-based scoring, not by excising a single token that defines the domain. Removing these documents would leave the model undertrained on the very terminology it must learn.

Why this answer

Keyword-based deletion of a domain-defining token is a destructive filter, not a quality filter. Effective NeMo data curation relies on deduplication, quality heuristics, and classifier scoring to remove low-value or redundant records while preserving coherent, in-domain technical content. The corpus exists to teach NVIDIA-specific concepts, so removing 'CUDA' would directly undermine the training objective.

Exam trap

The trap here is assuming that removing a vendor-specific keyword reduces bias, when it actually strips high-value domain signal that the continued pretraining corpus was built to provide.

13
MCQhard

When training a model for a highly technical domain with a scarcity of high-quality data, which data augmentation strategy is most likely to preserve the model's reliability?

A.Randomly masking 50% of the words in the training documents.
B.Using a smaller, unverified LLM to generate completely new technical scenarios.
C.Generating synthetic examples based on ground-truth technical documentation templates.
D.Translating the dataset into multiple languages using a generic online translator.
AnswerC

Templated generation ensures that the synthetic data adheres to the logical and structural rules of the domain. By basing synthetic examples on verified ground-truth templates, you maximize data variety while maintaining factual accuracy, which is the safest and most effective way to address data scarcity in technical domains.

Why this answer

In technical domains, synthetic data generation must be grounded in existing, verified documents to avoid 'hallucinating' technical facts. Using LLMs to paraphrase or summarize existing high-quality technical content while maintaining strict constraints ensures that the new data follows the same logic and terminology. This method expands the training set while minimizing the risk of introducing incorrect facts, which is essential for high-stakes technical domains.

Exam trap

Candidates often choose unconstrained generative augmentation, which introduces hallucinations. They fail to realize that grounding synthetic data in verified templates is the only way to maintain technical reliability.

14
MCQeasy

Which technique is most effective for mitigating data leakage during the training of an LLM on time-series-related document data?

A.Random shuffling of all documents regardless of their timestamps.
B.Chronological splitting of the dataset based on a cutoff date.
C.Oversampling the minority class in the training set to improve balance.
D.Applying aggressive data normalization to all numerical values.
AnswerB

Chronological splitting ensures that the model is trained exclusively on data from the past, while validation and testing are performed on subsequent time periods. This mirrors the real-world deployment environment, where the model must predict future events based only on information that has already occurred in history.

Why this answer

Preventing data leakage in temporal datasets requires strict chronological separation between training, validation, and testing sets. If future information is inadvertently included in the training data, the model will 'cheat' by observing future events during training, leading to artificially inflated performance metrics that do not generalize to actual production deployment scenarios where future data is unavailable.

Exam trap

Candidates often rely on random shuffling or standard cross-validation, failing to recognize that time-series text data requires strict chronological splitting to prevent future information leakage.

15
MCQmedium

You are preparing a large instruction-tuning dataset for an NVIDIA NeMo-based LLM. The raw data consists of user queries and assistant responses collected from a customer support system, stored as JSON lines. During preprocessing, you notice that many responses contain personally identifiable information (PII) such as names, email addresses, and phone numbers. You need to ensure the dataset is safe for training while preserving as much semantic content as possible. Which approach is most appropriate for handling PII in this dataset?

A.Leave the PII intact but exclude the dataset from production use and train only in a sandbox environment.
B.Hash each detected PII string with SHA-256 and substitute the hash into the text.
C.Apply a rule-based regex filter to remove any line containing PII, then train on the remaining data.
D.Use a named entity recognition (NER) model to detect PII and replace each entity with a generic placeholder (e.g., [NAME]).
AnswerD

Replacing PII with placeholders preserves sentence structure and semantic meaning while removing sensitive information. This is a standard de-identification technique that maintains data utility for instruction tuning. It also handles variations better than simple regex and is compatible with NeMo preprocessing pipelines.

Why this answer

Replacing PII with generic placeholders using an NER model is the most balanced approach: it removes sensitive information while preserving the linguistic patterns and context needed for instruction tuning. This method aligns with NVIDIA NeMo's data preprocessing recommendations for responsible AI. It also avoids the pitfalls of discarding data or introducing unintelligible tokens.

Exam trap

The trap here is assuming that regex-based removal is sufficient, but PII can be unstructured and context-dependent, making NER more reliable.

16
Multi-Selectmedium

When assessing the quality of a dataset for instruction fine-tuning, which TWO metrics or methods are considered most reliable for measuring dataset diversity?

Select 2 answers
A.Embedding-based clustering to visualize topical coverage across the corpus.
B.Calculating the total number of words in the dataset.
C.Perplexity distribution analysis across segments of the dataset.
D.Measuring the average response length for every instruction.
E.Checking if the dataset is solely in ASCII format.
AnswersA, C

Clustering document embeddings is a standard way to verify that the dataset covers a wide spectrum of topics. If the embeddings form only a few dense clusters, the dataset is likely too narrow. A diverse dataset should exhibit a broader distribution across the vector space, indicating varied content.

Why this answer

Dataset diversity ensures that the model encounters a broad range of topics and linguistic patterns, which is critical for generalization. Using embedding-based clustering allows for the identification of thematic coverage, while perplexity distribution analysis helps assess whether the dataset contains a balance of common and complex structures. These methods together provide a quantitative view of the data's breadth, reducing the risk of bias or overfitting.

Exam trap

Candidates often rely on simple text length or word count metrics, overlooking advanced quantitative methods like embedding-based clustering and perplexity distribution for measuring true dataset diversity.

17
MCQmedium

You are preparing a multilingual corpus for pre-training an LLM. The corpus contains documents in English, Spanish, and German, but the English portion is 80% of the data. You want to avoid the model becoming biased toward English. Which data preparation technique should you apply?

A.Remove all English documents to force multilingual learning
B.Translate all documents into English
C.Apply a language-specific tokenizer for each language
D.Oversample the non-English documents or undersample the English documents
AnswerD

Adjusting the sampling ratios by oversampling minority languages or undersampling the dominant language balances the effective training distribution. This reduces English bias and encourages the model to allocate capacity to Spanish and German. It is a standard technique in multilingual data preparation and can be implemented via NeMo Curator or custom sampling logic.

Why this answer

Balancing the language distribution by oversampling minority languages or undersampling the dominant one directly addresses the 80% English skew. This ensures the model sees a more representative mix during training, reducing bias toward English and improving performance on Spanish and German. It is a standard data preparation technique for multilingual corpora.

Exam trap

The trap here is thinking that tokenizer changes or translation solve imbalance, when the core issue is the proportion of training examples per language.

18
MCQmedium

You are using NVIDIA NeMo Curator to filter a 600 GB web-crawl corpus before pre-training. Your team wants to remove exact duplicates and near-duplicates to reduce memorization and speed up training. Which NeMo Curator stage should you apply?

A.HeuristicFilter stage with a quality threshold
B.DocumentDownloader stage followed by DocumentResolver
C.ExactDuplicates and FuzzyDuplicates stages
D.ClassifierFilter stage using a domain classifier
AnswerC

NeMo Curator provides ExactDuplicates and FuzzyDuplicates stages that identify redundant documents; ExactDuplicates catches byte-identical texts, while FuzzyDuplicates uses MinHash/LSH to find near-duplicates. Applying both in sequence on the 600 GB corpus removes redundant content, lowering memorization risk and reducing the effective training token count.

Why this answer

NeMo Curator's ExactDuplicates and FuzzyDuplicates stages are purpose-built for removing redundant documents. ExactDuplicates handles byte-level matches, while FuzzyDuplicates uses MinHash and LSH to catch near-identical texts. Together they reduce corpus size without discarding unique content, which directly lowers memorization risk and training cost on a large web-crawl dataset.

Exam trap

The trap here is assuming that any quality filter removes duplicates, when quality filters score documents individually rather than comparing them to each other.

19
MCQhard

A team is building a NeMo-based LLM pipeline and must tokenize a corpus that mixes English, Japanese, and Python source code. They plan to train a custom tokenizer with NVIDIA NeMo. Which tokenizer configuration best supports all three content types without excessive sequence length?

A.Byte-level Byte-Pair Encoding with a vocabulary of 64,000 to 128,000 tokens trained on a balanced sample of all three content types.
B.SentencePiece unigram tokenizer trained exclusively on Japanese text.
C.WordPiece tokenization trained only on the English portion of the corpus.
D.Byte-Pair Encoding with a vocabulary limited to 8,000 tokens.
AnswerA

Byte-level BPE avoids unknown tokens by operating on raw bytes, so any script or code symbol is representable. A large vocabulary of 64k-128k tokens captures common multilingual subwords and code identifiers, reducing sequence length. Training on a balanced sample ensures the merge rules reflect English, Japanese, and Python frequencies, which is exactly what this mixed corpus needs.

Why this answer

Byte-level BPE with a large vocabulary trained on a balanced multilingual and code sample handles arbitrary scripts and symbols without unknown tokens. It compresses common English words, Japanese subwords, and Python identifiers into single tokens, controlling sequence length. Training on all three distributions aligns merge statistics with actual usage, unlike English-only or Japanese-only training, and the 64k-128k range provides enough capacity for multilingual plus code coverage.

Exam trap

The trap here is choosing a tokenizer algorithm without checking whether its training data covers all content types, since out-of-vocabulary scripts and code inflate sequence length regardless of algorithm.

20
MCQmedium

You are fine-tuning a LLM on a large corpus of technical documentation. You notice the model struggles with domain-specific terminology that frequently appeared in the training data but was inconsistently formatted. Which data preparation technique best addresses this issue?

A.Increasing the batch size during the training phase.
B.Expanding the model's depth with additional transformer layers.
C.Applying consistent text normalization and domain-specific tokenization rules.
D.Reducing the learning rate to prevent overfitting on the noise.
AnswerC

Standardizing formatting and applying custom tokenization ensures that domain-specific terminology is tokenized consistently across the corpus. This alignment allows the model to build stronger semantic associations for technical terms, significantly reducing the probability of errors caused by variations in casing, punctuation, or special character usage.

Why this answer

Consistent normalization, such as lemmatization or standardizing case and special characters, ensures the tokenizer treats synonymous terms identically. In NVIDIA NeMo workflows, data quality directly impacts convergence speed and model accuracy. By reducing vocabulary noise during the preprocessing stage, the model can dedicate more capacity to learning semantic relationships rather than mapping variants of the same technical term to different embedding spaces, ultimately improving downstream performance on domain-specific benchmarks.

Exam trap

Candidates often focus on increasing model size or training epochs to fix terminology issues, ignoring that inconsistent tokenization and raw data noise are the root causes of poor performance.

21
Multi-Selectmedium

You are preparing a large instruction-tuning dataset for an LLM using NVIDIA NeMo. You need to ensure the dataset supports efficient training and evaluation. Which TWO steps are essential in the data preparation pipeline? (Choose two.)

Select 2 answers
A.Convert the dataset into a format compatible with NeMo, such as JSONL with input and output fields
B.Train a tokenizer from scratch on the dataset
C.Apply data augmentation using back-translation
D.Remove all punctuation from the text
E.Split the dataset into training, validation, and test sets
AnswersA, E

NeMo's instruction-tuning pipeline expects data in a structured format, commonly JSONL with fields like input and output or prompt and completion. Converting the raw dataset into this format ensures the data loader can parse and batch examples correctly. Without this step, training would fail or require custom parsing code, so it is essential.

Why this answer

Formatting the dataset into a NeMo-compatible structure such as JSONL with input and output fields ensures the data loader can consume it, while splitting into training, validation, and test sets enables proper model evaluation and hyperparameter tuning. These two steps are foundational; other activities like tokenizer training or augmentation are optional and context-dependent.

Exam trap

The trap here is treating optional enhancements like tokenizer retraining or augmentation as mandatory, when the truly essential steps are formatting and splitting.

22
MCQmedium

You are preparing a dataset for instruction fine-tuning an LLM using NVIDIA NeMo. The dataset contains pairs of instructions and responses, but you notice that some responses are significantly longer than others, and a few are extremely short (e.g., 'Yes' or 'No'). You want to ensure the model learns to generate appropriate-length responses. Which data preparation technique is most effective?

A.Augment the dataset by duplicating short responses to increase their frequency.
B.Filter out all responses shorter than 10 tokens to remove trivial answers.
C.Ensure the dataset has a balanced distribution of response lengths by sampling or resampling.
D.Include a length token or metadata in the input to indicate the desired response length.
AnswerC

Balancing response lengths helps the model learn to generate appropriate lengths based on the instruction. By avoiding overrepresentation of very short or very long responses, the model can better generalize. This can be done by stratified sampling or by capping extreme lengths, but maintaining a natural distribution is key.

Why this answer

A balanced distribution of response lengths allows the model to learn when to generate concise versus detailed answers. Overrepresentation of short responses may lead to truncation, while too many long responses can cause verbosity. Resampling to achieve a more uniform distribution helps the model internalize the relationship between instruction complexity and response length.

Exam trap

The trap here is thinking that removing short responses solves the issue, but it actually biases the model and removes valid training signals.

23
MCQhard

You are preparing a dataset for pretraining an LLM using NVIDIA NeMo's Megatron-LM. The dataset consists of JSONL files where each line contains a 'text' field. You need to convert these files into the binary format required by Megatron for efficient training. Which tool or method should you use to perform this conversion?

A.Use Apache Arrow to serialize the JSONL data into a columnar format and load it with a custom data loader.
B.Use the `jsonl_to_bin` utility provided by the Hugging Face Transformers library.
C.Use the `nemo_curator` library to convert JSONL to TFRecord format, then train directly on TFRecords.
D.Use the `preprocess_data.py` script from the Megatron repository to tokenize and create indexed binary files.
AnswerD

The `preprocess_data.py` script in Megatron-LM is specifically designed to tokenize JSONL text data and produce binary files with index mapping, which are required for efficient data loading during pretraining. It handles tokenization, concatenation, and sharding, making it the standard tool for this conversion.

Why this answer

Megatron-LM provides the `preprocess_data.py` script specifically for converting JSONL text data into indexed binary files that its data loader can read efficiently. This script handles tokenization, document concatenation, and sharding, ensuring compatibility with Megatron's training pipeline. It is the recommended and standard method for preparing data for Megatron-based pretraining.

Exam trap

The trap here is assuming that any serialization format like TFRecord or Arrow will work, but Megatron requires its own binary format produced by its preprocessing script.

24
MCQmedium

An enterprise is preparing a massive multi-terabyte corpus of specialized PDF documents containing technical schematics and dense tables for training a domain-specific LLM using NeMo Curator. During the document extraction pipeline, you notice that tabular data layouts are scrambled into unstructured string tokens, destroying spatial relationships. Which data preparation strategy should you implement first within NeMo Curator to preserve tabular integrity before tokenization?

A.Apply aggressive regex-based whitespace removal across all extracted string buffers to normalize sentence boundaries prior to embedding generation.
B.Deploy NeMo Curator's layout-aware PDF extraction modules featuring visual bounding box detection to parse tables into structured Markdown formatting.
C.Increase the chunk size parameter in the tokenizer configuration so that entire table blocks fit inside a single oversized context window.
D.Convert all PDF pages into low-resolution JPEG images to bypass text extraction errors and feed raw pixels directly into a text-only causal language model.
AnswerB

Layout-aware extraction uses visual bounding box detection to identify table cells and their spatial relationships, emitting structured Markdown rather than scrambled token strings. This preserves row and column integrity before tokenisation, directly addressing the destroyed tabular layouts described in the stem.

Why this answer

NeMo Curator provides specialized PDF extraction utilities that leverage advanced computer vision and layout parsing models to accurately identify bounding boxes for tables and figures. Preserving tabular markdown structures prevents spatial collapse, ensuring downstream tokenizers capture relational data accurately. This step is critical in domain-specific LLM training because scrambled tables introduce severe noise that degrades reasoning capabilities.

Exam trap

Candidates often assume standard text extraction tools are sufficient for all PDFs, overlooking that raw text extraction strips structural bounding box metadata essential for tabular layout preservation.

25
MCQeasy

A team is preparing a customer-support chat dataset for instruction fine-tuning of an LLM. The raw data contains HTML tags, inconsistent date formats, and emoji. Which data preparation step should be performed first to make the text usable for tokenization?

A.Normalize and clean the text
B.Tokenize the text with the model's tokenizer
C.Convert the text to lowercase
D.Split the dataset into train and validation sets
AnswerA

Normalization and cleaning remove HTML tags, standardize date formats, and handle emoji so the text is consistent before tokenization. This step ensures the tokenizer sees clean input, reducing noise in the training data and improving the quality of the instruction-tuning examples. It is the logical first step in the preparation pipeline.

Why this answer

Cleaning and normalizing the raw chat data first removes HTML, standardizes dates, and handles emoji, producing consistent text for tokenization. This order prevents noisy artifacts from becoming tokens and ensures both training and validation splits receive the same treatment. It is the foundational step before tokenization and dataset splitting.

Exam trap

The trap here is treating tokenization as a cleaning step, when tokenizers faithfully encode whatever noise is present in the input text.

26
MCQmedium

Refer to the exhibit. You are reviewing the configuration file for a data preprocessing pipeline. Why is the 'min_words' filter set to 50 in the context of LLM training?

A.It limits the memory footprint of the dataloader.
B.It enables faster tokenization by reducing the number of input files.
C.It removes low-quality snippets to ensure meaningful semantic context.
D.It forces the model to ignore PII-heavy documents.
AnswerC

Removing short sequences filters out noise like headers, footers, or incomplete sentences that provide little linguistic value. This ensures the model spends its training budget on high-quality text, improving its ability to learn complex long-range dependencies and overall coherence within the target language domain.

Why this answer

The 'min_words' filter removes low-information or malformed snippets that lack sufficient context for effective transformer learning. In large-scale training, such as those performed on NVIDIA H100 GPU clusters, including very short strings increases noise and consumes valuable compute cycles without contributing to meaningful semantic representation. Setting a minimum length ensures the model learns from coherent passages rather than disjointed fragments, promoting better structural understanding of the training corpus.

Exam trap

Candidates often assume filtering is purely for privacy, missing the technical reality that very short, low-information snippets introduce noise that degrades the model's ability to learn complex linguistic structures.

27
Multi-Selecthard

You are preparing a massive dataset for training a NeMo-based LLM. Which TWO data preprocessing steps are critical to prevent data leakage and ensure model quality?

Select 2 answers
A.Performing temporal or content-based splitting to isolate the validation set.
B.Implementing fuzzy deduplication to remove near-duplicate documents.
C.Increasing the frequency of data augmentation in the training pipeline.
D.Adding metadata tags to all training sequences for classification.
E.Converting all text to lowercase to increase vocabulary efficiency.
AnswersA, B

Isolating data based on timestamps or content clusters prevents the model from seeing future data or overlapping information during training. This ensures the evaluation set remains truly unseen, providing a realistic assessment of how the model will perform on new, unseen data in production environments.

Why this answer

Data leakage occurs when test data is inadvertently included in the training set, leading to inflated performance metrics. Deduplication is equally vital, as repetitive data causes the model to memorize samples rather than generalize. In NVIDIA workflows, these steps are typically performed via distributed scripts on the cluster before tokenization, ensuring that the model learns unique, non-overlapping information across all shards.

Exam trap

Test-takers often confuse basic random splitting with temporal or content-based splitting, missing the fact that standard random splits fail to prevent data leakage in LLM datasets.

28
MCQmedium

Why is it important to perform 'domain-specific' data cleaning when preparing a corpus for fine-tuning a medical LLM?

A.It helps to anonymize the patient data by replacing all medical terms with generic labels.
B.Generic cleaning often removes or modifies critical domain-specific nomenclature.
C.It ensures that the dataset size is significantly reduced to fit in GPU cache.
D.It automatically corrects all scientific inaccuracies present in the original documents.
AnswerB

Medical nomenclature is highly specialized and often uses abbreviations or symbols that generic cleaning scripts might identify as noise or formatting errors. By using domain-specific cleaning, developers ensure that these critical terms are preserved, which is essential for maintaining the accuracy of the model's domain knowledge.

Why this answer

Medical text contains specific jargon, abbreviations, and relationships that generic cleaning might misinterpret. Generic tools often remove or alter terms that are critical to medical context, such as drug names or procedural codes. By using domain-specific cleaning, you ensure that the model retains the precise terminology necessary for high-stakes, accurate clinical reasoning in medical applications.

Exam trap

Candidates often assume that generic cleaning tools are sufficient for all data types, overlooking the fact that medical terminology is fragile and easily destroyed by standard normalization or stop-word removal.

29
MCQhard

Refer to the exhibit. What is the impact of this filter on the training corpus?

A.It limits the vocabulary size to 20% of the original content.
B.It discards documents that are overly repetitive or lack linguistic richness.
C.It forces the model to use 20% more computation for tokenization.
D.It reduces the training corpus size by exactly 20%.
AnswerB

A low lexical diversity score indicates that a document uses a very small set of unique words relative to its total length. This is characteristic of repetitive or low-quality content. Filtering these out ensures the model learns from diverse, high-quality, and informative text, which improves overall model performance.

Why this answer

This filter removes documents with low lexical diversity, which often contain repetitive, low-value, or 'boilerplate' text. Such content provides little signal for the model to learn meaningful language patterns. By enforcing a minimum diversity threshold, you ensure the corpus consists of richer, more informative language, which typically leads to better convergence and higher quality output in the trained model.

Exam trap

Test-takers often misinterpret lexical diversity filters as removing long documents or high-frequency vocabulary, whereas they actually target repetitive, low-value boilerplate text.

30
MCQhard

You are building a pretraining dataset from a large collection of source-code repositories for an NVIDIA NeMo LLM. The data includes many files with licenses, generated code, and minified JavaScript. Which NeMo Curator-based approach best improves code data quality before tokenization?

A.Convert all code to a single programming language using an automated transpiler before tokenization.
B.Increase the model's context window to 32k tokens so that entire repositories fit in one sequence.
C.Tokenize all files with a byte-level BPE tokenizer and skip any further filtering.
D.Use language-specific filters that detect and remove minified files, license headers, and auto-generated code, then apply deduplication.
AnswerD

Minified JavaScript, license headers, and generated code are low-value or repetitive patterns that harm code pretraining. NeMo Curator supports custom filters for line length, ratio of alphanumeric characters, and detection of boilerplate. Deduplication removes repeated generated files. Applying these before tokenization ensures the model learns from meaningful code rather than noise, improving downstream code generation and understanding.

Why this answer

Language-specific filters combined with deduplication directly target the described quality issues. Minified files, license headers, and generated code are common in code corpora and can be detected with heuristics such as average line length, ratio of whitespace, and presence of standard license text. Deduplication removes repeated generated files.

This preprocessing ensures the pretraining corpus contains meaningful code, which improves the model's code capabilities.

Exam trap

The trap here is believing that a more powerful tokenizer or a larger context window can compensate for low-quality or repetitive code data.

31
MCQeasy

You are curating instruction-tuning data for an NVIDIA NIM-deployed LLM. The raw dataset contains many near-duplicate instruction-response pairs that differ only by punctuation and whitespace. Which data preparation step is most appropriate to remove these before fine-tuning?

A.Increase the batch size during fine-tuning to average out duplicate examples.
B.Convert all instructions to lowercase and remove all punctuation globally.
C.Use NVIDIA Triton Inference Server dynamic batching to filter duplicates at serving time.
D.Apply fuzzy deduplication using MinHash LSH over normalized instruction-response text.
AnswerD

Fuzzy deduplication with MinHash LSH is designed to catch near-duplicates that differ by punctuation, whitespace, or minor edits. Normalizing text before hashing ensures that trivial variations map to the same signature. Removing these duplicates prevents the model from overfitting to repeated examples and reduces wasted training compute, which is especially important when fine-tuning an instruction-following model on a curated dataset.

Why this answer

Fuzzy deduplication with MinHash LSH is the standard approach for near-duplicate removal in instruction-tuning datasets. It normalizes text for comparison, generates MinHash signatures, and uses locality-sensitive hashing to find similar pairs efficiently at scale. This directly removes the redundant examples described, improving training efficiency and reducing overfitting without altering the semantic content of the retained data.

Exam trap

The trap here is thinking that deduplication means exact string matching or global text normalization, when near-duplicates require similarity-based methods that preserve the original text.

32
MCQmedium

You are preparing a large corpus of customer support transcripts for continued pretraining of an NVIDIA NeMo Megatron model. The transcripts contain personally identifiable information such as names, email addresses, and account numbers, and company policy requires that this information be removed before training while preserving as much linguistic context as possible for the model to learn from. Which data preparation approach best satisfies both requirements?

A.Hash every token in the corpus so that the original text cannot be reconstructed, and train on the hashed sequences.
B.Replace all digits and capitalized words with a generic mask token across the entire corpus before training.
C.Detect PII spans with a combination of regular expressions and a named-entity recognition model, then replace each detected span with a consistent placeholder token that preserves the surrounding sentence structure.
D.Drop every transcript that contains any detected PII so that no sensitive information reaches the training pipeline.
AnswerC

Regexes reliably catch structured identifiers such as emails and account numbers, while a named-entity recognition model covers names and locations that patterns miss. Replacing spans with consistent placeholders removes the sensitive values but keeps sentence structure and surrounding context intact, which is exactly what continued pretraining needs to learn language patterns without memorizing private data.

Why this answer

Combining regex detection for structured identifiers with named-entity recognition for names and locations covers the PII surface area, and replacing detected spans with consistent placeholders removes sensitive values while preserving the sentence context that continued pretraining depends on. Dropping records, masking all digits and capitals, or hashing every token either destroys usable language data or fails to anonymize reliably.

Exam trap

The trap here is treating PII removal as a binary choice between deleting records and destroying all structure, when span-level replacement preserves the linguistic signal.

33
MCQhard

Refer to the exhibit. What is the intended outcome of this data cleaning configuration for a Large Language Model pre-training corpus?

A.It specifically targets the removal of personally identifiable information (PII).
B.It optimizes for the removal of low-quality or nonsensical text while minimizing redundancy.
C.It enforces a strict length-based chunking strategy for all documents.
D.It converts all text to a vector space representation before filtering.
AnswerB

The configuration uses perplexity filtering to identify incoherent content and length constraints to exclude short, low-information strings. The MinHash algorithm effectively manages the similarity threshold to eliminate near-duplicate documents. This combination ensures that the training dataset is concise, coherent, and free of redundant, low-value information inputs.

Why this answer

This configuration aims to remove low-quality text that fails to meet minimum length requirements or exhibits high perplexity (indicating gibberish or low coherence). Simultaneously, MinHash deduplication identifies and removes near-duplicate documents exceeding a 95% similarity threshold. This cleanup process is vital for pre-training, as it filters out low-value, noisy data that could impede model convergence and general quality during the training cycle.

Exam trap

Candidates often assume cleaning configurations only remove empty strings, failing to recognize the combined role of length/perplexity filters and MinHash deduplication in eliminating noise.

34
MCQmedium

Why is 'tokenization stability' a critical metric when preparing data for NVIDIA-based LLM deployment?

A.It guarantees that the model will always generate the same output for a given prompt.
B.It ensures that the GPU memory usage remains constant during the training process.
C.It prevents unexpected input shifts between training and inference environments.
D.It reduces the total number of parameters required for the embedding layer.
AnswerC

Stability ensures that the mapping between text and tokens remains consistent. If an inference pipeline tokenizes text differently than the training pipeline, the model encounters a distribution shift. This mismatch can result in degraded model performance, incorrect reasoning, or complete failure, making stability a foundational requirement for robust production systems.

Why this answer

Tokenization stability ensures that the same input text consistently maps to the same sequence of tokens across different environments or library versions. If tokenization is inconsistent, the model might receive unexpected inputs compared to what it observed during training, leading to severe performance degradation. For NVIDIA-based deployments, deterministic tokenization is essential for maintaining production-level reliability and predictable model behavior across various inference pipelines.

Exam trap

Test-takers frequently assume tokenization stability only affects processing speed, missing its critical role in preventing unexpected input shifts between training and inference environments.

35
MCQmedium

You are preparing a dataset of customer reviews for fine-tuning an LLM to generate concise summaries. The reviews are in multiple languages, but the target summaries must be in English. You have a limited budget for translation. Which data preparation step is most critical to ensure the fine-tuned model produces high-quality English summaries?

A.Translate all reviews into English and then train the model to summarize English text.
B.Ensure that each training example pairs a review in its original language with a high-quality English summary.
C.Use a multilingual LLM to translate all non-English reviews into English before training.
D.Filter the dataset to include only reviews originally written in English.
AnswerB

This approach directly trains the model to perform cross-lingual summarization: input in any language, output in English. It leverages the original text without translation errors and teaches the model to generate English summaries regardless of source language. This is the most effective strategy for the stated goal.

Why this answer

Pairing original-language reviews with English summaries directly trains the model for cross-lingual summarization, which is the end goal. This avoids translation errors and preserves the original semantics. It also prepares the model for real-world multilingual inputs, ensuring it can generate English summaries without an intermediate translation step.

Exam trap

The trap here is assuming that translating everything to English first is simpler, but it fails to train the model for multilingual input and may introduce translation artifacts.

36
MCQhard

When fine-tuning an LLM to follow specific safety protocols, why is the inclusion of 'adversarial' examples in the training data considered a best practice?

A.To artificially inflate the model's perplexity scores during evaluation.
B.To ensure the model learns to refuse requests that violate established safety policies.
C.To increase the model's fluency in generating complex technical instructions.
D.To allow the model to learn and reproduce the adversarial techniques during generation.
AnswerB

Adversarial training exposes the model to edge-case prompts designed to break safety constraints. By learning to handle these, the model becomes more robust against malicious attempts to bypass safety filters. This ensures consistent enforcement of safety protocols in production, making the model more secure and reliable for enterprise use.

Why this answer

Adversarial examples test the model's ability to maintain safety boundaries even when prompted with manipulative or malicious inputs. By training on these examples, the model learns to identify and refuse requests that violate safety protocols. This proactive preparation is essential for deploying LLMs in enterprise environments where maintaining strict safety and compliance standards is required, protecting against sophisticated prompt injection and bypass techniques.

Exam trap

Candidates often assume adversarial examples are only for improving accuracy or general performance, failing to recognize that they are specifically required to enforce safety boundaries and refusal behaviors in high-stakes environments.

37
Multi-Selectmedium

You are preparing a dataset for supervised fine-tuning (SFT) of an NVIDIA NeMo LLM to follow instructions. Which TWO data preparation practices are essential to ensure the model learns to generalize rather than memorize? (Choose two.)

Select 2 answers
A.Include a diverse set of instruction types and domains in the training data rather than repeating a few templates.
B.Increase the learning rate significantly to force the model to escape memorization of individual examples.
C.Duplicate the most common instruction-response pairs to reinforce the desired behavior.
D.Use only English instructions to simplify tokenization and avoid multilingual complexity.
E.Split the dataset into training, validation, and test sets with no overlapping instructions or responses.
AnswersA, E

Diversity in instruction types and domains encourages the model to learn the general pattern of following instructions rather than memorizing specific templates. If the training data is dominated by a few templates, the model may overfit to those formats and fail on novel instructions. A broad distribution of tasks improves zero-shot and few-shot generalization.

Why this answer

Disjoint data splits prevent leakage and give honest evaluation, while diverse instruction types and domains teach the model to follow instructions generally rather than memorize templates. Together, these practices reduce overfitting and improve generalization. The other options either fail to address memorization, actively encourage it, or unnecessarily restrict the data distribution.

Exam trap

The trap here is thinking that training hyperparameters like learning rate or duplicating data can substitute for proper data splitting and diversity when the goal is generalization.

38
MCQmedium

You are preparing a large instruction-tuning dataset with NVIDIA NeMo Curator. The dataset contains many near-duplicate instruction-response pairs that differ only in punctuation or minor wording. Which NeMo Curator stage should you apply to remove these near-duplicates before fine-tuning?

A.Fuzzy deduplication using MinHash locality-sensitive hashing.
B.Exact substring deduplication to remove overlapping text spans.
C.Heuristic filtering to remove short or malformed records.
D.Exact duplicate removal using a hash of the full instruction-response pair.
AnswerA

MinHash with locality-sensitive hashing groups documents by estimated Jaccard similarity, so near-duplicates with minor edits are clustered and one representative is kept. In NeMo Curator this is implemented as fuzzy deduplication and directly targets the scenario's near-duplicate instruction pairs. It scales to large datasets and removes redundancy that would otherwise bias fine-tuning.

Why this answer

Near-duplicate instruction-response pairs that differ only slightly require fuzzy matching, not exact hashing or heuristic quality filters. MinHash locality-sensitive hashing estimates Jaccard similarity between records and clusters near-identical ones, letting you retain a single representative. Applying fuzzy deduplication in NeMo Curator reduces redundancy, prevents the model from overfitting repeated patterns, and improves generalization on the instruction-tuning task.

Exam trap

The trap here is assuming that exact duplicate removal is sufficient, when minor punctuation or wording changes cause hashes to differ and near-duplicates to survive.

39
MCQmedium

You are preparing a multilingual corpus for pretraining with NVIDIA NeMo. The dataset contains documents in 40 languages, but the tokenizer was trained primarily on English. Which data preparation action best ensures that non-English text is represented efficiently during tokenization?

A.Increase the model's maximum sequence length to 8192 tokens to accommodate longer tokenized non-English text.
B.Train a new tokenizer on a balanced sample of all 40 languages using SentencePiece or Hugging Face tokenizers before pretraining.
C.Apply Unicode normalization form NFKC to all text and rely on the existing English tokenizer.
D.Translate all non-English documents to English using an NVIDIA NIM translation model before tokenization.
AnswerB

A tokenizer trained predominantly on English will split non-English words into many subword units, increasing sequence length and reducing effective context. Training a new tokenizer on a balanced multilingual sample ensures that each language has adequate vocabulary coverage, which lowers the number of tokens per document and improves training efficiency. This must be done before pretraining because the tokenizer defines the model's embedding space.

Why this answer

Training a new tokenizer on a balanced multilingual sample directly solves the vocabulary coverage problem. It ensures that frequent subwords in each language are included in the vocabulary, reducing token fragmentation and improving training efficiency. The other options either mask the symptom, alter the data inappropriately, or do not address tokenizer vocabulary at all.

Exam trap

The trap here is assuming that sequence length or Unicode normalization can compensate for a tokenizer that lacks vocabulary for the target languages.

40
MCQmedium

You are preparing a dataset of support tickets for a RAG system using NVIDIA NeMo. Many tickets are short and contain little context, which hurts retrieval quality. Which data preparation technique best improves retrieval by enriching each ticket with related information before embedding?

A.Increase the embedding dimension of the retrieval model.
B.Lower the similarity threshold used during retrieval.
C.Apply deduplication to remove similar tickets from the corpus.
D.Use an LLM to generate contextual summaries or metadata for each ticket and prepend them to the text before embedding.
AnswerD

Generating a summary or metadata adds missing context to each short ticket, so the embedded representation captures more relevant semantics and matches user queries better. Prepending this enriched text before embedding directly addresses the sparsity problem and improves retrieval, making it the appropriate technique for the scenario.

Why this answer

Short tickets embed poorly because they carry little semantic signal. Using an LLM to generate contextual summaries or metadata and prepending that text before embedding enriches each record, so the vector captures more of the ticket's intent and domain. This improves matching with user queries.

Deduplication, larger embeddings, and lower thresholds do not add missing context, so they cannot resolve the sparsity problem in this scenario.

Exam trap

The trap here is assuming that retrieval-time tuning, such as lowering a similarity threshold, can compensate for source documents that simply lack contextual content.

41
MCQhard

A team is preparing a mixed-language corpus for continued pretraining of an NVIDIA NeMo Megatron model. The corpus contains English, Japanese, and Arabic documents. Tokenizer analysis shows the current English-centric BPE vocabulary produces very long token sequences for Japanese and Arabic, inflating sequence length and compute cost. The team wants to reduce sequence length for non-English text without retraining the tokenizer from scratch and without degrading English performance. Which data preparation action best achieves this?

A.Lowercase and strip diacritics from all Japanese and Arabic text so that the English-centric vocabulary can represent it more efficiently.
B.Transliterate all Japanese and Arabic documents into Latin script using a standard romanization scheme before tokenization.
C.Increase the model's maximum sequence length and positional embedding size so that long token sequences for Japanese and Arabic fit without truncation.
D.Extend the existing BPE vocabulary with additional merges learned from a balanced multilingual sample, then retrain only the embedding and output layers while freezing the rest of the model.
AnswerD

Adding multilingual merges to the existing vocabulary reduces the number of tokens needed to represent Japanese and Arabic text, directly cutting sequence length and compute. Keeping the original English merges preserves English tokenization behavior, and retraining only the embedding and output layers adapts the new vocabulary entries without disturbing the pretrained transformer weights, which limits the risk of degrading English performance.

Why this answer

The high token counts for Japanese and Arabic stem from a vocabulary that lacks subword units for those scripts. Extending the BPE vocabulary with multilingual merges shortens their token sequences while preserving English merges, and retraining only the embedding and output layers adapts the new entries without perturbing the pretrained transformer. This lowers compute cost and keeps English behavior stable.

Exam trap

The trap here is treating long token sequences as a context-length problem to be solved by enlarging the window rather than as a vocabulary coverage problem rooted in the tokenizer.

Ready to test yourself?

Try a timed practice session using only Data Preparation questions.

CCNA Data Preparation Questions | Courseiva