NCP-GENL Fine-Tuning Practice Question
A developer needs to fine-tune a 7B LLM for a customer-support chatbot using NVIDIA NeMo. The dataset contains paired instructions and desired responses. Which data format should be used to prepare the dataset for supervised fine-tuning?
⚠ Common exam trap
Watch out — candidates often confuse the training configuration file with the training dataset, or assuming tokenized sequences alone are sufficient without input-output pairing.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
A JSONL file where each line contains an input text field and an output text field representing the instruction-response pair.
Supervised fine-tuning trains a model to produce a target response given an input instruction. The dataset must therefore contain aligned input-output pairs, and NeMo's data pipeline accepts JSONL records with distinct input and output fields. Files lacking response labels, pre-tokenized sequences without pairing, or pure configuration files cannot supply the supervision signal the training objective requires.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
A YAML configuration file listing hyperparameters and dataset paths only.
Why it's wrong here
A YAML file is used to configure the training job, including model, optimizer, and data paths, but it does not contain the training examples themselves. The scenario asks what format the dataset should take. Configuration files are necessary infrastructure, yet they cannot substitute for the actual instruction-response records that the model learns from during fine-tuning.
- ✗
A pickle file containing a Python list of tokenized integer sequences with no text labels.
Why it's wrong here
Pre-tokenized integer sequences without associated labels cannot directly serve supervised fine-tuning because the framework needs to align inputs with targets and apply the loss mask. While tokenization happens internally, providing only raw token IDs removes the semantic pairing of instruction and response. NeMo's data pipeline expects structured text fields it can tokenize and format consistently.
- ✓
A JSONL file where each line contains an input text field and an output text field representing the instruction-response pair.
Why this is correct
Supervised fine-tuning in NeMo expects paired examples that map an input prompt to a target completion. A JSONL file with input and output fields per line matches that structure and can be consumed by the NeMo data preprocessing pipeline. This format directly supports the instruction-response objective used for chatbot behavior, making it the appropriate choice for this scenario.
- ✗
A CSV file containing only the raw customer queries with no corresponding responses.
Why it's wrong here
Supervised fine-tuning requires target outputs to compute the loss against. A file containing only queries provides no response labels, so the model cannot learn the desired answer distribution. This format would be suitable for unlabeled pretraining or prompt-only inference, but it cannot drive the instruction-response training objective described in the scenario.
About these practice questions
This NCP-GENL question is part of Courseiva's 352-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official NVIDIA exam blueprint
This NCP-GENL practice question is part of Courseiva's free NVIDIA certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the NCP-GENL exam.