NCP-GENL Data Preparation Practice Question
You are preparing a dataset of customer reviews for fine-tuning an LLM to generate concise summaries. The reviews are in multiple languages, but the target summaries must be in English. You have a limited budget for translation. Which data preparation step is most critical to ensure the fine-tuned model produces high-quality English summaries?
⚠ Common exam trap
The trap here is assuming that translating everything to English first is simpler, but it fails to train the model for multilingual input and may introduce translation artifacts.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Ensure that each training example pairs a review in its original language with a high-quality English summary.
Pairing original-language reviews with English summaries directly trains the model for cross-lingual summarization, which is the end goal. This avoids translation errors and preserves the original semantics. It also prepares the model for real-world multilingual inputs, ensuring it can generate English summaries without an intermediate translation step.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Translate all reviews into English and then train the model to summarize English text.
Why it's wrong here
This reduces the task to monolingual summarization, which may not generalize to multilingual inputs at inference. It also incurs translation costs and potential errors. The model would not learn to handle non-English input directly, limiting its utility.
- ✓
Ensure that each training example pairs a review in its original language with a high-quality English summary.
Why this is correct
This approach directly trains the model to perform cross-lingual summarization: input in any language, output in English. It leverages the original text without translation errors and teaches the model to generate English summaries regardless of source language. This is the most effective strategy for the stated goal.
- ✗
Use a multilingual LLM to translate all non-English reviews into English before training.
Why it's wrong here
Machine translation can introduce errors and lose nuances, and the model may learn translation artifacts rather than summarization. It also increases cost and may not align with the goal of summarizing original content. Better to train on original language with English summaries if possible.
- ✗
Filter the dataset to include only reviews originally written in English.
Why it's wrong here
Filtering to English-only reduces dataset diversity and may not reflect the multilingual input the model will encounter in production. It also discards potentially valuable training signals. The requirement is to summarize any review into English, so multilingual data is beneficial.
About these practice questions
Courseiva writes every NCP-GENL question from scratch — 352 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official NVIDIA exam blueprint
This NCP-GENL practice question is part of Courseiva's free NVIDIA certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the NCP-GENL exam.