NCA-GENL Software Development Practice Question
Exhibit
config_policy: { 'max_tokens': 1024, 'temperature': 0.7, 'top_p': 0.9, 'presence_penalty': 0.0 }Refer to the exhibit. A developer wants to make the model's output more deterministic and focused on highly probable tokens. Which change should be made to the configuration policy?
⚠ Common exam trap
Candidates often confuse the effects of temperature and top-p, sometimes suggesting an increase in these values when the goal is to make the model output more predictable and focused.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Decrease temperature to 0.1 and top_p to 0.5.
Determinism in LLMs is controlled by the temperature and top-p sampling parameters. Lowering temperature reduces the randomness in the probability distribution, while lowering top-p restricts the sampling pool to a smaller subset of high-probability tokens. Adjusting these parameters is vital for applications requiring factual consistency and stable outputs, as they directly dictate the sampling strategy used during the token generation process in the inference engine.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Increase temperature to 1.5 and top_p to 1.0.
Why it's wrong here
Increasing the temperature to 1.5 makes the model more creative and random, leading to higher variance in outputs. Increasing top_p to 1.0 allows the model to sample from the entire probability distribution, which is the opposite of the desired goal for increasing output determinism.
- ✓
Decrease temperature to 0.1 and top_p to 0.5.
Why this is correct
A lower temperature of 0.1 makes the probability distribution sharper, favoring the most likely tokens. A lower top_p of 0.5 further restricts the model to only the most probable candidates, resulting in highly deterministic and focused outputs that are ideal for consistent, repetitive, or factual tasks.
- ✗
Set presence_penalty to 1.0.
Why it's wrong here
The presence_penalty parameter discourages the model from repeating tokens that have already appeared. While this affects the variety of the text, it does not directly control the determinism or the probability distribution of the next token, which is the primary driver of output stability.
- ✗
Increase max_tokens to 4096.
Why it's wrong here
Increasing the maximum number of tokens allows for longer responses but has no effect on the sampling strategy or the determinism of the generated text. The length of the output is independent of the probability distribution used to pick individual tokens during the generation process.
About these practice questions
This NCA-GENL question is part of Courseiva's 367-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official NVIDIA exam blueprint
This NCA-GENL practice question is part of Courseiva's free NVIDIA certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the NCA-GENL exam.