NCP-GENL Prompt Engineering Practice Question
A developer is building a customer support assistant using an NVIDIA NIM microservice for a Llama 3 model. The assistant must always respond in valid JSON with keys 'category' and 'urgency'. The model often returns conversational text instead. Which prompt engineering change most directly enforces the required output format?
⚠ Common exam trap
The trap here is assuming that simply asking the model for JSON or providing examples is enough, when only schema-guided decoding like guided_json can guarantee valid structure.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Append 'Return only JSON with keys category and urgency' and set the NIM request parameter 'guided_json' to the target schema.
Structured output requires both a clear instruction and a decoding-time constraint. NVIDIA NIM supports guided_json, which restricts token generation to a supplied JSON schema, guaranteeing valid keys and syntax. Pairing that with an explicit prompt instruction aligns the model's behavior with the schema. Few-shot examples or temperature changes alone cannot guarantee strict JSON compliance.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✓
Append 'Return only JSON with keys category and urgency' and set the NIM request parameter 'guided_json' to the target schema.
Why this is correct
Combining an explicit instruction with NVIDIA NIM's guided_json parameter constrains decoding to the provided JSON schema, so the model cannot emit conversational text. The instruction aligns the model's intent while guided_json enforces structural validity at generation time. This is the most direct way to guarantee the required keys and format in the response.
- ✗
Add 'Do not hallucinate' to the prompt and lower the max_tokens parameter to 50.
Why it's wrong here
Instructing the model not to hallucinate does not enforce output structure, and limiting max_tokens can truncate a valid JSON object before it closes. Neither technique binds the model to the 'category' and 'urgency' keys. The response may remain conversational or become incomplete, so this does not satisfy the formatting requirement.
- ✗
Add a system prompt that says 'You are a helpful assistant' and increase the temperature to 0.9.
Why it's wrong here
A generic system persona does not constrain the output structure, and raising temperature to 0.9 increases randomness, making JSON-only responses less likely. This approach may produce more varied natural language, which directly conflicts with the need for strict key-value JSON output. It fails because it neither specifies the schema nor reduces sampling variability.
- ✗
Use few-shot examples of JSON outputs and set top_p to 0.1 without any schema constraint.
Why it's wrong here
Few-shot examples and low top_p can nudge the model toward JSON-like output, but they do not guarantee valid JSON because decoding remains unconstrained. The model could still add prose, omit a key, or produce syntactically invalid JSON. Without a schema-enforced decoding mechanism, the format requirement is only probabilistic, not reliable.
About these practice questions
This NCP-GENL question is part of Courseiva's 352-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official NVIDIA exam blueprint
This NCP-GENL practice question is part of Courseiva's free NVIDIA certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the NCP-GENL exam.