Courseiva
Prompt Engineering →hardMultiple Choice

NCP-GENL Prompt Engineering Practice Question

A developer is using an NVIDIA NIM for a Llama 3.1 70B model to build a legal document review assistant. The model must answer questions based on a provided contract, but the contracts are often 50,000 tokens long, exceeding the model's 8,000-token context window. Which prompt engineering strategy is most appropriate to handle this constraint?

⚠ Common exam trap

The trap here is assuming that fine-tuning or summarization can easily overcome context window limits, when in fact retrieval-based approaches are the practical solution.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Use a retrieval-augmented generation (RAG) approach to fetch only the relevant sections of the contract and include them in the prompt.

RAG is the most appropriate strategy because it retrieves only the relevant sections of the long contract, fitting them into the model's context window. This allows the model to answer questions accurately without needing to process the entire document. It is a widely adopted prompt engineering pattern for handling context length limitations.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✓

    Use a retrieval-augmented generation (RAG) approach to fetch only the relevant sections of the contract and include them in the prompt.

    Why this is correct

    RAG involves retrieving the most relevant passages from the long document and including only those in the prompt. This reduces the context length to fit within the model's window while still providing the necessary information. It is the standard solution for handling documents that exceed the context limit.

  • ✗

    Split the contract into chunks and ask the model to summarize each chunk sequentially, then combine the summaries.

    Why it's wrong here

    Sequential summarization can lose important details and may not preserve the ability to answer specific questions. It is a form of hierarchical summarization, but it is not as precise as RAG for question answering. RAG directly retrieves the relevant information for the query, which is more efficient and accurate.

  • ✗

    Increase the model's context window by fine-tuning it on longer sequences.

    Why it's wrong here

    Fine-tuning to extend the context window is possible but requires significant resources and may not be supported by the NIM microservice. It is not a prompt engineering strategy. RAG is a more practical and immediate solution that does not require retraining.

  • ✗

    Use a sliding window approach where the model processes the contract in overlapping segments and maintains a memory of previous segments.

    Why it's wrong here

    Sliding window with memory is complex to implement and may still exceed the context window when maintaining memory. It is not a standard prompt engineering technique. RAG is simpler and more effective for question answering over long documents, as it only includes the relevant parts.

About these practice questions

One of 352 original NCP-GENL practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official NVIDIA exam blueprint

This NCP-GENL practice question is part of Courseiva's free NVIDIA certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the NCP-GENL exam.