Courseiva

NCA-GENL · topic practice

Software Development practice questions

This domain covers building and shipping generative AI applications with NVIDIA tooling: integrating LLMs into production services, defending against prompt injection, quantizing models, and tuning inference. Questions are scenario-based, asking you to pick the right software engineering pattern, TensorRT-LLM optimization, or deployment configuration setting for a described production problem.

Courseiva uses original exam-style practice questions designed for learning and revision. The goal is to understand the concepts, recognise exam patterns, and improve through explanations — not memorise copied exam dumps.

Editorial oversight:Johnson Ajibi· MSc IT Security, IEEE Senior Member
20 questionsDomain: Software Development

What the exam tests

What to know about Software Development

Be able to design a production LLM integration that resists prompt injection, select quantization precision from accuracy and latency needs, and tune a TensorRT-LLM RAG deployment. The most important thing is correctly separating untrusted input from system instructions and diagnosing where latency actually occurs.

Applying input validation and instruction/data separation patterns to block prompt injection in LLM services

Choosing FP16 versus INT8 quantization based on accuracy, latency, memory, and hardware support

Optimizing RAG retrieval latency using NVIDIA TensorRT-LLM and efficient embedding/vector search

Interpreting TensorRT-LLM deployment JSON options such as enable_cuda_graph for inference tuning

Watch out for

Common Software Development exam traps

  • ▸Treating prompt injection as a model problem instead of using software patterns like delimiting untrusted input and validating outputs
  • ▸Assuming INT8 always beats FP16, ignoring accuracy loss, calibration effort, and GPU support for the chosen precision
  • ▸Blaming generation for RAG latency when the bottleneck is document retrieval, embedding, or vector search before the LLM runs

Practice set

Software Development questions

20 questions · select your answer, then reveal the explanation

A developer is optimizing a retrieval-augmented generation (RAG) pipeline on NVIDIA NIM microservices. Which approach best minimizes latency while maintaining context relevance?

Refer to the exhibit. A developer notices that the inference performance is suboptimal. Based on the provided configuration, what is the most likely cause?

Exhibit

{
  "model": "llama-3-8b-instruct",
  "max_tokens": 512,
  "temperature": 0.7,
  "n_gpu_layers": 40,
  "stream": true
}

Which THREE techniques are standard for reducing hallucinations in an NVIDIA-powered RAG application?

Which component is responsible for converting raw input text into numerical embeddings suitable for an NVIDIA LLM?

Which TWO metrics are most critical for monitoring the health of a production-grade LLM inference service?

Refer to the exhibit. What is the primary benefit of using this specific combination of technologies for inference?

Exhibit

{
  "engine": "tensorrt-llm",
  "quantization": "awq",
  "batching": "continuous",
  "kv_cache_type": "paged"
}

A developer is optimizing a PyTorch model for NVIDIA TensorRT inference. Which step is essential to ensure the model benefits from FP16 precision kernels during the conversion process?

Refer to the exhibit. A developer observes that GPU utilization is low despite high request volume. Based on the provided Triton configuration, what is the most likely cause?

Exhibit

config.pbtxt: 
instance_group [
  {
    count: 2
    kind: KIND_GPU
  }
]

Which THREE steps are required when converting a standard Hugging Face model for use with the NVIDIA TensorRT-LLM library?

Which THREE conditions are necessary to enable TensorRT-LLM's 'In-flight Batching' feature?

Which technique should a developer use to reduce the memory footprint of an LLM while maintaining maximum compatibility with NVIDIA inference servers?

When evaluating an LLM's readiness for production, what role do 'benchmarks' play in the software development lifecycle?

Refer to the exhibit. An engineer is deploying a large language model using TensorRT-LLM. Despite the configuration, the deployment fails with an Out-of-Memory (OOM) error on the target nodes. Which adjustment is most appropriate to resolve this?

Exhibit

model_config.json: { 'quantization': 'int8_sq', 'tensor_parallel': 4, 'pipeline_parallel': 2, 'max_batch_size': 128 }

Which THREE techniques are commonly used to optimize inference latency for LLMs on NVIDIA GPUs?

Which technique allows for the concurrent execution of multiple model instances on a single NVIDIA GPU to maximize throughput?

A software engineer is optimizing a RAG pipeline using NVIDIA TensorRT-LLM. The model is currently experiencing latency bottlenecks during the prefill phase. Which action effectively addresses this?

An application developer needs to deploy a custom LoRA adapter for a Llama-3 model using NVIDIA Triton Inference Server. Which configuration ensures the model and adapter are loaded efficiently?

A developer is building an application with NeMo Framework and requires fine-tuning. Which TWO of the following steps are mandatory for setting up a distributed training environment?

A developer is optimizing a retrieval-augmented generation (RAG) pipeline. Which TWO factors significantly impact the latency of the retrieval phase in a production NVIDIA NIM deployment?

Refer to the exhibit. A developer encounters the following error during a RAG pipeline execution. Which action should the developer take to resolve this error while maintaining retrieval quality?

Exhibit

ERROR: [NIM_SERVER] Request failed. Context window exceeded for model: meta-llama-3-70b-instruct. Prompt + Retrieved Context length: 132,044 tokens. Max tokens: 128,000.

Free account

Track your progress over time

Create a free account to save your results and see which topics improve across sessions.

Focused Software Development sessions

Start a Software Development only practice session

Every question in these sessions is drawn from the Software Development domain — nothing else.

Related practice questions

Related NCA-GENL topic practice pages

Move into related areas when this topic feels solid.

Frequently asked questions

What does the NCA-GENL exam test about Software Development?
Be able to design a production LLM integration that resists prompt injection, select quantization precision from accuracy and latency needs, and tune a TensorRT-LLM RAG deployment. The most important thing is correctly separating untrusted input from system instructions and diagnosing where latency actually occurs.
How should I use these practice questions?
Select your answer before revealing the explanation. Then read why each option is right or wrong — this active recall approach builds retention far faster than re-reading notes.
Can I practise just Software Development questions in a focused session?
Yes — the session launcher on this page draws every question from the Software Development domain. Use a 10-question session first to gauge your baseline, then move to 20 or 30 once the weak spots are clear.
Where can I practise other NCA-GENL topics?
Use the topic links above to move to related areas, or go back to the NCA-GENL question bank to see all topics.
Are these real exam questions or dumps?
These are original practice questions written to test the same concepts the NCA-GENL exam covers. They are not copied from any real exam or dump site.