Courseiva

NCA-GENL · domain

Software Development

This domain covers building and shipping generative AI applications with NVIDIA tooling: integrating LLMs into production services, defending against prompt injection, quantizing models, and tuning inference. Questions are scenario-based, asking you to pick the right software engineering pattern, TensorRT-LLM optimization, or deployment configuration setting for a described production problem.

77 questions18 easy33 medium26 hard

Focused practice

Practice Software Development questions

Scored sessions drawing only from this domain — pick a length below.

Start 20-question practice test →

What this domain covers

What to know about Software Development

Be able to design a production LLM integration that resists prompt injection, select quantization precision from accuracy and latency needs, and tune a TensorRT-LLM RAG deployment. The most important thing is correctly separating untrusted input from system instructions and diagnosing where latency actually occurs.

Applying input validation and instruction/data separation patterns to block prompt injection in LLM services

Choosing FP16 versus INT8 quantization based on accuracy, latency, memory, and hardware support

Optimizing RAG retrieval latency using NVIDIA TensorRT-LLM and efficient embedding/vector search

Interpreting TensorRT-LLM deployment JSON options such as enable_cuda_graph for inference tuning

Watch out for

Common Software Development exam traps

  • ▸Treating prompt injection as a model problem instead of using software patterns like delimiting untrusted input and validating outputs
  • ▸Assuming INT8 always beats FP16, ignoring accuracy loss, calibration effort, and GPU support for the chosen precision
  • ▸Blaming generation for RAG latency when the bottleneck is document retrieval, embedding, or vector search before the LLM runs

Question index

All Software Development questions (77)

Click any question to see the full explanation, or start a practice session above.

1

A developer wants an LLM application to answer questions about an internal knowledge base that changes daily. Rather than retraining the model, they plan to retrieve relevant passages at query time and place them into the prompt. Which approach are they implementing?

Easy
2

A developer is writing an application that calls an NVIDIA-hosted LLM endpoint and needs to keep multi-turn context across several user messages. Which payload structure should the application send to the chat completions API?

Easy
3

A developer is packaging a generative AI application that must run inference on-premises with NVIDIA GPUs and also expose an OpenAI-compatible HTTP API so existing client code works unchanged. Which two components should the developer use together to meet these requirements? (Choose two.)

Medium
4

A developer needs to ensure that an LLM application remains deterministic across multiple runs. Which parameter configuration is most effective?

Medium
5

A team is building an internal document assistant and wants the model to answer only from an approved corpus of HR policy PDFs. They will deploy the model with NVIDIA NIM and control grounding at generation time by injecting retrieved passages into the prompt. Which parameter combination in the NIM chat completions request best enforces this grounding while keeping responses deterministic for audit logs?

Medium
6

Which component in the NVIDIA AI Enterprise stack is specifically designed to orchestrate the lifecycle of multi-model deployments on Kubernetes?

Medium
7

A team is serving a 70B-parameter LLM with TensorRT-LLM on a node with four GPUs. During load testing they observe that increasing concurrent requests improves throughput up to a point, then latency spikes sharply and GPU memory utilization sits near the limit. Profiling shows the KV cache is being paged out and recomputed. Which change most directly addresses this bottleneck?

Hard
8

What is the primary function of the NVIDIA NGC (NVIDIA GPU Cloud) registry in the software development lifecycle for Generative AI?

Easy
9

A developer is building a document-summarization service on NVIDIA NIM for LLMs and wants to stream partial tokens to the client while the model is still generating. The NIM endpoint exposes an OpenAI-compatible /chat/completions route. Which request parameter should the developer set to receive incremental token deltas rather than one complete response body?

Medium
10

A developer is deploying a large language model using NVIDIA TensorRT-LLM and wants to optimize inference for a production environment with limited GPU memory. Which two techniques can be used to reduce memory footprint while maintaining acceptable performance? (Choose two.)

Hard
11

A developer is using the NVIDIA NeMo framework to fine-tune a large language model. They want to reduce GPU memory usage during training without significantly sacrificing model quality. Which technique should they apply?

Medium
12

A team is serving a 70B-parameter model with NVIDIA Triton Inference Server and TensorRT-LLM. Under concurrent load, GPU memory is exhausted because each request reserves its own large KV cache. Which Triton feature should the team enable to share KV cache blocks across requests that have common prompt prefixes?

Hard
13

When fine-tuning a Large Language Model using PEFT (Parameter-Efficient Fine-Tuning) techniques like LoRA, what is the primary technical advantage being leveraged?

Easy
14

When designing a scalable inference microservice, which factor most significantly impacts the 'Time to First Token' (TTFT) for concurrent users?

Medium
15

A developer needs to serve a quantized Llama model on an NVIDIA GPU and wants the runtime to automatically select the fastest available execution kernels for the detected GPU architecture. Which approach aligns with the NVIDIA inference stack for this requirement?

Easy
16

Refer to the exhibit. A developer wants to make the model's output more deterministic and focused on highly probable tokens. Which change should be made to the configuration policy?

Hard
17

Refer to the exhibit. Given the current configuration, what is the primary risk during high-traffic bursts?

Hard
18

When integrating an LLM into a production application, you must protect against prompt injection. Which software engineering pattern is most effective for this purpose?

Hard
19

An application requires streaming responses from a deployed LLM. Which communication protocol is most suitable for minimizing latency and ensuring efficient data delivery in a real-time generative AI application?

Medium
20

Which NVIDIA framework is specifically designed to facilitate the deployment of optimized LLMs as microservices with standardized APIs?

Easy
21

A developer is building a retrieval-augmented generation service and needs to embed millions of document chunks and run low-latency similarity search over them on GPU. They want a library that handles both index construction and search with GPU acceleration. Which NVIDIA component should they use?

Medium
22

When integrating an LLM into an application using NVIDIA API endpoints, what is the primary purpose of the 'System' role in the messages payload?

Easy
23

When using NVIDIA Riva for speech-to-text applications, which component provides the real-time transcription service based on streaming audio inputs?

Easy
24

Which THREE factors should a developer consider when choosing between FP16 and INT8 quantization for a production LLM deployment?

Medium
25

During development of a RAG application using NVIDIA NeMo Guardrails, why is it important to define specific 'canonical forms' in the configuration?

Medium
26

When developing with NVIDIA NeMo, which component is primarily responsible for scaling the training of massive LLMs across multiple GPU nodes?

Easy
27

What is the primary purpose of using a Model Repository in the NVIDIA Triton Inference Server architecture?

Medium
28

A developer is writing an application that streams chat completions from an NVIDIA-hosted NIM endpoint. Users report that the interface freezes until the entire answer is ready, even though the endpoint supports token streaming. Which client-side change fixes the perceived latency?

Easy
29

A developer is packaging a NeMo-based LLM application into a container for deployment on an NVIDIA GPU node. Which two practices are required to ensure the container can access the GPU and run inference efficiently? (Choose two.)

Medium
30

A developer is writing a Python client for an NVIDIA-hosted LLM endpoint and needs the model to answer every request in a strict JSON schema without extra prose. Where should the formatting contract be expressed so it applies consistently across all requests from the service?

Medium
31

Refer to the exhibit. What is the technical implication of using the specified 'fp8' precision mode in this model configuration?

Hard
32

A developer is deploying a TensorRT-LLM optimized model on NVIDIA Triton Inference Server. They observe that the first inference request takes significantly longer than subsequent ones. Which Triton feature should they configure to reduce this initial latency?

Hard
33

Which NVIDIA SDK is specifically optimized for high-performance deep learning inference and supports the deployment of quantized models?

Easy
34

When implementing a Guardrails layer in a generative AI application, what is the primary goal regarding model output?

Medium
35

A developer is tuning a retrieval-augmented generation pipeline that uses NVIDIA NIM embeddings and a NIM LLM. Latency is dominated by embedding thousands of document chunks at query time because the team re-embeds the whole corpus on every request. Which change most directly fixes the architecture?

Medium
36

Which TWO of the following are primary benefits of using NVIDIA Triton Inference Server for deploying generative AI models?

Medium
37

A developer is using NVIDIA Triton Inference Server to deploy a TensorRT-LLM optimized model. The model must support multiple concurrent users with low latency. The developer notices that latency spikes when many requests arrive simultaneously. Which Triton feature should be configured to improve throughput while maintaining acceptable latency?

Hard
38

A developer is preparing a container for an LLM microservice that will run on an NVIDIA GPU node and must be deployable through NVIDIA NIM. They want the image to be portable across supported GPU generations while still using NVIDIA's optimized inference stack. Which two practices should they follow? (Choose two.)

Hard
39

A team fine-tunes a Llama-3 8B model with NVIDIA NeMo Framework and must ship an inference artifact that a C++ service can load without a Python runtime. They want maximum throughput on Hopper GPUs and plan to serve many concurrent requests with in-flight batching. Which artifact and runtime pairing best satisfies these constraints?

Hard
40

A developer is profiling a TensorRT-LLM serving deployment and notices that throughput collapses once concurrent requests exceed a small number of users, even though GPU compute utilization stays low. The model uses paged KV cache and continuous batching. Which factor most likely explains the bottleneck?

Hard
41

A developer is optimizing a retrieval-augmented generation (RAG) pipeline using NVIDIA TensorRT-LLM. They notice excessive latency during the document retrieval phase before the generation starts. Which optimization strategy is most effective for this bottleneck?

Medium
42

A developer is building a document summarization service using an NVIDIA NIM microservice for Llama-3. The service must process batches of 20 documents at once to maximize throughput. The NIM container is already running with default settings. Which API parameter should the developer configure to enable efficient batched inference?

Medium
43

A developer is using the NVIDIA API Catalog to test a Llama-3 model via its API endpoint. They need to send a request that includes a system prompt to set the model's behavior. Which component of the request payload is used to provide the system prompt?

Easy
44

A developer is using TensorRT-LLM to build a chatbot and wants to reduce the memory footprint of the KV cache during inference. Which technique should they use?

Hard
45

Which NVIDIA technology enables efficient cross-GPU communication during Tensor Parallelism for large-scale model inference?

Medium
46

Which TWO of the following practices are recommended when using NVIDIA Triton Inference Server to maximize throughput for a concurrent multi-model deployment?

Hard
47

Refer to the exhibit. An engineer receives this error during deployment. What is the most likely cause?

Hard
48

A developer is containerizing an inference service built with TensorRT-LLM and NVIDIA NIM for a Kubernetes cluster. They want the deployment to start reliably and use the GPU efficiently. Which two practices should they follow? (Choose two.)

Hard
49

A team fine-tunes a Llama model with NVIDIA NeMo Framework and must serve it behind an OpenAI-compatible endpoint with no Python glue code. They want the adapter weights kept separate from the base model so several adapters can share one loaded base. Which deployment approach fits these constraints?

Hard
50

Refer to the exhibit. An engineer is tuning a deployment config. Why is 'enable_cuda_graph' set to true in this JSON configuration?

Medium
51

A developer is debugging a RAG service where answers are correct in testing but degrade in production as the document corpus grows. Logs show retrieval returning chunks with high similarity scores that do not contain the answer. Which change most directly addresses the root cause?

Hard
52

A team is building a RAG assistant and wants to reduce hallucinated citations. They plan to have the LLM return structured output that names the source document chunk used for each claim. Which implementation strategy most directly improves the reliability of that structured output?

Hard
53

A developer is building a customer-support assistant on NVIDIA NIM microservices. After a model update, responses that previously arrived in under 300 ms now take over two seconds, and the streaming client shows a long pause before the first token appears. GPU utilization is low and the prompt template was not changed. Which action should the developer take first to diagnose the regression?

Medium
54

When deploying a Large Language Model using TensorRT-LLM, which TWO configuration factors must be tuned to maximize KV cache efficiency?

Hard
55

A developer is tuning a TensorRT-LLM deployment of a long-context chat model and observes that GPU memory is exhausted under concurrent requests, causing requests to be rejected. They want to reduce KV cache memory pressure without retraining the model. (Choose two.)

Hard
56

A developer is using the NVIDIA API Catalog to experiment with a hosted LLM. They want to send a prompt and receive a completion. Which endpoint should they use?

Easy
57

When utilizing NVIDIA NIM for deployment, why is it recommended to use a containerized environment?

Medium
58

A developer is evaluating a fine-tuned LLM with NVIDIA NeMo and observes that evaluation loss keeps decreasing while downstream task accuracy plateaus and then declines. Which action should the developer take to address this?

Hard
59

A developer is packaging a fine-tuned Llama 3 model as a TensorRT-LLM engine for an on-premises inference service. The model was trained with a custom tokenizer that adds four new special tokens beyond the base vocabulary. When the engine is built and the service is started, the model outputs garbled text and repeats the same fragment regardless of the prompt. Which action should the developer take to resolve this?

Medium
60

A team wants to load and run an optimized quantized LLM entirely inside a Python application with minimal dependencies, using a single high-level API that handles engine building and generation. They are not deploying a network service. Which component of the NVIDIA software stack is designed for this use case?

Easy
61

A developer has a working TensorRT-LLM engine and wants to expose it through NVIDIA Triton Inference Server so that multiple client applications can call it over HTTP and gRPC with a stable interface. Which Triton feature should they configure to serve the TensorRT-LLM engine as a backend?

Easy
62

Refer to the exhibit. If a developer increases the 'max_batch_size' in the JSON configuration, what is the primary expected trade-off in the system's performance metrics?

Medium
63

A developer is integrating a NeMo Guardrails configuration into an existing chatbot. They need to ensure that the LLM does not generate content related to unauthorized financial advice. Which mechanism should they implement to achieve this programmatic constraint?

Medium
64

A developer is preparing a RAG service that calls an NVIDIA-hosted LLM endpoint and must reduce hallucinations for questions whose answers are absent from the retrieved context. Which two practices should be applied in the application layer? (Choose two.)

Medium
65

A developer is packaging a generative AI application for NVIDIA AI Enterprise deployment on Kubernetes. The application must run an LLM served by NVIDIA NIM, an embedding model, and a vector database, and must support rolling upgrades without dropping in-flight inference requests. Which design choice best meets these requirements?

Hard
66

A developer is writing a Python service that calls an NVIDIA-hosted NIM endpoint for a Llama model. The service must recover gracefully when the endpoint returns HTTP 429 responses during peak traffic, without dropping user requests. Which implementation approach best satisfies this requirement?

Medium
67

A developer is writing a Python service that calls a locally hosted NVIDIA NIM microservice for a Llama model. They want to keep the client code portable so the same class can later target NVIDIA's hosted API endpoints without rewrites. Which client approach fits this goal?

Easy
68

A developer is building a generative AI application that uses an NVIDIA NIM microservice for a Llama 3 model. They need to persist the model's responses and associated metadata for later auditing. Which approach best integrates NIM with an external datastore?

Medium
69

When evaluating an LLM for a domain-specific task, why is 'Few-Shot Prompting' often superior to 'Zero-Shot Prompting'?

Medium
70

A developer is building a customer support chatbot using NVIDIA NIM microservices. The application must reliably return structured JSON containing 'intent' and 'confidence' fields for downstream ticket routing. Which approach should the developer use to constrain the model's output format?

Medium
71

Which memory management strategy in TensorRT-LLM is specifically designed to minimize fragmentation and allow for efficient KV cache allocation in multi-user environments?

Hard
72

A developer is building a customer support chatbot using NVIDIA NIM microservices. They need the model to always respond in a formal tone and never mention competitor products. Where should these directives be placed in the API request to ensure consistent behavior across all user interactions?

Medium
73

A developer is building a customer-support assistant that must retrieve answers only from an approved internal knowledge base and cite the source document for each reply. They are using NVIDIA NIM microservices for the LLM and an embedding model, and they need the application layer to enforce citation behavior and reject answers that are not grounded in retrieved passages. Which software development approach best enforces this grounding requirement?

Medium
74

A developer is using the NVIDIA Triton Inference Server to deploy a TensorRT-LLM optimized model. They want to send a request with multiple prompts to be processed in a single inference call. Which Triton feature should they use?

Easy
75

A developer is building a Retrieval-Augmented Generation (RAG) pipeline using NVIDIA NIM microservices. They need to ensure that the retriever returns the most relevant documents for a given query. Which two components should they optimize? (Choose two.)

Hard
76

A developer is debugging a TensorRT-LLM generation that intermittently produces truncated responses when many users submit long prompts concurrently. Logs show requests completing without errors, but outputs stop mid-sentence. Which configuration change is most likely to resolve this?

Hard
77

When deploying an LLM, what is the 'Time to First Token' (TTFT) metric used to measure?

Easy

Frequently asked questions

What does the Software Development domain cover on the NCA-GENL exam?
Be able to design a production LLM integration that resists prompt injection, select quantization precision from accuracy and latency needs, and tune a TensorRT-LLM RAG deployment. The most important thing is correctly separating untrusted input from system instructions and diagnosing where latency actually occurs.
How many questions are in this domain?
This page lists all 77 Software Development questions in the NCA-GENL question bank. The actual exam draws from this domain proportionally to its weighting in the official exam blueprint.
What is the best way to practise this domain?
Start with a short focused session (10 questions) to identify gaps, then work through explanations. Repeat with a longer session once the weak areas feel solid.
Can I practise only Software Development questions?
Yes — the session launcher on this page filters questions to this domain only. Choose any session length for inline explanations and scoring.
nvidia-nca-genl NVIDIA-NCA-GENL nca software development Practice Questions