NCA-GENL · domain
Software Development
This domain covers building and shipping generative AI applications with NVIDIA tooling: integrating LLMs into production services, defending against prompt injection, quantizing models, and tuning inference. Questions are scenario-based, asking you to pick the right software engineering pattern, TensorRT-LLM optimization, or deployment configuration setting for a described production problem.
Focused practice
Practice Software Development questions
Scored sessions drawing only from this domain — pick a length below.
Start 20-question practice test →What this domain covers
What to know about Software Development
Be able to design a production LLM integration that resists prompt injection, select quantization precision from accuracy and latency needs, and tune a TensorRT-LLM RAG deployment. The most important thing is correctly separating untrusted input from system instructions and diagnosing where latency actually occurs.
Applying input validation and instruction/data separation patterns to block prompt injection in LLM services
Choosing FP16 versus INT8 quantization based on accuracy, latency, memory, and hardware support
Optimizing RAG retrieval latency using NVIDIA TensorRT-LLM and efficient embedding/vector search
Interpreting TensorRT-LLM deployment JSON options such as enable_cuda_graph for inference tuning
Watch out for
Common Software Development exam traps
- ▸Treating prompt injection as a model problem instead of using software patterns like delimiting untrusted input and validating outputs
- ▸Assuming INT8 always beats FP16, ignoring accuracy loss, calibration effort, and GPU support for the chosen precision
- ▸Blaming generation for RAG latency when the bottleneck is document retrieval, embedding, or vector search before the LLM runs
Question index
All Software Development questions (77)
Click any question to see the full explanation, or start a practice session above.
A developer wants an LLM application to answer questions about an internal knowledge base that changes daily. Rather than retraining the model, they plan to retrieve relevant passages at query time and place them into the prompt. Which approach are they implementing?
Easy2A developer is writing an application that calls an NVIDIA-hosted LLM endpoint and needs to keep multi-turn context across several user messages. Which payload structure should the application send to the chat completions API?
Easy3A developer is packaging a generative AI application that must run inference on-premises with NVIDIA GPUs and also expose an OpenAI-compatible HTTP API so existing client code works unchanged. Which two components should the developer use together to meet these requirements? (Choose two.)
Medium4A developer needs to ensure that an LLM application remains deterministic across multiple runs. Which parameter configuration is most effective?
Medium5A team is building an internal document assistant and wants the model to answer only from an approved corpus of HR policy PDFs. They will deploy the model with NVIDIA NIM and control grounding at generation time by injecting retrieved passages into the prompt. Which parameter combination in the NIM chat completions request best enforces this grounding while keeping responses deterministic for audit logs?
Medium6Which component in the NVIDIA AI Enterprise stack is specifically designed to orchestrate the lifecycle of multi-model deployments on Kubernetes?
Medium7A team is serving a 70B-parameter LLM with TensorRT-LLM on a node with four GPUs. During load testing they observe that increasing concurrent requests improves throughput up to a point, then latency spikes sharply and GPU memory utilization sits near the limit. Profiling shows the KV cache is being paged out and recomputed. Which change most directly addresses this bottleneck?
Hard8What is the primary function of the NVIDIA NGC (NVIDIA GPU Cloud) registry in the software development lifecycle for Generative AI?
Easy9A developer is building a document-summarization service on NVIDIA NIM for LLMs and wants to stream partial tokens to the client while the model is still generating. The NIM endpoint exposes an OpenAI-compatible /chat/completions route. Which request parameter should the developer set to receive incremental token deltas rather than one complete response body?
Medium10A developer is deploying a large language model using NVIDIA TensorRT-LLM and wants to optimize inference for a production environment with limited GPU memory. Which two techniques can be used to reduce memory footprint while maintaining acceptable performance? (Choose two.)
Hard11A developer is using the NVIDIA NeMo framework to fine-tune a large language model. They want to reduce GPU memory usage during training without significantly sacrificing model quality. Which technique should they apply?
Medium12A team is serving a 70B-parameter model with NVIDIA Triton Inference Server and TensorRT-LLM. Under concurrent load, GPU memory is exhausted because each request reserves its own large KV cache. Which Triton feature should the team enable to share KV cache blocks across requests that have common prompt prefixes?
Hard13When fine-tuning a Large Language Model using PEFT (Parameter-Efficient Fine-Tuning) techniques like LoRA, what is the primary technical advantage being leveraged?
Easy14When designing a scalable inference microservice, which factor most significantly impacts the 'Time to First Token' (TTFT) for concurrent users?
Medium15A developer needs to serve a quantized Llama model on an NVIDIA GPU and wants the runtime to automatically select the fastest available execution kernels for the detected GPU architecture. Which approach aligns with the NVIDIA inference stack for this requirement?
Easy16Refer to the exhibit. A developer wants to make the model's output more deterministic and focused on highly probable tokens. Which change should be made to the configuration policy?
Hard17Refer to the exhibit. Given the current configuration, what is the primary risk during high-traffic bursts?
Hard18When integrating an LLM into a production application, you must protect against prompt injection. Which software engineering pattern is most effective for this purpose?
Hard19An application requires streaming responses from a deployed LLM. Which communication protocol is most suitable for minimizing latency and ensuring efficient data delivery in a real-time generative AI application?
Medium20Which NVIDIA framework is specifically designed to facilitate the deployment of optimized LLMs as microservices with standardized APIs?
Easy21A developer is building a retrieval-augmented generation service and needs to embed millions of document chunks and run low-latency similarity search over them on GPU. They want a library that handles both index construction and search with GPU acceleration. Which NVIDIA component should they use?
Medium22When integrating an LLM into an application using NVIDIA API endpoints, what is the primary purpose of the 'System' role in the messages payload?
Easy23When using NVIDIA Riva for speech-to-text applications, which component provides the real-time transcription service based on streaming audio inputs?
Easy24Which THREE factors should a developer consider when choosing between FP16 and INT8 quantization for a production LLM deployment?
Medium25During development of a RAG application using NVIDIA NeMo Guardrails, why is it important to define specific 'canonical forms' in the configuration?
Medium26When developing with NVIDIA NeMo, which component is primarily responsible for scaling the training of massive LLMs across multiple GPU nodes?
Easy27What is the primary purpose of using a Model Repository in the NVIDIA Triton Inference Server architecture?
Medium28A developer is writing an application that streams chat completions from an NVIDIA-hosted NIM endpoint. Users report that the interface freezes until the entire answer is ready, even though the endpoint supports token streaming. Which client-side change fixes the perceived latency?
Easy29A developer is packaging a NeMo-based LLM application into a container for deployment on an NVIDIA GPU node. Which two practices are required to ensure the container can access the GPU and run inference efficiently? (Choose two.)
Medium30A developer is writing a Python client for an NVIDIA-hosted LLM endpoint and needs the model to answer every request in a strict JSON schema without extra prose. Where should the formatting contract be expressed so it applies consistently across all requests from the service?
Medium31Refer to the exhibit. What is the technical implication of using the specified 'fp8' precision mode in this model configuration?
Hard32A developer is deploying a TensorRT-LLM optimized model on NVIDIA Triton Inference Server. They observe that the first inference request takes significantly longer than subsequent ones. Which Triton feature should they configure to reduce this initial latency?
Hard33Which NVIDIA SDK is specifically optimized for high-performance deep learning inference and supports the deployment of quantized models?
Easy34When implementing a Guardrails layer in a generative AI application, what is the primary goal regarding model output?
Medium35A developer is tuning a retrieval-augmented generation pipeline that uses NVIDIA NIM embeddings and a NIM LLM. Latency is dominated by embedding thousands of document chunks at query time because the team re-embeds the whole corpus on every request. Which change most directly fixes the architecture?
Medium36Which TWO of the following are primary benefits of using NVIDIA Triton Inference Server for deploying generative AI models?
Medium37A developer is using NVIDIA Triton Inference Server to deploy a TensorRT-LLM optimized model. The model must support multiple concurrent users with low latency. The developer notices that latency spikes when many requests arrive simultaneously. Which Triton feature should be configured to improve throughput while maintaining acceptable latency?
Hard38A developer is preparing a container for an LLM microservice that will run on an NVIDIA GPU node and must be deployable through NVIDIA NIM. They want the image to be portable across supported GPU generations while still using NVIDIA's optimized inference stack. Which two practices should they follow? (Choose two.)
Hard39A team fine-tunes a Llama-3 8B model with NVIDIA NeMo Framework and must ship an inference artifact that a C++ service can load without a Python runtime. They want maximum throughput on Hopper GPUs and plan to serve many concurrent requests with in-flight batching. Which artifact and runtime pairing best satisfies these constraints?
Hard40A developer is profiling a TensorRT-LLM serving deployment and notices that throughput collapses once concurrent requests exceed a small number of users, even though GPU compute utilization stays low. The model uses paged KV cache and continuous batching. Which factor most likely explains the bottleneck?
Hard41A developer is optimizing a retrieval-augmented generation (RAG) pipeline using NVIDIA TensorRT-LLM. They notice excessive latency during the document retrieval phase before the generation starts. Which optimization strategy is most effective for this bottleneck?
Medium42A developer is building a document summarization service using an NVIDIA NIM microservice for Llama-3. The service must process batches of 20 documents at once to maximize throughput. The NIM container is already running with default settings. Which API parameter should the developer configure to enable efficient batched inference?
Medium43A developer is using the NVIDIA API Catalog to test a Llama-3 model via its API endpoint. They need to send a request that includes a system prompt to set the model's behavior. Which component of the request payload is used to provide the system prompt?
Easy44A developer is using TensorRT-LLM to build a chatbot and wants to reduce the memory footprint of the KV cache during inference. Which technique should they use?
Hard45Which NVIDIA technology enables efficient cross-GPU communication during Tensor Parallelism for large-scale model inference?
Medium46Which TWO of the following practices are recommended when using NVIDIA Triton Inference Server to maximize throughput for a concurrent multi-model deployment?
Hard47Refer to the exhibit. An engineer receives this error during deployment. What is the most likely cause?
Hard48A developer is containerizing an inference service built with TensorRT-LLM and NVIDIA NIM for a Kubernetes cluster. They want the deployment to start reliably and use the GPU efficiently. Which two practices should they follow? (Choose two.)
Hard49A team fine-tunes a Llama model with NVIDIA NeMo Framework and must serve it behind an OpenAI-compatible endpoint with no Python glue code. They want the adapter weights kept separate from the base model so several adapters can share one loaded base. Which deployment approach fits these constraints?
Hard50Refer to the exhibit. An engineer is tuning a deployment config. Why is 'enable_cuda_graph' set to true in this JSON configuration?
Medium51A developer is debugging a RAG service where answers are correct in testing but degrade in production as the document corpus grows. Logs show retrieval returning chunks with high similarity scores that do not contain the answer. Which change most directly addresses the root cause?
Hard52A team is building a RAG assistant and wants to reduce hallucinated citations. They plan to have the LLM return structured output that names the source document chunk used for each claim. Which implementation strategy most directly improves the reliability of that structured output?
Hard53A developer is building a customer-support assistant on NVIDIA NIM microservices. After a model update, responses that previously arrived in under 300 ms now take over two seconds, and the streaming client shows a long pause before the first token appears. GPU utilization is low and the prompt template was not changed. Which action should the developer take first to diagnose the regression?
Medium54When deploying a Large Language Model using TensorRT-LLM, which TWO configuration factors must be tuned to maximize KV cache efficiency?
Hard55A developer is tuning a TensorRT-LLM deployment of a long-context chat model and observes that GPU memory is exhausted under concurrent requests, causing requests to be rejected. They want to reduce KV cache memory pressure without retraining the model. (Choose two.)
Hard56A developer is using the NVIDIA API Catalog to experiment with a hosted LLM. They want to send a prompt and receive a completion. Which endpoint should they use?
Easy57When utilizing NVIDIA NIM for deployment, why is it recommended to use a containerized environment?
Medium58A developer is evaluating a fine-tuned LLM with NVIDIA NeMo and observes that evaluation loss keeps decreasing while downstream task accuracy plateaus and then declines. Which action should the developer take to address this?
Hard59A developer is packaging a fine-tuned Llama 3 model as a TensorRT-LLM engine for an on-premises inference service. The model was trained with a custom tokenizer that adds four new special tokens beyond the base vocabulary. When the engine is built and the service is started, the model outputs garbled text and repeats the same fragment regardless of the prompt. Which action should the developer take to resolve this?
Medium60A team wants to load and run an optimized quantized LLM entirely inside a Python application with minimal dependencies, using a single high-level API that handles engine building and generation. They are not deploying a network service. Which component of the NVIDIA software stack is designed for this use case?
Easy61A developer has a working TensorRT-LLM engine and wants to expose it through NVIDIA Triton Inference Server so that multiple client applications can call it over HTTP and gRPC with a stable interface. Which Triton feature should they configure to serve the TensorRT-LLM engine as a backend?
Easy62Refer to the exhibit. If a developer increases the 'max_batch_size' in the JSON configuration, what is the primary expected trade-off in the system's performance metrics?
Medium63A developer is integrating a NeMo Guardrails configuration into an existing chatbot. They need to ensure that the LLM does not generate content related to unauthorized financial advice. Which mechanism should they implement to achieve this programmatic constraint?
Medium64A developer is preparing a RAG service that calls an NVIDIA-hosted LLM endpoint and must reduce hallucinations for questions whose answers are absent from the retrieved context. Which two practices should be applied in the application layer? (Choose two.)
Medium65A developer is packaging a generative AI application for NVIDIA AI Enterprise deployment on Kubernetes. The application must run an LLM served by NVIDIA NIM, an embedding model, and a vector database, and must support rolling upgrades without dropping in-flight inference requests. Which design choice best meets these requirements?
Hard66A developer is writing a Python service that calls an NVIDIA-hosted NIM endpoint for a Llama model. The service must recover gracefully when the endpoint returns HTTP 429 responses during peak traffic, without dropping user requests. Which implementation approach best satisfies this requirement?
Medium67A developer is writing a Python service that calls a locally hosted NVIDIA NIM microservice for a Llama model. They want to keep the client code portable so the same class can later target NVIDIA's hosted API endpoints without rewrites. Which client approach fits this goal?
Easy68A developer is building a generative AI application that uses an NVIDIA NIM microservice for a Llama 3 model. They need to persist the model's responses and associated metadata for later auditing. Which approach best integrates NIM with an external datastore?
Medium69When evaluating an LLM for a domain-specific task, why is 'Few-Shot Prompting' often superior to 'Zero-Shot Prompting'?
Medium70A developer is building a customer support chatbot using NVIDIA NIM microservices. The application must reliably return structured JSON containing 'intent' and 'confidence' fields for downstream ticket routing. Which approach should the developer use to constrain the model's output format?
Medium71Which memory management strategy in TensorRT-LLM is specifically designed to minimize fragmentation and allow for efficient KV cache allocation in multi-user environments?
Hard72A developer is building a customer support chatbot using NVIDIA NIM microservices. They need the model to always respond in a formal tone and never mention competitor products. Where should these directives be placed in the API request to ensure consistent behavior across all user interactions?
Medium73A developer is building a customer-support assistant that must retrieve answers only from an approved internal knowledge base and cite the source document for each reply. They are using NVIDIA NIM microservices for the LLM and an embedding model, and they need the application layer to enforce citation behavior and reject answers that are not grounded in retrieved passages. Which software development approach best enforces this grounding requirement?
Medium74A developer is using the NVIDIA Triton Inference Server to deploy a TensorRT-LLM optimized model. They want to send a request with multiple prompts to be processed in a single inference call. Which Triton feature should they use?
Easy75A developer is building a Retrieval-Augmented Generation (RAG) pipeline using NVIDIA NIM microservices. They need to ensure that the retriever returns the most relevant documents for a given query. Which two components should they optimize? (Choose two.)
Hard76A developer is debugging a TensorRT-LLM generation that intermittently produces truncated responses when many users submit long prompts concurrently. Logs show requests completing without errors, but outputs stop mid-sentence. Which configuration change is most likely to resolve this?
Hard77When deploying an LLM, what is the 'Time to First Token' (TTFT) metric used to measure?
EasyOther domains
All NCA-GENL exam domains
Frequently asked questions
- What does the Software Development domain cover on the NCA-GENL exam?
- Be able to design a production LLM integration that resists prompt injection, select quantization precision from accuracy and latency needs, and tune a TensorRT-LLM RAG deployment. The most important thing is correctly separating untrusted input from system instructions and diagnosing where latency actually occurs.
- How many questions are in this domain?
- This page lists all 77 Software Development questions in the NCA-GENL question bank. The actual exam draws from this domain proportionally to its weighting in the official exam blueprint.
- What is the best way to practise this domain?
- Start with a short focused session (10 questions) to identify gaps, then work through explanations. Repeat with a longer session once the weak areas feel solid.
- Can I practise only Software Development questions?
- Yes — the session launcher on this page filters questions to this domain only. Choose any session length for inline explanations and scoring.