Be able to design a production LLM integration that resists prompt injection, select quantization precision from accuracy and latency needs, and tune a TensorRT-LLM RAG deployment. The most important thing is correctly separating untrusted input from system instructions and diagnosing where latency actually occurs.
Start practicing
Software Development — choose a session length
Free · No account required
Domain overview
This domain covers building and shipping generative AI applications with NVIDIA tooling: integrating LLMs into production services, defending against prompt injection, quantizing models, and tuning inference. Questions are scenario-based, asking you to pick the right software engineering pattern, TensorRT-LLM optimization, or deployment configuration setting for a described production problem.
Exam objectives
Applying input validation and instruction/data separation patterns to block prompt injection in LLM services
Choosing FP16 versus INT8 quantization based on accuracy, latency, memory, and hardware support
Optimizing RAG retrieval latency using NVIDIA TensorRT-LLM and efficient embedding/vector search
Interpreting TensorRT-LLM deployment JSON options such as enable_cuda_graph for inference tuning
Treating prompt injection as a model problem instead of using software patterns like delimiting untrusted input and validating outputs
Assuming INT8 always beats FP16, ignoring accuracy loss, calibration effort, and GPU support for the chosen precision
Blaming generation for RAG latency when the bottleneck is document retrieval, embedding, or vector search before the LLM runs
Click any question to see the full explanation and answer options, or start a focused practice session above.
When deploying a Large Language Model using TensorRT-LLM, which TWO configuration factors must be tuned to maximize KV cache efficiency?
2When integrating an LLM into an application using NVIDIA API endpoints, what is the primary purpose of the 'System' role in the messages payload?
3A developer needs to ensure that an LLM application remains deterministic across multiple runs. Which parameter configuration is most effective?
4When designing a scalable inference microservice, which factor most significantly impacts the 'Time to First Token' (TTFT) for concurrent users?
5Refer to the exhibit. Given the current configuration, what is the primary risk during high-traffic bursts?
6When utilizing NVIDIA NIM for deployment, why is it recommended to use a containerized environment?
7Which NVIDIA technology enables efficient cross-GPU communication during Tensor Parallelism for large-scale model inference?
8When evaluating an LLM for a domain-specific task, why is 'Few-Shot Prompting' often superior to 'Zero-Shot Prompting'?
9Which NVIDIA SDK is specifically optimized for high-performance deep learning inference and supports the deployment of quantized models?
10Which TWO of the following practices are recommended when using NVIDIA Triton Inference Server to maximize throughput for a concurrent multi-model deployment?
11When developing with NVIDIA NeMo, which component is primarily responsible for scaling the training of massive LLMs across multiple GPU nodes?
12During development of a RAG application using NVIDIA NeMo Guardrails, why is it important to define specific 'canonical forms' in the configuration?
13Which memory management strategy in TensorRT-LLM is specifically designed to minimize fragmentation and allow for efficient KV cache allocation in multi-user environments?
14When using NVIDIA Riva for speech-to-text applications, which component provides the real-time transcription service based on streaming audio inputs?
15Refer to the exhibit. If a developer increases the 'max_batch_size' in the JSON configuration, what is the primary expected trade-off in the system's performance metrics?
16What is the primary function of the NVIDIA NGC (NVIDIA GPU Cloud) registry in the software development lifecycle for Generative AI?
17Refer to the exhibit. What is the technical implication of using the specified 'fp8' precision mode in this model configuration?
18A developer is optimizing a retrieval-augmented generation (RAG) pipeline using NVIDIA TensorRT-LLM. They notice excessive latency during the document retrieval phase before the generation starts. Which optimization strategy is most effective for this bottleneck?
19Which TWO of the following are primary benefits of using NVIDIA Triton Inference Server for deploying generative AI models?
20When fine-tuning a Large Language Model using PEFT (Parameter-Efficient Fine-Tuning) techniques like LoRA, what is the primary technical advantage being leveraged?
21An application requires streaming responses from a deployed LLM. Which communication protocol is most suitable for minimizing latency and ensuring efficient data delivery in a real-time generative AI application?
22When implementing a Guardrails layer in a generative AI application, what is the primary goal regarding model output?
23Refer to the exhibit. A developer wants to make the model's output more deterministic and focused on highly probable tokens. Which change should be made to the configuration policy?
24Which component in the NVIDIA AI Enterprise stack is specifically designed to orchestrate the lifecycle of multi-model deployments on Kubernetes?
25Which THREE factors should a developer consider when choosing between FP16 and INT8 quantization for a production LLM deployment?
26What is the primary purpose of using a Model Repository in the NVIDIA Triton Inference Server architecture?
27Refer to the exhibit. An engineer receives this error during deployment. What is the most likely cause?
28When deploying an LLM, what is the 'Time to First Token' (TTFT) metric used to measure?
29Refer to the exhibit. An engineer is tuning a deployment config. Why is 'enable_cuda_graph' set to true in this JSON configuration?
30When integrating an LLM into a production application, you must protect against prompt injection. Which software engineering pattern is most effective for this purpose?
31A developer is integrating a NeMo Guardrails configuration into an existing chatbot. They need to ensure that the LLM does not generate content related to unauthorized financial advice. Which mechanism should they implement to achieve this programmatic constraint?
32Which NVIDIA framework is specifically designed to facilitate the deployment of optimized LLMs as microservices with standardized APIs?
33A developer is building a customer support chatbot using NVIDIA NIM microservices. They need the model to always respond in a formal tone and never mention competitor products. Where should these directives be placed in the API request to ensure consistent behavior across all user interactions?
34A developer is building a document-summarization service on NVIDIA NIM for LLMs and wants to stream partial tokens to the client while the model is still generating. The NIM endpoint exposes an OpenAI-compatible /chat/completions route. Which request parameter should the developer set to receive incremental token deltas rather than one complete response body?
35A developer is building a customer-support assistant on NVIDIA NIM microservices. After a model update, responses that previously arrived in under 300 ms now take over two seconds, and the streaming client shows a long pause before the first token appears. GPU utilization is low and the prompt template was not changed. Which action should the developer take first to diagnose the regression?
36A developer is deploying a TensorRT-LLM optimized model on NVIDIA Triton Inference Server. They observe that the first inference request takes significantly longer than subsequent ones. Which Triton feature should they configure to reduce this initial latency?
37A developer is building a document summarization service using an NVIDIA NIM microservice for Llama-3. The service must process batches of 20 documents at once to maximize throughput. The NIM container is already running with default settings. Which API parameter should the developer configure to enable efficient batched inference?
38A team fine-tunes a Llama-3 8B model with NVIDIA NeMo Framework and must ship an inference artifact that a C++ service can load without a Python runtime. They want maximum throughput on Hopper GPUs and plan to serve many concurrent requests with in-flight batching. Which artifact and runtime pairing best satisfies these constraints?
39A developer is building a generative AI application that uses an NVIDIA NIM microservice for a Llama 3 model. They need to persist the model's responses and associated metadata for later auditing. Which approach best integrates NIM with an external datastore?
40A developer is using the NVIDIA NeMo framework to fine-tune a large language model. They want to reduce GPU memory usage during training without significantly sacrificing model quality. Which technique should they apply?
41A team fine-tunes a Llama model with NVIDIA NeMo Framework and must serve it behind an OpenAI-compatible endpoint with no Python glue code. They want the adapter weights kept separate from the base model so several adapters can share one loaded base. Which deployment approach fits these constraints?
42A developer is building a customer support chatbot using NVIDIA NIM microservices. The application must reliably return structured JSON containing 'intent' and 'confidence' fields for downstream ticket routing. Which approach should the developer use to constrain the model's output format?
43A developer is using NVIDIA Triton Inference Server to deploy a TensorRT-LLM optimized model. The model must support multiple concurrent users with low latency. The developer notices that latency spikes when many requests arrive simultaneously. Which Triton feature should be configured to improve throughput while maintaining acceptable latency?
44A developer is using the NVIDIA Triton Inference Server to deploy a TensorRT-LLM optimized model. They want to send a request with multiple prompts to be processed in a single inference call. Which Triton feature should they use?
45A developer is packaging a fine-tuned Llama 3 model as a TensorRT-LLM engine for an on-premises inference service. The model was trained with a custom tokenizer that adds four new special tokens beyond the base vocabulary. When the engine is built and the service is started, the model outputs garbled text and repeats the same fragment regardless of the prompt. Which action should the developer take to resolve this?
46A developer wants an LLM application to answer questions about an internal knowledge base that changes daily. Rather than retraining the model, they plan to retrieve relevant passages at query time and place them into the prompt. Which approach are they implementing?
47A developer is building a customer-support assistant that must retrieve answers only from an approved internal knowledge base and cite the source document for each reply. They are using NVIDIA NIM microservices for the LLM and an embedding model, and they need the application layer to enforce citation behavior and reject answers that are not grounded in retrieved passages. Which software development approach best enforces this grounding requirement?
48A developer is writing a Python service that calls an NVIDIA-hosted NIM endpoint for a Llama model. The service must recover gracefully when the endpoint returns HTTP 429 responses during peak traffic, without dropping user requests. Which implementation approach best satisfies this requirement?
49A team is building an internal document assistant and wants the model to answer only from an approved corpus of HR policy PDFs. They will deploy the model with NVIDIA NIM and control grounding at generation time by injecting retrieved passages into the prompt. Which parameter combination in the NIM chat completions request best enforces this grounding while keeping responses deterministic for audit logs?
50A developer is preparing a RAG service that calls an NVIDIA-hosted LLM endpoint and must reduce hallucinations for questions whose answers are absent from the retrieved context. Which two practices should be applied in the application layer? (Choose two.)
51A team is serving a 70B-parameter model with NVIDIA Triton Inference Server and TensorRT-LLM. Under concurrent load, GPU memory is exhausted because each request reserves its own large KV cache. Which Triton feature should the team enable to share KV cache blocks across requests that have common prompt prefixes?
52A developer is building a Retrieval-Augmented Generation (RAG) pipeline using NVIDIA NIM microservices. They need to ensure that the retriever returns the most relevant documents for a given query. Which two components should they optimize? (Choose two.)
53A team wants to load and run an optimized quantized LLM entirely inside a Python application with minimal dependencies, using a single high-level API that handles engine building and generation. They are not deploying a network service. Which component of the NVIDIA software stack is designed for this use case?
54A developer is preparing a container for an LLM microservice that will run on an NVIDIA GPU node and must be deployable through NVIDIA NIM. They want the image to be portable across supported GPU generations while still using NVIDIA's optimized inference stack. Which two practices should they follow? (Choose two.)
55A team is building a RAG assistant and wants to reduce hallucinated citations. They plan to have the LLM return structured output that names the source document chunk used for each claim. Which implementation strategy most directly improves the reliability of that structured output?
56A developer is packaging a NeMo-based LLM application into a container for deployment on an NVIDIA GPU node. Which two practices are required to ensure the container can access the GPU and run inference efficiently? (Choose two.)
57A developer is writing an application that streams chat completions from an NVIDIA-hosted NIM endpoint. Users report that the interface freezes until the entire answer is ready, even though the endpoint supports token streaming. Which client-side change fixes the perceived latency?
58A team is serving a 70B-parameter LLM with TensorRT-LLM on a node with four GPUs. During load testing they observe that increasing concurrent requests improves throughput up to a point, then latency spikes sharply and GPU memory utilization sits near the limit. Profiling shows the KV cache is being paged out and recomputed. Which change most directly addresses this bottleneck?
59A developer is using the NVIDIA API Catalog to experiment with a hosted LLM. They want to send a prompt and receive a completion. Which endpoint should they use?
60A developer is writing a Python service that calls a locally hosted NVIDIA NIM microservice for a Llama model. They want to keep the client code portable so the same class can later target NVIDIA's hosted API endpoints without rewrites. Which client approach fits this goal?
61A developer is tuning a TensorRT-LLM deployment of a long-context chat model and observes that GPU memory is exhausted under concurrent requests, causing requests to be rejected. They want to reduce KV cache memory pressure without retraining the model. (Choose two.)
62A developer is writing an application that calls an NVIDIA-hosted LLM endpoint and needs to keep multi-turn context across several user messages. Which payload structure should the application send to the chat completions API?
63A developer is deploying a large language model using NVIDIA TensorRT-LLM and wants to optimize inference for a production environment with limited GPU memory. Which two techniques can be used to reduce memory footprint while maintaining acceptable performance? (Choose two.)
64A developer is writing a Python client for an NVIDIA-hosted LLM endpoint and needs the model to answer every request in a strict JSON schema without extra prose. Where should the formatting contract be expressed so it applies consistently across all requests from the service?
65A developer needs to serve a quantized Llama model on an NVIDIA GPU and wants the runtime to automatically select the fastest available execution kernels for the detected GPU architecture. Which approach aligns with the NVIDIA inference stack for this requirement?
66A developer is containerizing an inference service built with TensorRT-LLM and NVIDIA NIM for a Kubernetes cluster. They want the deployment to start reliably and use the GPU efficiently. Which two practices should they follow? (Choose two.)
67A developer has a working TensorRT-LLM engine and wants to expose it through NVIDIA Triton Inference Server so that multiple client applications can call it over HTTP and gRPC with a stable interface. Which Triton feature should they configure to serve the TensorRT-LLM engine as a backend?
68A developer is using TensorRT-LLM to build a chatbot and wants to reduce the memory footprint of the KV cache during inference. Which technique should they use?
69A developer is evaluating a fine-tuned LLM with NVIDIA NeMo and observes that evaluation loss keeps decreasing while downstream task accuracy plateaus and then declines. Which action should the developer take to address this?
70A developer is building a retrieval-augmented generation service and needs to embed millions of document chunks and run low-latency similarity search over them on GPU. They want a library that handles both index construction and search with GPU acceleration. Which NVIDIA component should they use?
71A developer is using the NVIDIA API Catalog to test a Llama-3 model via its API endpoint. They need to send a request that includes a system prompt to set the model's behavior. Which component of the request payload is used to provide the system prompt?
72A developer is packaging a generative AI application that must run inference on-premises with NVIDIA GPUs and also expose an OpenAI-compatible HTTP API so existing client code works unchanged. Which two components should the developer use together to meet these requirements? (Choose two.)
73A developer is tuning a retrieval-augmented generation pipeline that uses NVIDIA NIM embeddings and a NIM LLM. Latency is dominated by embedding thousands of document chunks at query time because the team re-embeds the whole corpus on every request. Which change most directly fixes the architecture?
74A developer is packaging a generative AI application for NVIDIA AI Enterprise deployment on Kubernetes. The application must run an LLM served by NVIDIA NIM, an embedding model, and a vector database, and must support rolling upgrades without dropping in-flight inference requests. Which design choice best meets these requirements?
75A developer is debugging a TensorRT-LLM generation that intermittently produces truncated responses when many users submit long prompts concurrently. Logs show requests completing without errors, but outputs stop mid-sentence. Which configuration change is most likely to resolve this?
76A developer is debugging a RAG service where answers are correct in testing but degrade in production as the document corpus grows. Logs show retrieval returning chunks with high similarity scores that do not contain the answer. Which change most directly addresses the root cause?
77A developer is profiling a TensorRT-LLM serving deployment and notices that throughput collapses once concurrent requests exceed a small number of users, even though GPU compute utilization stays low. The model uses paged KV cache and continuous batching. Which factor most likely explains the bottleneck?
Be able to design a production LLM integration that resists prompt injection, select quantization precision from accuracy and latency needs, and tune a TensorRT-LLM RAG deployment. The most important thing is correctly separating untrusted input from system instructions and diagnosing where latency actually occurs.
The Courseiva NCA-GENL question bank contains 77 questions in the Software Development domain. Click any question to see the full explanation and answer breakdown.
Start with a 10-question focused session to identify your baseline accuracy in this domain. Read every explanation — even for questions you answer correctly — to understand the reasoning. Once you score consistently above 80%, move to a 20–30 question session to confirm depth before moving to the next domain.
Yes — the session launcher on this page draws questions exclusively from the Software Development domain. Choose 10, 20, 30, or 50 questions for a focused session, or click individual questions to review them one by one.
Save your results, see per-domain analytics, and get readiness scores — free, for every certification.
Sign Up FreeFree forever · Every certification included