Courseiva

CCNA Nca Software Development Questions

75 of 77 questions · Page 1/2 · Nca Software Development topic · Answers revealed

1
MCQeasy

A developer wants an LLM application to answer questions about an internal knowledge base that changes daily. Rather than retraining the model, they plan to retrieve relevant passages at query time and place them into the prompt. Which approach are they implementing?

A.Few-shot prompting, where several worked examples are prepended to every request.
B.Retrieval-augmented generation, where an embedding index supplies context for the prompt.
C.Supervised fine-tuning, where labeled question-answer pairs update the model weights.
D.Parameter-efficient fine-tuning with LoRA adapters trained on the knowledge base.
AnswerB

Retrieval-augmented generation separates knowledge from model weights: documents are embedded into a vector index, the query retrieves the closest passages, and those passages are inserted into the prompt. Because the index can be refreshed whenever documents change, answers stay current without any retraining or fine-tuning cycle.

Why this answer

Answering over a frequently changing corpus without retraining is the defining use case for retrieval-augmented generation: an embedding index is refreshed as documents change, and retrieved passages are injected into the prompt at query time. Fine-tuning variants and few-shot prompting all fix knowledge at build time or prompt-authoring time.

Exam trap

The trap here is reading "internal knowledge base" and jumping to fine-tuning, when the daily-change requirement rules out any approach that stores knowledge in model weights.

2
MCQeasy

A developer is writing an application that calls an NVIDIA-hosted LLM endpoint and needs to keep multi-turn context across several user messages. Which payload structure should the application send to the chat completions API?

A.A 'messages' array of objects, each with a 'role' and 'content', ordered from system to user and assistant turns.
B.A 'context' object mapping user IDs to their previous messages.
C.A 'history' array containing only the assistant's prior responses.
D.A single 'prompt' string containing all previous turns concatenated.
AnswerA

The chat completions API expects a messages array where each entry carries a role such as system, user, or assistant and its content. Preserving this ordered structure gives the model clear turn boundaries, enabling correct multi-turn context and instruction adherence across the conversation.

Why this answer

Chat completions APIs from NVIDIA expect an ordered messages array with role and content fields, preserving system, user, and assistant turns. This structure gives the model the turn boundaries it needs for coherent multi-turn dialogue, unlike a flat prompt string or incomplete history.

Exam trap

The trap here is assuming the endpoint accepts a single concatenated prompt string, when the chat schema requires role-tagged messages to preserve turn structure.

3
Multi-Selectmedium

A developer is packaging a generative AI application that must run inference on-premises with NVIDIA GPUs and also expose an OpenAI-compatible HTTP API so existing client code works unchanged. Which two components should the developer use together to meet these requirements? (Choose two.)

Select 2 answers
A.NVIDIA NeMo Guardrails as the component that translates OpenAI API calls into model inputs.
B.NVIDIA TensorRT as the reverse proxy that routes client requests to model replicas.
C.NVIDIA CUDA Toolkit as the HTTP server exposing the compatible endpoints.
D.NVIDIA Triton Inference Server as the serving runtime hosting the NIM model artifacts.
E.NVIDIA NIM for the model microservice providing the OpenAI-compatible inference endpoints.
AnswersD, E

Triton Inference Server is the runtime that hosts and executes the model artifacts behind the compatible API, handling batching, concurrency, and GPU scheduling. NIM microservices are delivered as Triton-hosted deployments, so pairing the compatible API layer with Triton as the execution engine matches both the on-premises GPU requirement and the API requirement.

Why this answer

The requirement pairs an on-premises GPU deployment with OpenAI-compatible APIs. NIM supplies optimized model microservices with those compatible endpoints, while Triton Inference Server provides the runtime that executes the model artifacts and manages batching and GPU resources. Together they deliver both the API contract and the execution environment without changes to existing client code.

Exam trap

The trap here is assuming that an optimization library or a guardrails framework can substitute for the serving runtime and API layer that actually expose compatible endpoints.

4
MCQmedium

A developer needs to ensure that an LLM application remains deterministic across multiple runs. Which parameter configuration is most effective?

A.Increase top_p to 1.0 and set temperature to 0.7.
B.Set temperature to 0.0 and define a fixed seed.
C.Disable streaming and use a large batch size for inference.
D.Apply top_k filtering with a value of 50.
AnswerB

Temperature 0.0 forces the model to choose the most likely token (greedy decoding), while a fixed seed ensures the underlying noise in the sampling process remains constant. Combined, these create a highly deterministic environment where inputs consistently map to the same output tokens, satisfying the requirement.

Why this answer

Determinism in LLMs is achieved by minimizing the stochastic nature of the generation process. By setting the temperature to zero and fixing the seed, the model consistently follows the same probability path. This is vital for debugging, testing, and production scenarios where identical inputs must yield identical outputs, ensuring reliability in complex automated workflows and compliance with validation requirements.

Exam trap

Candidates often only set the temperature to zero while forgetting to fix the random seed, leading to non-deterministic behavior stemming from initialization or floating-point non-associativity.

5
MCQmedium

A team is building an internal document assistant and wants the model to answer only from an approved corpus of HR policy PDFs. They will deploy the model with NVIDIA NIM and control grounding at generation time by injecting retrieved passages into the prompt. Which parameter combination in the NIM chat completions request best enforces this grounding while keeping responses deterministic for audit logs?

A.Set temperature to 0 and provide the retrieved passages inside the system and user messages as the only context, instructing the model to refuse when the answer is not present.
B.Set temperature to 1.5 and rely on the model's pretrained HR knowledge, then post-filter answers with a regex for policy numbers.
C.Set top_p to 0.1 and pass the PDFs as base64 file attachments in the request body so NIM parses them automatically.
D.Set presence_penalty to 2.0 and include only the document titles so the model is discouraged from inventing policy details.
AnswerA

Setting temperature to 0 makes sampling greedy and repeatable for audit logs, and placing the retrieved passages directly in the system and user messages is exactly how retrieval-augmented grounding is enforced at generation time with an OpenAI-compatible NIM endpoint. The explicit refusal instruction constrains the model to the supplied HR corpus, which is the requirement here.

Why this answer

Grounding in a NIM-hosted model is achieved by supplying the approved passages as prompt context and constraining behavior through the system message, while temperature zero gives repeatable outputs for audit. The other choices either increase randomness, assume server-side PDF parsing that does not exist in the chat completions API, or apply sampling penalties that do not limit the model to the HR corpus.

Exam trap

The trap here is assuming that a sampling knob such as top_p or presence_penalty can enforce grounding, when grounding actually comes from the retrieved context placed in the prompt.

6
MCQmedium

Which component in the NVIDIA AI Enterprise stack is specifically designed to orchestrate the lifecycle of multi-model deployments on Kubernetes?

A.NVIDIA CUDA Toolkit.
B.NVIDIA Triton Inference Server with Kubernetes Operator.
C.NVIDIA TensorRT-LLM library.
D.NVIDIA NeMo Framework.
AnswerB

The Triton Operator for Kubernetes automates the deployment, scaling, and lifecycle management of Triton instances. This allows developers to handle complex deployments, model updates, and resource allocation across a cluster, ensuring that generative models are available, performant, and correctly configured in a production-ready environment.

Why this answer

NVIDIA Triton, when integrated with Kubernetes using tools like the Triton Operator, provides the necessary orchestration for scaling, health monitoring, and lifecycle management. This orchestration is essential for maintaining high availability and efficient resource distribution in large-scale AI deployments, allowing developers to manage complex, multi-model architectures with standardized workflows that integrate seamlessly into existing DevOps CI/CD pipelines for AI applications.

Exam trap

Candidates often select general Kubernetes tools like 'kubectl' or 'Helm' alone, failing to realize that the Triton Operator is the specific component required for lifecycle management of AI models.

7
MCQhard

A team is serving a 70B-parameter LLM with TensorRT-LLM on a node with four GPUs. During load testing they observe that increasing concurrent requests improves throughput up to a point, then latency spikes sharply and GPU memory utilization sits near the limit. Profiling shows the KV cache is being paged out and recomputed. Which change most directly addresses this bottleneck?

A.Enable in-flight batching and increase the maximum batch size so more requests share each forward pass.
B.Switch the deployment from tensor parallelism across four GPUs to pipeline parallelism to reduce per-GPU memory pressure.
C.Reduce the maximum sequence length and configure a KV cache size that fits in remaining GPU memory, potentially with quantized cache.
D.Increase the tensor-parallel degree beyond four GPUs so the model weights and cache are spread across more devices.
AnswerC

The paging and recomputation indicate that the KV cache exceeds available memory as concurrency rises. Capping maximum sequence length and sizing the cache explicitly, optionally with FP8 or INT8 KV cache quantization, keeps the working set resident and removes the recompute penalty. This directly targets the profiled bottleneck while preserving the existing four-GPU tensor-parallel layout.

Why this answer

Sharp latency growth with memory near the limit and profiler evidence of KV cache paging and recomputation means the cache working set no longer fits. Capping maximum sequence length and explicitly sizing the KV cache, optionally with quantized cache, keeps the cache resident and eliminates recomputation. In-flight batching remains useful, but it must operate within a cache budget that the deployment actually fits.

Exam trap

The trap here is treating a throughput plateau as a batching problem when the profiler shows cache eviction and recomputation.

8
MCQeasy

What is the primary function of the NVIDIA NGC (NVIDIA GPU Cloud) registry in the software development lifecycle for Generative AI?

A.To act as a remote compiler for Python code.
B.To provide optimized containers and pre-trained models.
C.To provide a hosted Kubernetes cluster for execution.
D.To manage source code version control for teams.
AnswerB

NGC provides optimized Docker containers that contain all necessary drivers, libraries, and frameworks. It also hosts pre-trained AI models. This ecosystem allows developers to quickly bootstrap their projects with high-quality, pre-tested software components, significantly reducing the time required to reach a functional production-grade AI solution.

Why this answer

NGC acts as a centralized repository for pre-trained models, optimized containers, and Helm charts. It simplifies the development process by providing vetted, ready-to-deploy environments that are already optimized for NVIDIA GPUs. For developers, this eliminates the 'dependency hell' of setting up complex AI stacks manually, ensuring that the software runs optimally on the underlying hardware from the moment it is deployed in a production cluster.

Exam trap

Test-takers frequently confuse NGC with general-purpose cloud storage providers or raw compute orchestration platforms, overlooking its core role in distributing pre-optimized AI assets.

9
MCQmedium

A developer is building a document-summarization service on NVIDIA NIM for LLMs and wants to stream partial tokens to the client while the model is still generating. The NIM endpoint exposes an OpenAI-compatible /chat/completions route. Which request parameter should the developer set to receive incremental token deltas rather than one complete response body?

A.Set "logprobs": true and reconstruct partial text from the token log probabilities.
B.Set the SSE header "Accept: text/event-stream" on the request without changing the body.
C.Set "n": 2 so the endpoint returns two candidate completions the client can merge.
D.Set "stream": true in the JSON request body and consume the server-sent event chunks.
AnswerD

The OpenAI-compatible chat completions schema used by NIM for LLMs accepts a boolean stream field. When it is true, the server returns text/event-stream data chunks containing incremental choices delta objects terminated by a data: [DONE] sentinel, which is exactly the incremental token delivery the summarization UI needs.

Why this answer

Streaming on an OpenAI-compatible NIM chat completions endpoint is controlled by the stream boolean in the request body, which switches the transport to server-sent events carrying incremental delta objects. The other parameters alter candidate count, content negotiation, or scoring metadata, and none of them cause tokens to be emitted before generation completes.

Exam trap

The trap here is assuming that an Accept header or a parameter like n or logprobs changes response timing, when only the stream field actually enables incremental token delivery.

10
Multi-Selecthard

A developer is deploying a large language model using NVIDIA TensorRT-LLM and wants to optimize inference for a production environment with limited GPU memory. Which two techniques can be used to reduce memory footprint while maintaining acceptable performance? (Choose two.)

Select 2 answers
A.Using a larger batch size to amortize memory overhead across more requests.
B.Enabling paged KV cache and tuning the maximum sequence length to a realistic value.
C.Increasing the tensor parallelism degree to spread the model across more GPUs.
D.Quantization of model weights to INT8 or FP8 using TensorRT-LLM's quantization toolkit.
E.Storing the model weights on CPU and offloading them to GPU on demand.
AnswersB, D

Paged KV cache in TensorRT-LLM manages key-value cache memory in fixed-size blocks, reducing fragmentation and allowing more efficient memory usage. Tuning the maximum sequence length to a realistic value prevents over-allocation of KV cache for sequences that are never reached. Together, these techniques can significantly reduce memory footprint while maintaining performance for typical workloads.

Why this answer

Quantization and paged KV cache with tuned sequence length are effective techniques to reduce GPU memory footprint in TensorRT-LLM. Quantization lowers precision of weights, and paged KV cache optimizes memory allocation for key-value pairs. Both help fit larger models on limited GPUs while preserving performance.

Other options either increase memory usage or introduce latency, making them unsuitable for the scenario.

Exam trap

The trap here is thinking that increasing tensor parallelism or batch size reduces memory footprint, when they actually distribute or increase memory usage, respectively.

11
MCQmedium

A developer is using the NVIDIA NeMo framework to fine-tune a large language model. They want to reduce GPU memory usage during training without significantly sacrificing model quality. Which technique should they apply?

A.Increase the micro batch size
B.Disable mixed precision training
C.Enable gradient checkpointing
D.Use a higher learning rate
AnswerC

Gradient checkpointing trades compute for memory by not storing all intermediate activations during the forward pass, recomputing them during backward. This significantly reduces memory usage, allowing larger models or batch sizes, with only a modest increase in training time and negligible impact on model quality.

Why this answer

Gradient checkpointing reduces memory by storing only a subset of activations and recomputing the rest during backpropagation. This allows training larger models or using bigger batches on the same GPU. Other options either increase memory usage or do not affect it, making gradient checkpointing the correct choice for memory-constrained fine-tuning.

Exam trap

The trap here is thinking that increasing batch size or disabling mixed precision would help memory, when they actually increase it.

12
MCQhard

A team is serving a 70B-parameter model with NVIDIA Triton Inference Server and TensorRT-LLM. Under concurrent load, GPU memory is exhausted because each request reserves its own large KV cache. Which Triton feature should the team enable to share KV cache blocks across requests that have common prompt prefixes?

A.Dynamic batching with a larger maximum batch size.
B.Model instance groups with multiple instances per GPU.
C.Sequence batching with a higher maximum queue delay.
D.Paged KV cache with prefix caching enabled in the TensorRT-LLM backend.
AnswerD

TensorRT-LLM's paged KV cache stores cache in fixed-size blocks, and prefix caching reuses blocks whose token prefixes match earlier requests. In Triton, enabling this in the backend configuration lets many concurrent requests share common system prompts, drastically cutting GPU memory and raising throughput for the 70B deployment.

Why this answer

Prefix caching combined with paged KV cache in the TensorRT-LLM backend lets Triton reuse cache blocks for requests sharing a prompt prefix, directly reducing memory consumed by duplicate caches. Batching, sequence scheduling, and instance groups tune throughput or state handling but do not deduplicate KV cache blocks.

Exam trap

The trap here is confusing batching or instance tuning, which improve utilization, with prefix caching, which is what actually shares KV cache memory across requests.

13
MCQeasy

When fine-tuning a Large Language Model using PEFT (Parameter-Efficient Fine-Tuning) techniques like LoRA, what is the primary technical advantage being leveraged?

A.Full parameter updates are performed on every layer.
B.Weight matrices are frozen and low-rank adaptors are trained.
C.The entire model is converted to INT8 precision during training.
D.All model activations are stored in CPU memory.
AnswerB

LoRA freezes the original model weights and adds small, trainable rank decomposition matrices to transformer layers. This drastically reduces the total count of trainable parameters, leading to much lower VRAM usage during the training process while maintaining high performance on the target downstream tasks.

Why this answer

LoRA injects trainable rank decomposition matrices into the transformer layers while keeping pre-trained weights frozen. This approach significantly reduces the number of parameters requiring gradient updates, which saves memory and compute. This is essential for developers working with limited GPU hardware, allowing them to adapt massive models to specific tasks without full-parameter fine-tuning, which would be computationally prohibitive for most enterprise-grade infrastructure deployments.

Exam trap

Students frequently assume PEFT updates all model weights with a smaller learning rate, misunderstanding that base weights remain completely frozen.

14
MCQmedium

When designing a scalable inference microservice, which factor most significantly impacts the 'Time to First Token' (TTFT) for concurrent users?

A.The total number of output tokens generated.
B.The prefill phase efficiency and GPU throughput.
C.The network latency between the user and the load balancer.
D.The number of active vector databases connected to the service.
AnswerB

The prefill phase is where the model processes the prompt and calculates the initial KV cache. High GPU throughput and optimized kernels (like those in TensorRT-LLM) directly reduce the duration of this compute-heavy phase, which is the primary contributor to the time elapsed before the first token appears.

Why this answer

TTFT is primarily driven by the time taken to process the input prompt and initialize the decoding state. In a multi-user environment, resource contention and efficient scheduling become the main inhibitors. Optimizing the prefill phase, which is compute-bound, is critical for achieving low TTFT, allowing users to perceive the application as responsive even during high load periods.

Exam trap

Candidates often focus on decoding speed or output tokens, failing to realize that TTFT is primarily constrained by the compute-heavy prefill phase of the prompt.

15
MCQeasy

A developer needs to serve a quantized Llama model on an NVIDIA GPU and wants the runtime to automatically select the fastest available execution kernels for the detected GPU architecture. Which approach aligns with the NVIDIA inference stack for this requirement?

A.Write custom CUDA kernels for each target GPU and dispatch them manually at runtime.
B.Convert the model to ONNX and rely on the default CPU execution provider.
C.Run the model in eager execution mode on PyTorch without any compilation step.
D.Compile the model with TensorRT-LLM to produce an engine that contains architecture-tuned kernels and let the runtime pick the best ones for the GPU present.
AnswerD

TensorRT-LLM builds engines containing kernels tuned for the target GPU architecture, and the runtime selects among available tactics at execution time. Building for the deployment architecture ensures the best kernels are present, so the developer gets optimized execution without writing kernels. This directly matches the requirement for automatic selection of the fastest kernels on the detected GPU.

Why this answer

TensorRT-LLM compiles model graphs into engines containing kernels tuned for specific GPU architectures, and at runtime it selects from the available tactics for the hardware present. Building the engine for the deployment GPU ensures the fastest kernels are candidates. This gives the developer optimized execution without hand-authoring CUDA code or losing performance to eager framework dispatch.

Exam trap

The trap here is believing that any GPU-enabled framework automatically performs architecture-aware kernel selection, when that behavior comes from an ahead-of-time compiled inference engine.

16
MCQhard

Refer to the exhibit. A developer wants to make the model's output more deterministic and focused on highly probable tokens. Which change should be made to the configuration policy?

A.Increase temperature to 1.5 and top_p to 1.0.
B.Decrease temperature to 0.1 and top_p to 0.5.
C.Set presence_penalty to 1.0.
D.Increase max_tokens to 4096.
AnswerB

A lower temperature of 0.1 makes the probability distribution sharper, favoring the most likely tokens. A lower top_p of 0.5 further restricts the model to only the most probable candidates, resulting in highly deterministic and focused outputs that are ideal for consistent, repetitive, or factual tasks.

Why this answer

Determinism in LLMs is controlled by the temperature and top-p sampling parameters. Lowering temperature reduces the randomness in the probability distribution, while lowering top-p restricts the sampling pool to a smaller subset of high-probability tokens. Adjusting these parameters is vital for applications requiring factual consistency and stable outputs, as they directly dictate the sampling strategy used during the token generation process in the inference engine.

Exam trap

Candidates often confuse the effects of temperature and top-p, sometimes suggesting an increase in these values when the goal is to make the model output more predictable and focused.

17
MCQhard

Refer to the exhibit. Given the current configuration, what is the primary risk during high-traffic bursts?

A.FP8 quantization will cause excessive model hallucinations.
B.Head-of-line blocking due to the FCFS scheduling policy.
C.The request limit of 128 is too low to saturate the GPU.
D.The strict policy setting prevents dynamic batching.
AnswerB

FCFS processes requests in the order they arrive regardless of their compute requirements. A single long-generation request will occupy GPU resources, causing all subsequent, potentially short requests to wait. This leads to poor overall system latency and creates a bottleneck during high-traffic bursts, significantly degrading user experience.

Why this answer

The 'fcfs' (First-Come, First-Served) scheduling policy is suboptimal for LLM inference because it lacks priority handling. In high-traffic scenarios, long requests block shorter ones, leading to 'head-of-line blocking'. This causes latency spikes for all users.

For robust production systems, implementing continuous batching or priority-aware scheduling is necessary to ensure that the GPU utilization remains high while maintaining acceptable response times for individual requests.

Exam trap

Students often assume that simple queueing algorithms like First-Come-First-Served work efficiently for LLMs, forgetting that varying token generation lengths cause severe head-of-line blocking under heavy load.

18
MCQhard

When integrating an LLM into a production application, you must protect against prompt injection. Which software engineering pattern is most effective for this purpose?

A.Hard-coding all possible malicious user inputs.
B.Implementing an asynchronous guardrail input filter.
C.Increasing the model's temperature parameter.
D.Caching all prompts in a local database.
AnswerB

An input filter acts as a gateway that checks prompts against semantic and structural safety rules. By using guardrails, developers can programmatically block or sanitize malicious instructions before they reach the LLM, ensuring a secure interaction loop for every user.

Why this answer

Using a guardrail pattern, such as NVIDIA NeMo Guardrails, allows developers to intercept inputs and outputs to validate them against safety policies. This protects the LLM from malicious prompts and prevents the generation of harmful content. Guardrails are fundamental in enterprise AI software development because they provide a programmatic layer of control that enforces safety, compliance, and reliability, regardless of the underlying LLM's inherent behavior.

Exam trap

Candidates often confuse static prompt engineering rules or fine-tuning with runtime defense mechanisms, missing that prompt injection requires active interception filters.

19
MCQmedium

An application requires streaming responses from a deployed LLM. Which communication protocol is most suitable for minimizing latency and ensuring efficient data delivery in a real-time generative AI application?

A.HTTP/1.1 REST without streaming support.
B.gRPC with server-side streaming.
C.Standard SMTP email protocols.
D.Synchronous SQL query polling.
AnswerB

gRPC uses HTTP/2 for transport, which supports server-side streaming. This allows tokens to be sent back as soon as they are generated, minimizing the time-to-first-token and creating a fluid, real-time experience for the end-user. It is the preferred method for high-performance communication in modern AI stacks.

Why this answer

Streaming generative responses requires a protocol that supports persistent connections and low-overhead message framing. gRPC with server-side streaming is the industry standard for high-performance AI inference, as it facilitates efficient binary serialization and maintains low latency across the network. Using this protocol is critical for user-facing applications where perceived latency is directly tied to the speed at which text tokens appear on the screen.

Exam trap

Candidates often confuse gRPC with standard HTTP/1.1 REST APIs, forgetting that REST lacks native bidirectional streaming capabilities and incurs higher overhead for continuous token transmission in real-time generative applications.

20
MCQeasy

Which NVIDIA framework is specifically designed to facilitate the deployment of optimized LLMs as microservices with standardized APIs?

A.NVIDIA TensorRT-LLM
B.NVIDIA NIM
C.NVIDIA CUDA-X
D.NVIDIA NeMo
AnswerB

NVIDIA NIM provides the infrastructure to deploy AI models as containerized microservices with uniform APIs. It abstracts the underlying hardware and software optimizations, enabling developers to serve models consistently across different environments while maintaining enterprise-grade performance and ease of integration into existing software stacks.

Why this answer

NVIDIA NIM (NVIDIA Inference Microservices) is the standard framework for packaging optimized AI models as containerized microservices. By providing standardized APIs, NIM allows developers to easily integrate state-of-the-art models into their applications. This abstraction simplifies the deployment lifecycle, ensuring that developers can focus on building application features rather than managing the complexities of model optimization, serving infrastructure, and hardware-specific performance tuning.

Exam trap

Candidates often confuse NVIDIA NIM with generic container orchestration tools like Kubernetes or model training frameworks like PyTorch, failing to see the 'microservice' and 'API' distinction.

21
MCQmedium

A developer is building a retrieval-augmented generation service and needs to embed millions of document chunks and run low-latency similarity search over them on GPU. They want a library that handles both index construction and search with GPU acceleration. Which NVIDIA component should they use?

A.NVIDIA cuVS, a GPU-accelerated library for vector search and clustering that builds and queries similarity indexes.
B.NVIDIA NeMo Retriever embedding microservices for generating document and query embeddings at scale.
C.NVIDIA Triton Inference Server with an ensemble pipeline that chains preprocessing and postprocessing steps.
D.NVIDIA TensorRT-LLM for compiling the embedding model into an optimized engine.
AnswerA

cuVS provides GPU-accelerated vector search and clustering primitives, including index build and nearest-neighbor query paths. It is designed exactly for large-scale embedding search where low latency matters, covering both index construction and search. That matches the developer's dual requirement of embedding millions of chunks and serving fast similarity lookups on GPU.

Why this answer

The task splits into encoding and retrieval, and the retrieval half demands an index plus fast nearest-neighbor search. cuVS supplies GPU-accelerated vector search and clustering, handling index build and query for large embedding sets. Embedding microservices, TensorRT-LLM, and Triton each address adjacent concerns such as encoding, generation optimization, or orchestration, but none provides the vector index and search capability the scenario requires.

Exam trap

The trap here is conflating embedding generation with vector search, assuming that a service which produces embeddings also performs similarity retrieval over them.

22
MCQeasy

When integrating an LLM into an application using NVIDIA API endpoints, what is the primary purpose of the 'System' role in the messages payload?

A.To provide the user with a history of previous successful interactions.
B.To define the model's behavioral guidelines and operational constraints.
C.To act as a buffer for temporary memory storage during inference.
D.To specify the hardware architecture used for the inference request.
AnswerB

The System role is explicitly designed to set the stage for the model's persona, functional boundaries, and safety policies. By defining these at the start, developers ensure the model adheres to application requirements regardless of user input, providing a stable foundation for the conversation context.

Why this answer

The 'System' role provides foundational instructions that dictate the model's behavior, tone, and constraints throughout the conversation. It is a critical component in ensuring that the AI remains within the expected parameters of the application, such as maintaining a specific persona or following strict safety guidelines. Mastering this role is fundamental for developers to ensure consistent and high-quality outputs across different user interactions.

Exam trap

Test-takers often confuse the system role with user prompts or few-shot examples, failing to recognize its unique function in enforcing global behavioral constraints.

23
MCQeasy

When using NVIDIA Riva for speech-to-text applications, which component provides the real-time transcription service based on streaming audio inputs?

A.Riva Speech API (ASR).
B.The TensorRT Model Optimizer.
C.The NVIDIA NeMo Guardrails Engine.
D.The Triton Model Repository.
AnswerA

The Riva Speech API provides the ASR (Automatic Speech Recognition) services necessary for real-time transcription. It is built to handle streaming data, allowing the server to process chunks of audio as they arrive from the client, which is the fundamental requirement for live transcription and voice-interactive system deployments.

Why this answer

Riva Speech API is designed for high-performance, low-latency streaming applications. It uses optimized neural networks to transcribe audio in real-time as it arrives. By leveraging GPU acceleration, Riva ensures that the transcription remains synchronized with the audio feed, which is critical for live transcription services, voice assistants, and other interactive applications that require immediate feedback from the system to the user.

Exam trap

Candidates often select general Riva services or non-streaming components, failing to identify the specific Riva Speech API (ASR) which is architected for low-latency, real-time streaming transcription tasks.

24
MCQmedium

Which THREE factors should a developer consider when choosing between FP16 and INT8 quantization for a production LLM deployment?

A.Hardware support for specific tensor core operations.
B.The potential impact on model perplexity or accuracy.
C.The availability of sufficient cooling for the server.
D.The memory overhead of the model weights.
E.The compatibility with standard CSV file formats.
AnswerA, B, D

NVIDIA GPUs have varying support for different precision types in their Tensor Cores. Choosing a precision format that the underlying hardware can accelerate natively is crucial for achieving maximum throughput, as using non-optimized formats results in significant performance degradation during the inference execution phase.

Why this answer

Choosing the correct precision balance is critical for optimizing LLM performance. FP16 offers a good compromise between quality and speed, while INT8 provides significant memory savings and increased throughput at the cost of potential precision loss. Understanding these trade-offs allows developers to align their model deployment with hardware constraints and accuracy requirements, ensuring the application maintains acceptable quality while meeting performance targets for end-user response times.

Exam trap

Candidates often assume INT8 is always superior because it is faster, ignoring that hardware support for specific tensor operations and perplexity degradation are critical constraints that can disqualify INT8 for certain use cases.

25
MCQmedium

During development of a RAG application using NVIDIA NeMo Guardrails, why is it important to define specific 'canonical forms' in the configuration?

A.To increase the training speed of the underlying LLM.
B.To standardize user inputs for easier rule matching.
C.To force the LLM to output only JSON-formatted data.
D.To bypass the authentication module of the application.
AnswerB

Canonical forms map semantically similar user expressions into a single, standardized identifier. This allows developers to write rules based on these canonical forms, which makes the Guardrails configuration much cleaner and easier to maintain, ensuring the system handles variations in natural language input consistently and reliably.

Why this answer

Canonical forms in NeMo Guardrails act as a bridge between diverse user inputs and the system's internal logic. By mapping many variations of a user query to a single canonical form, developers simplify the dialog management rules. This ensures consistent responses and predictable guardrail behavior, which is crucial for maintaining safety and accuracy in enterprise LLM applications where ambiguous user intent could lead to potentially dangerous or incorrect AI outputs.

Exam trap

Candidates often assume canonical forms are meant for changing the LLM's core architecture or fine-tuning weights, rather than simply normalizing varied user inputs for reliable rule matching.

26
MCQeasy

When developing with NVIDIA NeMo, which component is primarily responsible for scaling the training of massive LLMs across multiple GPU nodes?

A.The NeMo Data Augmentation Module.
B.PyTorch Lightning and Megatron-Core.
C.The TensorRT Model Parser.
D.The NVIDIA Driver API.
AnswerB

NeMo integrates PyTorch Lightning for training orchestration and Megatron-Core for handling model-parallelism primitives. This combination allows NeMo to efficiently distribute model weights and activation states across multiple GPUs, which is the foundational requirement for training very large language models that exceed the capacity of single hardware devices.

Why this answer

NVIDIA NeMo leverages PyTorch Lightning and the Megatron-Core library to handle distributed training complexities. This framework is essential because LLMs are too large to fit into a single GPU's memory. By using techniques like tensor parallelism, pipeline parallelism, and data parallelism, NeMo enables developers to train models with hundreds of billions of parameters efficiently across clusters, ensuring consistent performance and scalability in high-performance computing environments.

Exam trap

Candidates often credit standard distributed data-parallel frameworks alone, forgetting that massive LLMs require specialized tensor and pipeline parallelism libraries like Megatron-Core.

27
MCQmedium

What is the primary purpose of using a Model Repository in the NVIDIA Triton Inference Server architecture?

A.To store raw training data for future fine-tuning.
B.To manage model versioning and configuration dynamically.
C.To act as a high-performance vector database.
D.To serve as a GUI for end-user model interactions.
AnswerB

The repository allows for structured versioning of models and their associated config files. Triton can automatically detect changes in the directory, enabling hot-swapping of models, version management, and clean deployment cycles without needing to restart the inference server, which is vital for high-availability systems.

Why this answer

The Model Repository is a centralized, organized directory structure that Triton monitors to load and serve models. It provides a standardized interface for managing model versions, configurations, and dependencies. This structure is critical for version control, allowing developers to roll back models, conduct A/B testing, and ensure consistent deployment across multiple environments, which is essential for maintaining system stability and reliability in production software development.

Exam trap

Candidates assume the repository is for model training or weight storage. They fail to understand that Triton's repository is specifically a deployment-focused directory structure for versioning and runtime configuration management.

28
MCQeasy

A developer is writing an application that streams chat completions from an NVIDIA-hosted NIM endpoint. Users report that the interface freezes until the entire answer is ready, even though the endpoint supports token streaming. Which client-side change fixes the perceived latency?

A.Send the request from a background thread and update the user interface only after the full response has been parsed.
B.Lower the maximum token limit so the model finishes generating the complete answer sooner.
C.Set the request parameter that disables streaming to false and read server-sent events incrementally as each chunk arrives.
D.Enable streaming on the request and render each incremental delta as it is received instead of awaiting the full response body.
AnswerD

Streaming returns the completion as a sequence of incremental deltas over a long-lived connection, so the client can paint tokens as they arrive. The user perceives the first token within the prefill time rather than after full generation. This directly removes the freeze and is the standard fix for chat interfaces against streaming-capable endpoints.

Why this answer

Streaming endpoints deliver the completion incrementally, so the client must both request streaming and render each delta as it arrives. That shifts perceived latency from full generation time down to time to first token. Capping tokens, disabling streaming, or merely threading the call all leave the user staring at a blank interface until the whole answer exists.

Exam trap

The trap here is confusing non-blocking execution with incremental rendering: a background thread keeps the app responsive but still shows nothing until the full response is parsed.

29
Multi-Selectmedium

A developer is packaging a NeMo-based LLM application into a container for deployment on an NVIDIA GPU node. Which two practices are required to ensure the container can access the GPU and run inference efficiently? (Choose two.)

Select 2 answers
A.Install the NVIDIA Container Toolkit on the host and run the container with the --gpus flag or equivalent runtime configuration.
B.Pin the CUDA runtime and relevant libraries in the image to versions compatible with the host driver and NeMo release.
C.Set the container's shared memory size to the minimum allowed value to conserve RAM.
D.Disable NVIDIA persistence mode on the host to reduce memory overhead.
E.Bake the full NVIDIA data center driver into the container image.
AnswersA, B

The NVIDIA Container Toolkit exposes host GPU devices and driver libraries to containers. Without it and the corresponding runtime flag, the container sees no GPU and falls back to CPU or fails. This is a prerequisite for any GPU-accelerated NeMo inference workload on the node.

Why this answer

GPU access in containers requires the NVIDIA Container Toolkit on the host plus a runtime flag, and efficient NeMo inference requires matching CUDA runtime, driver, and framework versions. Bundling drivers, disabling persistence mode, or shrinking shared memory does not enable GPU access and can degrade performance.

Exam trap

The trap here is thinking the driver belongs inside the image, when in reality the host driver is injected by the NVIDIA Container Toolkit and only the user-space CUDA runtime should be pinned in the image.

30
MCQmedium

A developer is writing a Python client for an NVIDIA-hosted LLM endpoint and needs the model to answer every request in a strict JSON schema without extra prose. Where should the formatting contract be expressed so it applies consistently across all requests from the service?

A.In the assistant-role message, since the assistant produces the JSON output.
B.In a system-role message describing the required schema and the rule against extra prose.
C.In an HTTP header on the request, such as X-Response-Schema.
D.Appended to each user message as a trailing reminder sentence.
AnswerB

The system role carries instructions that apply to the whole conversation and take precedence over user turns, making it the right place for a persistent formatting contract. Every request built by the service reuses the same system message, so schema and no-prose rules stay consistent without repeating them in each user prompt.

Why this answer

A service-wide formatting contract belongs in the system-role message, which the endpoint treats as standing instructions for the whole conversation and which the client can reuse on every call. User-turn reminders, assistant-role content, and custom headers either dilute the instruction, misrepresent the turn, or are simply ignored.

Exam trap

The trap here is reaching for a custom HTTP header or an assistant message as a configuration channel, when behavioral instructions are only honored in the system and user roles.

31
MCQhard

Refer to the exhibit. What is the technical implication of using the specified 'fp8' precision mode in this model configuration?

A.It will disable Tensor Core utilization on the GPU.
B.It requires the model to be retrained from scratch.
C.It provides a better balance of accuracy and speed than INT8.
D.It is only compatible with CPU-based inference engines.
AnswerC

FP8 offers a wider dynamic range than INT8, which is fixed-point. This makes FP8 much more resilient to accuracy loss during quantization. It delivers the speed and memory efficiency benefits of low-bit arithmetic while maintaining performance levels closer to full FP16 or FP32 implementations.

Why this answer

FP8 precision takes advantage of the hardware-native support for 8-bit floating-point math in the latest NVIDIA GPU architectures (like Hopper). This mode provides a higher dynamic range than INT8 quantization, making it easier to maintain model accuracy while achieving superior throughput and memory efficiency. It is the current state-of-the-art for high-performance LLM deployment, balancing the need for speed with the requirement for high-fidelity generative output in large-scale enterprise services.

Exam trap

Candidates often assume FP8 is just a version of INT8. They fail to realize FP8 offers a superior dynamic range, which is why it is preferred for LLMs over the more restrictive integer formats.

32
MCQhard

A developer is deploying a TensorRT-LLM optimized model on NVIDIA Triton Inference Server. They observe that the first inference request takes significantly longer than subsequent ones. Which Triton feature should they configure to reduce this initial latency?

A.Instance groups
B.Dynamic batching
C.Model warmup
D.Response cache
AnswerC

Model warmup runs dummy inference requests during model loading to initialize CUDA contexts, allocate memory, and compile kernels. This moves the overhead from the first real request to the loading phase, reducing initial latency. Configuring warmup in the model's config.pbtxt is the standard solution for this scenario.

Why this answer

Model warmup explicitly runs dummy inferences at load time, forcing the model to initialize CUDA contexts, load weights, and compile kernels before any real request arrives. This eliminates the cold-start penalty observed on the first request. Other features like dynamic batching or caching improve different aspects of performance but not initial latency.

Exam trap

The trap here is confusing throughput optimizations like dynamic batching with latency-hiding techniques like warmup, which target different phases of the request lifecycle.

33
MCQeasy

Which NVIDIA SDK is specifically optimized for high-performance deep learning inference and supports the deployment of quantized models?

A.NVIDIA CUDA Toolkit.
B.NVIDIA TensorRT.
C.NVIDIA cuDNN.
D.NVIDIA NCCL.
AnswerB

TensorRT is the dedicated SDK for high-performance inference. It excels at optimizing model graphs, applying quantization, and selecting the most efficient CUDA kernels for specific hardware. It is the core tool for developers needing to maximize throughput and minimize latency for production-ready AI applications on NVIDIA hardware.

Why this answer

TensorRT is the industry-standard SDK for optimizing deep learning inference on NVIDIA GPUs. It provides advanced techniques like layer fusion, precision calibration (FP8, INT8), and kernel auto-tuning. For developers, mastering TensorRT is essential to transition from research code to high-speed, scalable production deployments, ensuring that models operate at peak efficiency while respecting the strict latency requirements of modern enterprise applications.

Exam trap

Exam takers frequently mix up training frameworks with dedicated inference SDKs, incorrectly choosing training-centric libraries when the question specifically asks for high-performance deployment optimization.

34
MCQmedium

When implementing a Guardrails layer in a generative AI application, what is the primary goal regarding model output?

A.To increase the total number of tokens generated per second.
B.To sanitize and validate content against predefined policies.
C.To compress the output text into a smaller format.
D.To provide persistent long-term memory for the LLM.
AnswerB

The primary role of guardrails is to check the output for prohibited content, tone issues, or factual inaccuracies based on organizational policy. By validating the response in real-time, the application ensures that the generative model adheres to safety standards before exposing the user to the content.

Why this answer

Guardrails are implemented to intercept and validate LLM outputs to ensure they align with safety, toxicity, and quality standards. This is essential for enterprise safety, preventing the model from generating harmful, inaccurate, or biased content before it reaches the end user. By establishing this layer, developers create a robust feedback loop that protects the application's reputation while maintaining the flexibility of the underlying generative model.

Exam trap

Candidates sometimes confuse Guardrails with 'model fine-tuning' or 'prompt engineering,' failing to recognize that Guardrails specifically act as an external validation layer to enforce safety policies on generated content.

35
MCQmedium

A developer is tuning a retrieval-augmented generation pipeline that uses NVIDIA NIM embeddings and a NIM LLM. Latency is dominated by embedding thousands of document chunks at query time because the team re-embeds the whole corpus on every request. Which change most directly fixes the architecture?

A.Cache the final LLM answers in memory and serve repeats, leaving the per-request embedding of the corpus unchanged.
B.Switch the embedding NIM to a larger model with higher dimensionality so fewer chunks are needed to cover the corpus.
C.Increase the LLM's max_tokens so it can summarize the raw documents directly instead of retrieving chunks.
D.Precompute and persist chunk embeddings in a vector index once, then embed only the incoming query and retrieve nearest neighbors at request time.
AnswerD

The latency comes from re-embedding the full corpus per request, which is wasteful because document chunks are static. Generating embeddings once and storing them in a vector index means each query only needs a single embedding call plus a nearest-neighbor lookup. This is the standard RAG pattern and removes the dominant cost while preserving retrieval quality.

Why this answer

When latency is driven by re-embedding static documents per request, the correct remedy is to embed the corpus once, store the vectors in an index, and embed only the query at serving time. Larger embedding models, longer generation, and answer caches do not remove the redundant full-corpus embedding work and can even increase cost or staleness.

Exam trap

The trap here is treating an answer cache or a bigger embedding model as the fix, when the real problem is embedding static documents repeatedly instead of once.

36
MCQmedium

Which TWO of the following are primary benefits of using NVIDIA Triton Inference Server for deploying generative AI models?

A.It enforces a strict requirement to use only PyTorch-based model backends.
B.It provides native support for dynamic batching of inference requests.
C.It offers built-in support for concurrent model execution.
D.It automatically converts all models to the proprietary CUDA-Graph format.
E.It removes the need for any containerization technology.
AnswerB, C

Dynamic batching allows Triton to combine individual inference requests into a single batch, significantly improving throughput on GPUs. This feature is essential for production environments where requests arrive asynchronously, as it maximizes the utilization of hardware resources compared to processing requests one by one.

Why this answer

NVIDIA Triton simplifies the deployment of complex AI models by supporting multiple frameworks and optimizing inference. Its ability to handle concurrent model execution and dynamic batching is critical for maximizing GPU utilization in production environments. Mastering these features allows developers to build scalable, high-performance pipelines that serve multiple models efficiently without needing custom backend infrastructure for every specific model architecture.

Exam trap

Test-takers frequently select general cloud orchestration benefits instead of Triton-specific inference acceleration features like native dynamic batching and concurrent multi-model execution.

37
MCQhard

A developer is using NVIDIA Triton Inference Server to deploy a TensorRT-LLM optimized model. The model must support multiple concurrent users with low latency. The developer notices that latency spikes when many requests arrive simultaneously. Which Triton feature should be configured to improve throughput while maintaining acceptable latency?

A.Sequence batching to maintain state across multiple inference requests for stateful models.
B.Model warmup with sample inputs to preload CUDA kernels and reduce cold-start latency.
C.Dynamic batching with a `max_batch_size` and `preferred_batch_size` in the model configuration.
D.Instance groups with multiple model instances per GPU to increase parallelism.
AnswerC

Triton's dynamic batching automatically groups incoming inference requests into batches to improve GPU utilization and throughput. By setting `max_batch_size` and `preferred_batch_size`, the server can form batches that fit within latency constraints. For TensorRT-LLM models, this is crucial for handling concurrent users efficiently, as it reduces the number of forward passes and amortizes overhead, directly addressing latency spikes under load.

Why this answer

Triton's dynamic batching is essential for handling concurrent requests efficiently. By configuring `max_batch_size` and `preferred_batch_size`, the server can group multiple inference requests into a single batch, reducing the number of forward passes and improving GPU utilization. This leads to higher throughput and lower latency under load, which is critical for serving TensorRT-LLM models to multiple users.

Other features like warmup or multiple instances address different aspects but not the core batching need.

Exam trap

The trap here is confusing model warmup or instance groups with dynamic batching, which is the primary mechanism to handle concurrent request throughput in Triton.

38
Multi-Selecthard

A developer is preparing a container for an LLM microservice that will run on an NVIDIA GPU node and must be deployable through NVIDIA NIM. They want the image to be portable across supported GPU generations while still using NVIDIA's optimized inference stack. Which two practices should they follow? (Choose two.)

Select 2 answers
A.Base the service on an NGC-hosted NIM container image and pin a specific version tag.
B.Pass the GPU through with the NVIDIA Container Toolkit so the container can see the device.
C.Compile the model into a TensorRT engine for one exact GPU SKU and ship only that engine.
D.Install a full CUDA toolkit and matching driver inside the image for maximum portability.
E.Bake the model weights directly into the image layer during the build step.
AnswersA, B

NGC-published NIM images already bundle the optimized runtime, CUDA libraries, and model-serving layer that NIM expects. Pinning an explicit version tag instead of latest keeps deployments reproducible and lets the team roll forward deliberately when a new image is validated against their traffic.

Why this answer

Portable NIM deployment rests on using the NGC-published, version-pinned NIM image and exposing the GPU through the NVIDIA Container Toolkit. Bundling weights or drivers into the image, or shipping one SKU-specific TensorRT engine, either breaks the supported deployment model or destroys the cross-generation portability the team requires.

Exam trap

The trap here is assuming that installing drivers or precompiling engines inside the image increases portability, when both actually tie the artifact to a specific host or GPU SKU.

39
MCQhard

A team fine-tunes a Llama-3 8B model with NVIDIA NeMo Framework and must ship an inference artifact that a C++ service can load without a Python runtime. They want maximum throughput on Hopper GPUs and plan to serve many concurrent requests with in-flight batching. Which artifact and runtime pairing best satisfies these constraints?

A.Convert the checkpoint to ONNX and execute it with the ONNX Runtime CUDA execution provider.
B.Save the fine-tuned checkpoint in NeMo .nemo format and load it with the NeMo inference toolkit in a Python worker.
C.Export a TensorRT-LLM engine and run it through the TensorRT-LLM C++ runtime with in-flight batching enabled.
D.Merge the LoRA weights into the base model and serve the merged Hugging Face checkpoint with the standard transformers pipeline.
AnswerC

TensorRT-LLM compiles the fine-tuned weights into a GPU-specific engine and ships a C++ runtime that supports in-flight (continuous) batching, so new requests join an active batch without waiting. This removes the Python dependency, targets Hopper kernels at compile time, and is the intended path for high-throughput production serving.

Why this answer

The requirement combination of no Python runtime, Hopper-tuned kernels, and in-flight batching points to a TensorRT-LLM engine executed by its C++ runtime. Checkpoint formats, ONNX, and the transformers pipeline each fail at least one of the three constraints, most commonly the C++ deployment or the continuous batching capability.

Exam trap

The trap here is treating any GPU-capable runtime as equivalent, when only the TensorRT-LLM C++ path combines a Python-free deployment with in-flight batching on Hopper.

40
MCQhard

A developer is profiling a TensorRT-LLM serving deployment and notices that throughput collapses once concurrent requests exceed a small number of users, even though GPU compute utilization stays low. The model uses paged KV cache and continuous batching. Which factor most likely explains the bottleneck?

A.The tokenizer runs on the GPU and competes with the model for SM cycles, capping the batch size.
B.TensorRT-LLM lacks continuous batching support and therefore serializes every request regardless of the KV cache size.
C.The GPU is thermally throttled, which reduces clock speed and therefore limits throughput at high concurrency.
D.The KV cache memory pool is too small, so the scheduler cannot admit more concurrent sequences and requests queue while the GPU idles.
AnswerD

With paged KV cache, each active sequence consumes blocks from a fixed GPU memory pool. If the pool is undersized, the scheduler must limit the number of in-flight sequences, so additional requests wait even though SM compute is underused. Low compute utilization combined with throughput saturation at low concurrency is the signature of KV cache capacity, not arithmetic throughput.

Why this answer

Paged KV cache allocates GPU memory in blocks per sequence, and the in-flight batch size is bounded by how many blocks the memory pool can hold. When that pool is too small, the scheduler queues new requests and the GPU sits underutilized, which matches the described symptom. Missing batching, thermal throttling, and GPU tokenization do not explain low compute utilization with a hard concurrency ceiling.

Exam trap

The trap here is equating low GPU utilization with a compute problem, when low utilization plus a low concurrency ceiling usually points to KV cache memory capacity.

41
MCQmedium

A developer is optimizing a retrieval-augmented generation (RAG) pipeline using NVIDIA TensorRT-LLM. They notice excessive latency during the document retrieval phase before the generation starts. Which optimization strategy is most effective for this bottleneck?

A.Increase the number of transformer layers in the LLM.
B.Implement a GPU-accelerated vector database for similarity search.
C.Use float64 precision for all vector embedding calculations.
D.Decrease the context window size of the retrieval model.
AnswerB

Moving from CPU-based vector indexing to GPU-accelerated solutions like Faiss on NVIDIA hardware drastically reduces search latency. This optimization allows for parallel processing of vector embeddings, which is crucial when handling large datasets in RAG pipelines, effectively offloading the retrieval bottleneck from the host CPU.

Why this answer

Latency in RAG pipelines often stems from inefficient embedding lookups or serial processing. Moving vector search to a GPU-accelerated database or utilizing a cross-encoder for re-ranking ensures the model receives highly relevant chunks. This approach balances retrieval precision with speed, ensuring the LLM receives context without stalling the inference server, which is critical for real-time generative applications.

Exam trap

Candidates often assume the bottleneck is in the LLM generation itself, failing to realize that slow retrieval (embedding search) is a common, distinct performance killer in RAG pipelines.

42
MCQmedium

A developer is building a document summarization service using an NVIDIA NIM microservice for Llama-3. The service must process batches of 20 documents at once to maximize throughput. The NIM container is already running with default settings. Which API parameter should the developer configure to enable efficient batched inference?

A.Set the `batch_size` field in the request payload to 20, allowing the NIM to process all documents in a single forward pass.
B.Enable the `tensor_parallel` option in the request headers and set it to 20, so the NIM distributes the batch across multiple GPUs.
C.Use the `stream` parameter set to `true` and send each document as a separate request; the NIM will automatically coalesce them into a batch on the server side.
D.Configure the `max_batch_size` parameter when starting the NIM container, and send multiple prompts in a single request using the `prompt` field as an array.
AnswerD

NVIDIA NIM microservices support dynamic batching at the server level, controlled by the `max_batch_size` parameter set during container launch. To process a batch, the client sends a single request with the `prompt` field as an array of strings. The server then groups these into a batch, improving GPU utilization and throughput for document summarization.

Why this answer

To efficiently process a batch of documents with an NVIDIA NIM microservice, the developer must ensure the server is configured with an appropriate `max_batch_size` at launch. Then, the client should send a single request containing an array of prompts. This leverages the NIM's dynamic batching capabilities, which group concurrent requests or batched prompts to maximize GPU utilization and throughput.

Other options either misuse parameters or misunderstand the batching mechanism.

Exam trap

The trap here is assuming that batching is controlled by a client-side parameter like `batch_size` in the request body, rather than by server-side configuration and sending an array of prompts.

43
MCQeasy

A developer is using the NVIDIA API Catalog to test a Llama-3 model via its API endpoint. They need to send a request that includes a system prompt to set the model's behavior. Which component of the request payload is used to provide the system prompt?

A.A message with `role` set to `system` within the `messages` array.
B.The `prompt` field, prefixed with `System:` followed by the instruction.
C.The `system` field at the top level of the JSON payload.
D.The `instructions` parameter in the request headers.
AnswerA

The NVIDIA API Catalog endpoints for chat models expect a `messages` array containing objects with `role` and `content`. To provide a system prompt, include a message with `role` set to `system` and the desired instruction in `content`. This is a standard format for chat-based LLM APIs and ensures the model interprets the instruction as a system-level directive.

Why this answer

In the NVIDIA API Catalog, chat models accept a `messages` array where each message has a `role` and `content`. To set a system prompt, include a message with `role` set to `system`. This is consistent with common chat completion APIs and ensures the model treats the content as a system-level instruction that guides its responses throughout the conversation.

Exam trap

The trap here is assuming a top-level `system` field or a special header, rather than using the standard `messages` array with a system role.

44
MCQhard

A developer is using TensorRT-LLM to build a chatbot and wants to reduce the memory footprint of the KV cache during inference. Which technique should they use?

A.Disable quantization
B.Increase the beam width
C.Use a larger batch size
D.Enable paged KV cache
AnswerD

Paged KV cache, a feature in TensorRT-LLM, manages the KV cache in fixed-size blocks (pages) that can be allocated and freed dynamically. This reduces memory fragmentation and allows more efficient memory usage, especially for variable-length sequences. It is specifically designed to reduce KV cache memory footprint, making it the correct choice.

Why this answer

Paged KV cache is a memory management technique in TensorRT-LLM that divides the KV cache into fixed-size blocks, allowing non-contiguous storage and dynamic allocation. This reduces fragmentation and enables more sequences to fit in memory. It is the recommended approach to minimize KV cache memory footprint during inference.

Exam trap

The trap here is assuming that increasing batch size or beam width improves efficiency, but both increase KV cache memory usage.

45
MCQmedium

Which NVIDIA technology enables efficient cross-GPU communication during Tensor Parallelism for large-scale model inference?

A.CUDA Stream Multi-Processor (SM) scheduling.
B.NVLink.
C.cuDNN Convolutional Layers.
D.TensorRT-LLM PagedAttention.
AnswerB

NVLink provides a high-speed direct interconnect between GPUs, allowing them to share data much faster than standard PCIe. This low-latency communication is essential for Tensor Parallelism, where GPU layers must constantly exchange activations to perform synchronized matrix multiplications, directly impacting inference speed in multi-GPU configurations.

Why this answer

Tensor Parallelism involves splitting model layers across multiple GPUs, which requires high-bandwidth, low-latency communication to synchronize activation values. NVIDIA's NVLink and NVSwitch technologies are designed specifically for this purpose, overcoming the limitations of standard PCIe buses. This is critical for maintaining performance in models that are too large to fit into a single GPU, enabling seamless, high-speed distributed computation.

Exam trap

Candidates often select general networking protocols like Ethernet or standard PCIe, ignoring that NVLink is the specific NVIDIA interconnect required for high-bandwidth communication between GPUs in parallel settings.

46
Multi-Selecthard

Which TWO of the following practices are recommended when using NVIDIA Triton Inference Server to maximize throughput for a concurrent multi-model deployment?

Select 2 answers
A.Configure multiple instance groups per model for concurrency.
B.Enable dynamic batching in the model configuration.
C.Disable all logging to reduce disk I/O latency.
D.Set the batch size to 1 for all incoming requests.
E.Use the default backend for every model type.
AnswersA, B

Creating multiple instance groups allows the server to spawn several concurrent execution units for a single model. This capability is vital for parallelizing requests across multiple GPU streams, effectively hiding latency and increasing total throughput when the hardware has sufficient spare capacity to handle additional workloads.

Why this answer

Optimizing Triton involves managing hardware resources and execution concurrency. By configuring instance groups, developers can ensure that multiple model instances are pre-loaded to saturate GPU compute. Additionally, using dynamic batching allows the server to aggregate individual requests into larger batches, which is essential for maximizing GPU utilization during high-traffic periods.

These strategies are fundamental for scaling LLM services in production environments where resource contention is a primary concern.

Exam trap

Test-takers frequently confuse dynamic batching with static batching or mistakenly think that adding more hardware nodes alone replaces the need for optimized instance group configurations on a single server instance.

47
MCQhard

Refer to the exhibit. An engineer receives this error during deployment. What is the most likely cause?

A.The model file is corrupted and missing.
B.Another process is already using the target GPU.
C.The system lacks the required RAM for the CPU.
D.The model version is incompatible with the server.
AnswerB

This error is the standard response when the GPU is locked by another application or container. In a multi-tenant environment, this often happens if resources are not correctly partitioned, preventing the current inference service from acquiring the device handle required to load the model into VRAM.

Why this answer

The error indicates that another process or container has locked the GPU device, preventing the Triton server from initializing the model. In production environments, managing GPU resource allocation is crucial. If multiple processes compete for the same hardware without proper resource isolation, initialization will fail, causing downtime.

Resolving this requires checking for conflicting processes or ensuring proper container resource limits to prevent GPU resource contention.

Exam trap

Candidates often mistake hardware memory exhaustion errors for software configuration issues or missing library dependencies, ignoring the direct system message indicating that another active process currently occupies the target GPU.

48
Multi-Selecthard

A developer is containerizing an inference service built with TensorRT-LLM and NVIDIA NIM for a Kubernetes cluster. They want the deployment to start reliably and use the GPU efficiently. Which two practices should they follow? (Choose two.)

Select 2 answers
A.Request nvidia.com/gpu resources in the pod spec and ensure the NVIDIA GPU Operator or device plugin is installed so the scheduler can allocate the device.
B.Disable the NVIDIA container runtime so the pod uses the host's default runtime and accesses the GPU through /dev/nvidia devices directly.
C.Set the pod's restartPolicy to Always and rely on CrashLoopBackOff to warm the GPU driver on each restart.
D.Mount the TensorRT-LLM engine files and model weights from a persistent volume or baked image layer rather than downloading them at pod startup.
E.Pin the pod to a CPU-only node pool and use CUDA emulation to run the TensorRT-LLM engine.
AnswersA, D

Kubernetes does not expose GPUs to pods unless the NVIDIA device plugin or GPU Operator advertises nvidia.com/gpu as a schedulable resource. Without that resource request, the container starts without a GPU and inference fails. Declaring the resource also lets the scheduler place the pod on a node with a free device, which is fundamental to reliable startup.

Why this answer

Reliable GPU-backed startup requires that Kubernetes can schedule the device through the NVIDIA device plugin or GPU Operator, and that the heavy engine artifacts are already present locally rather than downloaded at boot. Restart loops, bypassing the NVIDIA container runtime, and CPU-only emulation all prevent the engine from loading or waste resources instead of improving reliability and GPU efficiency.

Exam trap

The trap here is thinking that a restart policy or CrashLoopBackOff can substitute for proper GPU resource advertisement and pre-staged engine files.

49
MCQhard

A team fine-tunes a Llama model with NVIDIA NeMo Framework and must serve it behind an OpenAI-compatible endpoint with no Python glue code. They want the adapter weights kept separate from the base model so several adapters can share one loaded base. Which deployment approach fits these constraints?

A.Export the merged model to a TensorRT-LLM engine and serve it through NVIDIA Triton Inference Server with the TensorRT-LLM backend.
B.Package the fine-tuned model inside a custom FastAPI service that wraps the NeMo checkpoint and exposes chat completions.
C.Deploy the base model on NVIDIA NIM and load the fine-tuned adapter at runtime through multi-LoRA support in the NIM container.
D.Convert the adapter to ONNX and deploy it with NVIDIA Triton Inference Server using the Python backend for each adapter.
AnswerC

NIM microservices expose OpenAI-compatible APIs without custom glue code, and their multi-LoRA capability lets a single loaded base model host multiple adapters selected per request. This keeps adapter weights separate from the base, satisfies the shared-base requirement, and avoids building per-adapter engines. It directly matches every stated constraint.

Why this answer

The constraints are an OpenAI-compatible endpoint, no custom Python service code, and adapter weights stored separately from a shared base model. NIM's multi-LoRA serving satisfies all three: the container exposes standard APIs and loads adapters against one resident base. Merged engines, custom wrappers, and per-adapter Python backends each break at least one stated requirement.

Exam trap

The trap here is treating any OpenAI-compatible endpoint as sufficient and overlooking that merging the adapter destroys the ability to share one base across several adapters.

50
MCQmedium

Refer to the exhibit. An engineer is tuning a deployment config. Why is 'enable_cuda_graph' set to true in this JSON configuration?

A.To increase the maximum batch size to 256.
B.To reduce CPU overhead during repetitive kernel launches.
C.To convert the model precision from fp16 to fp8.
D.To force the model to use the CPU for inference.
AnswerB

By capturing the graph of operations, the driver can execute them with a single launch command. This avoids the overhead of traversing the command queue for every operation, which is highly beneficial for LLM inference where the execution pattern is consistent.

Why this answer

CUDA Graphs capture a sequence of GPU work as a single graph, reducing CPU overhead associated with kernel launches. This is critical for LLMs where many small kernel calls can lead to CPU-bound execution. In high-performance generative AI scenarios, reducing launch latency is essential to ensure that the GPU remains saturated with work, thereby maximizing tokens-per-second and reducing total request latency for end-users.

Exam trap

Candidates often assume CUDA Graphs are used for distributed multi-node communication or automatic mixed precision, missing their actual purpose of eliminating CPU launch overhead.

51
MCQhard

A developer is debugging a RAG service where answers are correct in testing but degrade in production as the document corpus grows. Logs show retrieval returning chunks with high similarity scores that do not contain the answer. Which change most directly addresses the root cause?

A.Increase the embedding model's output dimension to improve similarity precision.
B.Increase the number of retrieved chunks passed to the LLM from five to twenty.
C.Replace dense embedding retrieval with a hybrid approach that combines dense vectors and keyword matching, then rerank the merged candidates.
D.Lower the similarity threshold so fewer chunks are returned to the model.
AnswerC

High similarity scores without the answer indicate dense embeddings are matching topical similarity rather than the specific terms that carry the answer. Hybrid retrieval adds lexical matching so exact names, numbers, and rare tokens are captured, and a reranker reorders the merged set by relevance to the actual query. This directly improves precision at the top of the retrieved list, which is what the failure mode requires.

Why this answer

Confident but non-answering chunks point to dense retrieval optimizing topical similarity instead of matching the specific terms that carry the answer. Hybrid retrieval adds lexical matching for exact names, numbers, and rare tokens, and reranking the merged candidates pushes truly relevant chunks to the top. This improves precision where it matters, at the head of the retrieved list passed to the model.

Exam trap

The trap here is reading a high similarity score as evidence of relevance, when dense similarity can be high for passages that share topic but not the answer.

52
MCQhard

A team is building a RAG assistant and wants to reduce hallucinated citations. They plan to have the LLM return structured output that names the source document chunk used for each claim. Which implementation strategy most directly improves the reliability of that structured output?

A.Retrieve more chunks and place them all in the prompt without identifiers.
B.Increase the model temperature so the model explores more citation candidates.
C.Ask the model to cite sources in a free-text bibliography appended after the answer.
D.Pass the retrieved chunks with stable identifiers in the prompt and constrain generation to a JSON schema that references those identifiers.
AnswerD

Giving the model explicit, stable chunk identifiers in context makes attribution a selection task rather than an invention task. Constraining decoding to a JSON schema that requires an identifier field prevents free-form citation text and makes the output machine-checkable. The application can then verify that every cited identifier exists in the retrieved set, catching unsupported claims before they reach the user.

Why this answer

Reliable attribution comes from making the citation a constrained choice among known identifiers. Supplying stable chunk IDs in context and enforcing a JSON schema means the model selects from real evidence instead of generating citation text. The application can then validate each returned identifier against the retrieved set, converting an unverifiable prose claim into a programmatically checkable reference.

Exam trap

The trap here is assuming that asking the model politely for citations produces trustworthy citations, when only constrained decoding against known identifiers makes them verifiable.

53
MCQmedium

A developer is building a customer-support assistant on NVIDIA NIM microservices. After a model update, responses that previously arrived in under 300 ms now take over two seconds, and the streaming client shows a long pause before the first token appears. GPU utilization is low and the prompt template was not changed. Which action should the developer take first to diagnose the regression?

A.Compare the time to first token (TTFT) and inter-token latency measured before and after the update to isolate whether the delay is in prefill or decode.
B.Increase the tensor-parallel degree of the NIM container so the model weights are sharded across more GPUs.
C.Switch the client from streaming to a single non-streaming request so the full response is timed end to end.
D.Enable dynamic batching in the NIM microservice so concurrent requests are grouped into larger inference batches.
AnswerA

Splitting latency into time to first token and inter-token latency immediately separates prefill (prompt processing) from decode (token generation). A long pause before the first token with fast subsequent tokens points at prefill, while slow steady output points at decode. This measurement requires no code change and directly targets the reported symptom, making it the correct first diagnostic step for the NIM deployment.

Why this answer

The reported symptom is a latency regression with low GPU utilization and unchanged prompts, so the developer must localize the delay before changing deployment configuration. Instrumenting time to first token and inter-token latency cleanly separates prompt prefill from token decode and reveals which phase regressed. Batching, parallelism, and disabling streaming either mask the signal or alter behavior, so measurement comes first.

Exam trap

The trap here is assuming that low GPU utilization automatically means the fix is more GPUs or tensor parallelism, when the real question is which inference phase became slow.

54
Multi-Selecthard

When deploying a Large Language Model using TensorRT-LLM, which TWO configuration factors must be tuned to maximize KV cache efficiency?

Select 2 answers
A.KV cache block size.
B.Model quantization bit-width.
C.Maximum number of blocks.
D.Input prompt token limit.
E.GPU clock speed frequency.
AnswersA, C

The KV cache block size determines how memory is partitioned for attention heads. Setting an optimal block size balances memory fragmentation and allocation overhead, preventing wasted VRAM that would otherwise be unusable for storing new tokens, thus maximizing the total sequence capacity available for concurrent users.

Why this answer

Maximizing KV cache efficiency is essential for scaling inference performance and reducing memory fragmentation. By tuning the block size and the number of blocks allocated, developers prevent out-of-memory errors and ensure smooth handling of long-context requests. This is a foundational skill for engineers deploying generative models in production, as it directly influences the number of concurrent requests the GPU can manage effectively during peak usage.

Exam trap

Candidates often focus on model weights or batch sizes, overlooking that the KV cache is the primary bottleneck for memory consumption and concurrent request handling in TensorRT-LLM.

55
Multi-Selecthard

A developer is tuning a TensorRT-LLM deployment of a long-context chat model and observes that GPU memory is exhausted under concurrent requests, causing requests to be rejected. They want to reduce KV cache memory pressure without retraining the model. (Choose two.)

Select 2 answers
A.Reduce the model's hidden dimension by pruning attention heads and rebuild the engine from the pruned checkpoint.
B.Enable paged KV cache so cache blocks are allocated dynamically instead of reserving a contiguous buffer per sequence.
C.Apply KV cache quantization so cached keys and values are stored in a lower-precision format such as INT8 or FP8.
D.Disable continuous batching so each request is processed to completion before the next one begins.
E.Increase the maximum number of batched tokens so more requests are processed simultaneously in a single forward pass.
AnswersB, C

Paged KV cache breaks the cache into blocks allocated on demand and shared through a block manager, which greatly reduces internal and external fragmentation compared with reserving a maximum-length contiguous buffer for every sequence. Under concurrent long-context traffic this raises the number of sequences that fit in the same memory budget, directly relieving the exhaustion the developer observes without any retraining.

Why this answer

KV cache memory dominates long-context serving, so the effective levers are how cache memory is allocated and how many bytes each cached token consumes. Paged KV cache removes fragmentation by allocating blocks on demand, and KV cache quantization reduces per-token storage precision. Both are runtime or build-time configuration changes that need no retraining, unlike architectural pruning, and both increase the number of concurrent sequences the GPU can hold.

Exam trap

The trap here is treating increased batching or disabled batching as memory optimizations, when both actually change concurrency in ways that do not reduce the KV cache footprint per sequence.

56
MCQeasy

A developer is using the NVIDIA API Catalog to experiment with a hosted LLM. They want to send a prompt and receive a completion. Which endpoint should they use?

A./v1/chat/completions
B./v1/models
C./v1/completions
D./v1/embeddings
AnswerA

The /v1/chat/completions endpoint is the standard OpenAI-compatible API for chat-based interactions. It accepts a messages array and returns model-generated responses. NVIDIA API Catalog endpoints follow this convention, making it the correct choice for sending prompts and receiving completions.

Why this answer

The /v1/chat/completions endpoint is designed for conversational LLMs and accepts a structured messages array. It is the primary interface for interacting with hosted models on NVIDIA API Catalog. Other endpoints serve different purposes like listing models or generating embeddings, and the legacy completions endpoint is less suitable for chat-based models.

Exam trap

The trap here is assuming any completions endpoint works, but the chat completions endpoint is the correct one for modern conversational LLMs.

57
MCQmedium

When utilizing NVIDIA NIM for deployment, why is it recommended to use a containerized environment?

A.To bypass the need for CUDA drivers on the host machine.
B.To ensure dependency consistency and portability.
C.To improve the GPU's clock speed by optimizing kernel distribution.
D.To automatically optimize the model's weight distribution for multi-GPU setups.
AnswerB

AI models require specific versions of libraries (e.g., specific CUDA versions). Containerization bundles these dependencies, ensuring that the environment is reproducible and portable. This eliminates version conflicts and configuration drift, allowing the same microservice to run reliably across local workstations, testing clusters, and cloud production environments.

Why this answer

Containerization provides a consistent runtime environment across development, testing, and production, which is crucial for managing the complex dependencies of AI stacks like CUDA, cuDNN, and TensorRT. This approach minimizes 'works on my machine' issues and ensures that the model performance remains identical regardless of the underlying host configuration, making it the industry standard for deploying high-performance generative AI models at scale.

Exam trap

Test-takers often guess that containerization is primarily used for security isolation or cloud billing simplification, missing its core role in resolving complex CUDA and driver dependency issues.

58
MCQhard

A developer is evaluating a fine-tuned LLM with NVIDIA NeMo and observes that evaluation loss keeps decreasing while downstream task accuracy plateaus and then declines. Which action should the developer take to address this?

A.Extend training for more epochs so the loss can converge further.
B.Increase the learning rate to escape the plateau.
C.Apply early stopping based on the downstream validation metric and consider regularization such as dropout or weight decay.
D.Switch the evaluation metric to training loss for consistency.
AnswerC

Falling loss with declining task accuracy indicates overfitting to the training distribution. Early stopping on the validation metric halts training at the best generalizing point, while dropout or weight decay constrain the model. Together they restore alignment between optimization loss and real task performance.

Why this answer

A widening gap between decreasing evaluation loss and declining downstream accuracy is the classic signature of overfitting. Early stopping on the task metric preserves the best generalizing checkpoint, and regularization techniques such as dropout or weight decay constrain the model so optimization progress translates into real task gains.

Exam trap

The trap here is treating falling loss as proof of improvement, when declining task accuracy reveals overfitting that more training or a higher learning rate would worsen.

59
MCQmedium

A developer is packaging a fine-tuned Llama 3 model as a TensorRT-LLM engine for an on-premises inference service. The model was trained with a custom tokenizer that adds four new special tokens beyond the base vocabulary. When the engine is built and the service is started, the model outputs garbled text and repeats the same fragment regardless of the prompt. Which action should the developer take to resolve this?

A.Convert the fine-tuned checkpoint back to a Hugging Face format and serve it through a generic Python HTTP wrapper instead of a compiled engine.
B.Increase the engine's maximum batch size and rebuild the engine so the additional special tokens can be processed in parallel.
C.Rebuild the TensorRT-LLM engine using the fine-tuned checkpoint's updated tokenizer vocabulary and matching special-token IDs.
D.Set the sampling temperature to zero and add a repetition penalty in the runtime generation config to stabilize the output.
AnswerC

The garbled, repeating output is the classic signature of an input token ID mapping mismatch. The fine-tuned checkpoint extends the vocabulary by four tokens, so the engine must be built with that same tokenizer configuration and the same special-token ID assignments. Building the engine from the updated tokenizer keeps the prompt-to-embedding mapping consistent with training and restores coherent generation.

Why this answer

Garbled, looping text from a fine-tuned model almost always traces back to a tokenizer or vocabulary mismatch between training and inference. Because the fine-tuned checkpoint added special tokens, the TensorRT-LLM engine must be compiled against the checkpoint's tokenizer metadata so token IDs align with the embedding table. Rebuilding the engine with the updated vocabulary is the only option that restores correct prompt encoding.

Exam trap

The trap here is blaming generation settings such as temperature or repetition penalty for output corruption that actually originates from a tokenizer vocabulary mismatch.

60
MCQeasy

A team wants to load and run an optimized quantized LLM entirely inside a Python application with minimal dependencies, using a single high-level API that handles engine building and generation. They are not deploying a network service. Which component of the NVIDIA software stack is designed for this use case?

A.NVIDIA NIM microservices, which package models as containerized endpoints with standardized inference APIs.
B.NVIDIA DALI, which provides a data loading and augmentation library primarily for computer vision preprocessing pipelines.
C.NVIDIA Triton Inference Server, which loads models from a model repository and exposes HTTP and gRPC endpoints.
D.NVIDIA TensorRT-LLM's Python LLM API, which wraps engine build and generation behind a high-level runtime interface.
AnswerD

The TensorRT-LLM Python LLM API is purpose-built for embedding optimized inference directly in a Python process. It exposes a high-level interface that handles engine construction and token generation without requiring a separate server. That matches the requirement of running a quantized model in-process with minimal dependencies and no network service, unlike serving-oriented components.

Why this answer

When the goal is to run an optimized, quantized LLM inside a Python process without standing up a service, the TensorRT-LLM Python LLM API is the fit. It abstracts engine building and token generation behind a high-level interface, avoiding client-server overhead. Serving platforms such as Triton or NIM and preprocessing libraries such as DALI solve different problems and would add unnecessary infrastructure or miss the requirement.

Exam trap

The trap here is assuming any NVIDIA inference component can run in-process, when serving platforms like Triton and NIM inherently require a network service.

61
MCQeasy

A developer has a working TensorRT-LLM engine and wants to expose it through NVIDIA Triton Inference Server so that multiple client applications can call it over HTTP and gRPC with a stable interface. Which Triton feature should they configure to serve the TensorRT-LLM engine as a backend?

A.An ensemble model that chains a preprocessing model, the TensorRT-LLM engine, and a postprocessing model.
B.A Python backend model that loads the engine and implements the inference logic in a custom script.
C.A custom C++ backend compiled against the TensorRT-LLM libraries and registered with Triton at startup.
D.The TensorRT-LLM backend, declaring the engine in the model repository with a config.pbtxt that sets the backend and engine directory.
AnswerD

Triton's TensorRT-LLM backend is purpose-built to load TensorRT-LLM engines and expose them through the standard HTTP and gRPC inference protocols. Placing the engine in the model repository and pointing config.pbtxt at the backend and engine directory gives clients a stable interface plus optimized batching and KV cache handling. This is the intended production path for serving TensorRT-LLM engines.

Why this answer

Triton's TensorRT-LLM backend is designed specifically to load TensorRT-LLM engines from the model repository and serve them over the standard HTTP and gRPC endpoints. Declaring the engine directory and backend in config.pbtxt gives clients a stable interface while retaining optimized in-flight batching and KV cache management. Custom backends and ensembles add complexity without providing the required serving capability.

Exam trap

The trap here is reaching for a custom or Python wrapper when an official Triton backend already targets TensorRT-LLM engines.

62
MCQmedium

Refer to the exhibit. If a developer increases the 'max_batch_size' in the JSON configuration, what is the primary expected trade-off in the system's performance metrics?

A.Reduced total system throughput.
B.Increased GPU memory fragmentation.
C.Higher latency for individual requests.
D.Improved accuracy of the underlying model.
AnswerC

As the maximum batch size increases, the system may wait longer to accumulate enough requests to fill the batch. This increased wait time at the start of the inference pipeline results in higher latency for the first few requests, which is a standard trade-off for maximizing overall throughput.

Why this answer

Increasing the batch size allows the GPU to process more requests in parallel, which typically increases total throughput. However, this often leads to higher individual request latency for earlier requests as they wait for the buffer to fill. This is a classic throughput-vs-latency trade-off in GPU computing.

For LLMs, this balance is crucial, as too large a batch can lead to memory exhaustion or unacceptable delays for users waiting for the initial tokens.

Exam trap

Candidates often confuse throughput with latency, mistakenly believing that increasing the maximum batch size improves responsiveness for every individual user rather than causing initial requests to queue longer.

63
MCQmedium

A developer is integrating a NeMo Guardrails configuration into an existing chatbot. They need to ensure that the LLM does not generate content related to unauthorized financial advice. Which mechanism should they implement to achieve this programmatic constraint?

A.Increase the temperature parameter to 1.0 to improve model creativity.
B.Utilize Colang to define input rails that intercept and block restricted topics.
C.Enable standard logging in the inference server to identify bad prompts.
D.Apply a secondary LLM to re-write every user prompt to be neutral.
AnswerB

Colang allows developers to define specific flows that detect intent and trigger predefined responses when sensitive topics are identified. By intercepting the user input before it reaches the LLM, the system can effectively block restricted queries, providing a deterministic layer of protection that standard prompting cannot guarantee reliably.

Why this answer

NeMo Guardrails uses Colang to define flows and dialogue rails that intercept and inspect the interaction between the user and the LLM. By defining specific canonical forms and guardrail flows, developers can force the model to refuse prompts that trigger sensitive topics. This approach is essential for production-grade applications where ensuring model safety and regulatory compliance is mandatory for deploying generative AI safely within enterprise environments.

Exam trap

Candidates often confuse NeMo Guardrails with vector database filtering or simple prompt engineering. They assume that adding 'do not give advice' to the system prompt is sufficient, ignoring the need for programmatic validation.

64
Multi-Selectmedium

A developer is preparing a RAG service that calls an NVIDIA-hosted LLM endpoint and must reduce hallucinations for questions whose answers are absent from the retrieved context. Which two practices should be applied in the application layer? (Choose two.)

Select 2 answers
A.Instruct the model in the system prompt to answer only from the supplied context and to state explicitly when the context is insufficient.
B.Increase the number of retrieved chunks to the model's full context limit so the answer is guaranteed to be somewhere in the prompt.
C.Cache every generated response and replay it for semantically similar questions to keep the answers consistent over time.
D.Apply a minimum relevance score to retrieved chunks and skip generation entirely when no chunk clears the threshold.
E.Set the sampling temperature to its maximum so the model explores a wider range of candidate answers and avoids repeating a single phrasing.
AnswersA, D

A grounding instruction constrains generation to the retrieved passages and gives the model a sanctioned way to decline, which is the cheapest and most direct hallucination control. It works with any hosted endpoint because it lives entirely in the request payload. Combined with a threshold on retrieval scores, it prevents the model from inventing an answer when nothing relevant was retrieved.

Why this answer

Hallucination in RAG is controlled by constraining generation and by refusing to generate when retrieval fails. A grounding system prompt gives the model permission to abstain and limits it to supplied evidence, while a relevance threshold stops irrelevant chunks from ever reaching the prompt. Together they address both the model's behavior and the quality of its input.

Exam trap

The trap here is assuming that more retrieved context or higher sampling temperature improves answer quality, when both actually increase the chance of unsupported output.

65
MCQhard

A developer is packaging a generative AI application for NVIDIA AI Enterprise deployment on Kubernetes. The application must run an LLM served by NVIDIA NIM, an embedding model, and a vector database, and must support rolling upgrades without dropping in-flight inference requests. Which design choice best meets these requirements?

A.Run each component as its own Kubernetes Deployment with readiness probes, multiple replicas, and a Service, using rolling update strategy and preStop hooks for graceful drain.
B.Use a single Deployment with a Recreate strategy so old pods are fully terminated before new ones start, guaranteeing no version mixing.
C.Deploy all three components in a single pod so they share a network namespace and can be upgraded together atomically.
D.Place the NIM LLM and embedding model behind one Service and rely on client-side retries to mask any requests lost during upgrades.
AnswerA

Separate Deployments let each tier scale and upgrade independently, readiness probes gate traffic until a replica is warm, and rolling updates replace pods gradually so capacity is maintained. A preStop hook plus termination grace period lets in-flight requests finish before shutdown. This directly satisfies the rolling-upgrade and no-dropped-requests requirements across the LLM, embedding, and database components.

Why this answer

Independent Deployments per component, each with readiness probes, multiple replicas, and a stable Service, allow rolling updates that replace pods gradually while healthy capacity remains. A preStop hook combined with a termination grace period lets the NIM LLM finish in-flight generations before the process exits. This architecture also lets the LLM, embedding, and vector database tiers scale separately according to their distinct resource needs.

Exam trap

The trap here is assuming that atomic or all-at-once upgrades simplify operations, when they actually guarantee downtime for in-flight inference.

66
MCQmedium

A developer is writing a Python service that calls an NVIDIA-hosted NIM endpoint for a Llama model. The service must recover gracefully when the endpoint returns HTTP 429 responses during peak traffic, without dropping user requests. Which implementation approach best satisfies this requirement?

A.Retry the request in a tight loop immediately after each 429 response until it succeeds.
B.Increase the client's request timeout value so the 429 responses stop occurring.
C.Implement exponential backoff with jitter and a bounded retry count, then surface a fallback error to the caller.
D.Switch the endpoint URL to a different NIM model and continue sending the same request payload unchanged.
AnswerC

HTTP 429 signals rate limiting, so the client should wait progressively longer between attempts and randomize the delay with jitter to avoid synchronized retry storms from many clients. A bounded retry count prevents indefinite blocking, and a defined fallback keeps the service responsive when retries are exhausted. This is standard resilient client behavior for hosted inference endpoints.

Why this answer

Rate-limit responses require the client to slow down and retry deliberately. Exponential backoff with jitter spreads retry attempts so multiple clients do not collide, and a bounded retry count keeps latency predictable. Because some requests may still fail after all retries, returning a controlled fallback error preserves a defined service contract for callers rather than hanging or crashing.

Exam trap

The trap here is treating a 429 as a transient network fault that should be retried instantly rather than as an explicit signal to reduce request rate.

67
MCQeasy

A developer is writing a Python service that calls a locally hosted NVIDIA NIM microservice for a Llama model. They want to keep the client code portable so the same class can later target NVIDIA's hosted API endpoints without rewrites. Which client approach fits this goal?

A.Invoke nvidia-smi from a subprocess and parse GPU utilization to infer model outputs.
B.Use the Triton client library with the gRPC protocol and hardcode the model name and version in every call site.
C.Call the TensorRT-LLM C++ runtime directly from Python with ctypes and load the engine file inside the service.
D.Use the OpenAI Python client pointed at the NIM base_url, since NIM exposes OpenAI-compatible /v1/chat/completions routes.
AnswerD

NIM microservices deliberately expose OpenAI-compatible endpoints, so an OpenAI-style client with a configurable base_url works against the local container and can be repointed at hosted endpoints by changing only the URL and key. This preserves portability and avoids vendor-specific SDK code, which is exactly what the developer wants for later migration.

Why this answer

Portability comes from targeting the OpenAI-compatible HTTP interface that NIM exposes, letting a single client class switch between a local container URL and NVIDIA hosted endpoints by configuration alone. Direct TensorRT-LLM bindings and the Triton gRPC client both couple code to a specific runtime or protocol, while GPU telemetry tools cannot serve inference requests.

Exam trap

The trap here is assuming that NVIDIA NIM requires a proprietary SDK, when its chat surface is intentionally OpenAI-compatible and configurable by base URL.

68
MCQmedium

A developer is building a generative AI application that uses an NVIDIA NIM microservice for a Llama 3 model. They need to persist the model's responses and associated metadata for later auditing. Which approach best integrates NIM with an external datastore?

A.Use the NIM's built-in PostgreSQL connector by setting the NIM_PG_CONNECTION environment variable.
B.Call the NIM's inference endpoint from the application code, then write the response and metadata to the datastore using the application's own database client.
C.Mount a shared volume into the NIM container and have the NIM write response files that the application later reads.
D.Enable the NIM's audit logging feature and configure it to forward logs directly to the datastore.
AnswerB

NIM exposes an HTTP/REST inference endpoint; the application is responsible for capturing the response and any metadata, then persisting it using its own database client. This is the standard integration pattern because NIM is stateless and does not manage application-level persistence, making this approach correct and flexible.

Why this answer

NIM microservices are designed to be stateless inference endpoints. The application that calls the NIM API is responsible for handling the response and any associated metadata. To persist this information, the application should use its own database client to write to the external datastore.

This separation of concerns ensures scalability and allows the developer to choose the appropriate datastore and schema.

Exam trap

The trap here is assuming that NIM includes built-in persistence or logging connectors, when actually persistence must be implemented in the calling application.

69
MCQmedium

When evaluating an LLM for a domain-specific task, why is 'Few-Shot Prompting' often superior to 'Zero-Shot Prompting'?

A.It decreases the number of tokens processed per request.
B.It provides in-context learning examples to guide the model.
C.It forces the model to ignore its internal pre-trained knowledge.
D.It is the only method to ensure the model produces non-hallucinated results.
AnswerB

By presenting examples (shots) within the prompt, the model uses them to understand the pattern or task structure. This in-context learning is highly effective for specialized tasks where the model needs to adapt to specific user requirements or output formats without requiring expensive fine-tuning.

Why this answer

Few-shot prompting provides the model with concrete examples of the desired input-output mapping, which acts as a guide for structure and reasoning. This reduces ambiguity and aligns the model's output with application-specific requirements. It is a fundamental technique for improving model precision in specialized domains where the model needs to understand unique formatting or logical patterns that are not explicitly defined in its base training data.

Exam trap

Candidates often believe few-shot prompting modifies the underlying model weights permanently, confusing it with parameter fine-tuning.

70
MCQmedium

A developer is building a customer support chatbot using NVIDIA NIM microservices. The application must reliably return structured JSON containing 'intent' and 'confidence' fields for downstream ticket routing. Which approach should the developer use to constrain the model's output format?

A.Set the 'temperature' parameter to 0 in the NIM API request.
B.Use NIM's guided decoding with a JSON schema supplied in the request.
C.Add a system prompt instructing the model to 'always respond in JSON'.
D.Increase 'max_tokens' so the model has room to complete the JSON.
AnswerB

NVIDIA NIM microservices support guided decoding that accepts a JSON schema, constraining token generation so the result validates against the schema. This guarantees the 'intent' and 'confidence' fields appear with correct types, eliminating brittle post-processing and retry loops in the ticket-routing pipeline.

Why this answer

Reliable structured output in NIM comes from guided decoding against a JSON schema, which constrains the decoder so every response validates. Temperature, token limits, and prompt wording influence style or length but cannot guarantee the required 'intent' and 'confidence' fields, making them unsuitable for automated ticket routing.

Exam trap

The trap here is assuming that a strict system prompt or temperature 0 guarantees valid JSON, when only schema-guided decoding actually constrains token generation.

71
MCQhard

Which memory management strategy in TensorRT-LLM is specifically designed to minimize fragmentation and allow for efficient KV cache allocation in multi-user environments?

A.Static contiguous memory allocation.
B.PagedAttention.
C.Global unified memory pooling.
D.Dynamic weight quantization.
AnswerB

PagedAttention manages the KV cache by dividing it into small blocks, allowing for non-contiguous storage. This design eliminates the internal fragmentation associated with static allocation, enabling the server to store more sequences concurrently and effectively increasing the overall capacity of the system for handling multiple user requests.

Why this answer

PagedAttention is the core innovation here. It manages the Key-Value (KV) cache in fixed-size blocks, similar to how an operating system manages virtual memory. By treating the KV cache as non-contiguous memory, TensorRT-LLM can efficiently allocate and reclaim space on the fly.

This prevents the memory fragmentation that usually occurs with static allocation, allowing for higher concurrency and supporting more simultaneous users without running out of GPU memory.

Exam trap

Candidates tend to confuse generic memory optimization techniques with PagedAttention, failing to recognize how virtual memory principles apply specifically to non-contiguous KV cache allocation.

72
MCQmedium

A developer is building a customer support chatbot using NVIDIA NIM microservices. They need the model to always respond in a formal tone and never mention competitor products. Where should these directives be placed in the API request to ensure consistent behavior across all user interactions?

A.In the 'assistant' role message pre-filled with a sample response
B.In the 'temperature' parameter set to a low value
C.In the 'system' role message at the beginning of the messages array
D.In the 'user' role message appended to each turn
AnswerC

The system message sets global instructions that the model must follow for all subsequent turns. Placing tone and content restrictions there ensures they apply uniformly, regardless of user input. This is the standard mechanism in NVIDIA NIM chat completions for persistent behavioral constraints, making it the correct choice.

Why this answer

The system role message provides persistent, high-priority instructions that shape the model's behavior throughout the conversation. By placing tone and content restrictions there, the developer ensures they are applied consistently to every response. Other roles or parameters do not carry the same authoritative weight for enforcing global rules.

Exam trap

The trap here is assuming that any message in the conversation can enforce global rules equally, when only the system role carries that persistent authority.

73
MCQmedium

A developer is building a customer-support assistant that must retrieve answers only from an approved internal knowledge base and cite the source document for each reply. They are using NVIDIA NIM microservices for the LLM and an embedding model, and they need the application layer to enforce citation behavior and reject answers that are not grounded in retrieved passages. Which software development approach best enforces this grounding requirement?

A.Post-process the model output with a separate NLI-style entailment check that compares each generated claim against the retrieved passages, and suppress any response whose claims are not entailed.
B.Fine-tune the NIM-hosted LLM on the internal knowledge base so that all answers are memorized in the model weights and retrieval becomes unnecessary.
C.Rely on the system prompt to instruct the model to answer only from context and to include citations, and ship the assistant once spot checks look acceptable.
D.Increase the model's temperature parameter so the assistant paraphrases the retrieved passages more freely and avoids repeating source wording verbatim.
AnswerA

Adding an entailment-verification step after generation directly enforces grounding: each claim is checked against the retrieved passages, and unsupported text is suppressed before it reaches the user. This works at the application layer with the NIM LLM and embedding endpoints and produces the citation guarantee the support assistant requires. It does not depend on the model voluntarily obeying instructions.

Why this answer

Grounding must be enforced programmatically in the application layer rather than trusted to the model's instruction-following. After the NIM LLM generates a draft answer from retrieved passages, an entailment or claim-verification step confirms that every statement is supported by those passages; unsupported claims are removed or the whole response is rejected. This yields auditable citations and prevents ungrounded content from reaching the user.

Exam trap

The trap here is assuming that a well-written system prompt is sufficient to guarantee grounded, cited answers in production.

74
MCQeasy

A developer is using the NVIDIA Triton Inference Server to deploy a TensorRT-LLM optimized model. They want to send a request with multiple prompts to be processed in a single inference call. Which Triton feature should they use?

A.Ensemble models
B.Request batching
C.Sequence batching
D.Dynamic batching
AnswerB

Triton Inference Server supports request batching, where a single client request can contain multiple input items (e.g., multiple prompts) that are processed together in one inference call. This is exactly what the developer needs to submit multiple prompts efficiently. The server will handle the batching internally if the model supports it.

Why this answer

Triton Inference Server allows clients to send a single request containing multiple input elements, which are then processed together as a batch. This is often referred to as request batching or client-side batching. For TensorRT-LLM models, the model's input tensor can have a batch dimension, so the client can populate it with multiple prompts.

This reduces the number of network round trips and can improve GPU utilization.

Exam trap

The trap here is confusing server-side dynamic batching with client-side request batching; only the latter lets you put multiple prompts into one request.

75
Multi-Selecthard

A developer is building a Retrieval-Augmented Generation (RAG) pipeline using NVIDIA NIM microservices. They need to ensure that the retriever returns the most relevant documents for a given query. Which two components should they optimize? (Choose two.)

Select 2 answers
A.The LLM's temperature setting
B.The embedding model used to vectorize documents and queries
C.The vector database index and similarity search algorithm
D.The tokenizer's vocabulary size
E.The number of attention heads in the LLM
AnswersB, C

The embedding model determines the quality of semantic representations. A model fine-tuned for the domain or a state-of-the-art model like NVIDIA's NV-Embed can significantly improve retrieval relevance. Choosing the right embedding model is foundational to RAG performance, making this a correct choice.

Why this answer

Retrieval quality in RAG depends heavily on the embedding model that encodes semantic meaning and the vector database that efficiently finds nearest neighbors. Optimizing both ensures that the most relevant documents are surfaced. Other parameters like temperature or attention heads affect generation, not retrieval, and tokenizer vocabulary is fixed.

Exam trap

The trap here is conflating generation parameters with retrieval components, leading to choices that do not affect which documents are retrieved.

Page 1 of 2 · 77 questions totalNext →

Ready to test yourself?

Try a timed practice session using only Nca Software Development questions.