Courseiva

CCNA Model Deployment Questions

37 questions · Model Deployment · All types, answers revealed

1
MCQeasy

An engineer is preparing an ensemble in NVIDIA Triton Inference Server that chains a Python preprocessing model with a TensorRT-LLM backend model. The preprocessing model must run on CPU while the LLM must run on GPU, and the ensemble must expose a single HTTP endpoint. Which configuration is required to make the ensemble execute correctly?

A.Write a custom backend that loads both models and orchestrates them in C++
B.Merge both models into a single TensorRT-LLM engine using a custom plugin
C.Use Triton's 'instance_group' to pin both models to the same GPU and rely on default routing
D.Define each step as a separate Triton model and reference them in a platform 'ensemble' model's ensemble_scheduling block
AnswerD

Triton ensembles are defined by a model with platform 'ensemble', which contains an ensemble_scheduling section listing each step, its model name, version, and input/output mappings. This lets Triton route tensors between the CPU Python model and the GPU TensorRT-LLM model while exposing one endpoint.

Why this answer

Triton ensembles use a model with platform 'ensemble' and an ensemble_scheduling block that lists each step with model name, version, and tensor mappings. This enables chaining a CPU Python preprocessing model with a GPU TensorRT-LLM model behind a single endpoint, which is exactly the requirement.

Exam trap

The trap here is reaching for custom code or instance_group tuning when Triton's built-in ensemble scheduling already solves multi-model chaining.

2
MCQeasy

A startup is deploying a small LLM for a chatbot on a single NVIDIA L4 GPU using NVIDIA Triton Inference Server. They want to ensure the model is automatically loaded when Triton starts and can be updated without restarting the server. Which Triton feature should they configure?

A.Instance groups with multiple model instances per GPU.
B.Model repository with version policy and model control API.
C.Dynamic batching with preferred batch sizes.
D.Ensemble scheduler with a custom preprocessing model.
AnswerB

Triton's model repository stores model files and versions. By configuring the version policy (e.g., latest), Triton can automatically load the latest version at startup. The model control API allows loading and unloading models at runtime without restarting the server. This meets both requirements: automatic loading and dynamic updates.

Why this answer

The model repository with version policy allows Triton to automatically load models at startup, and the model control API enables loading and unloading models at runtime. This combination provides both automatic initialization and dynamic updates without server restarts. Other options address inference scheduling or parallelism, not model lifecycle management.

Exam trap

The trap here is confusing the ensemble scheduler with model management, when ensembles are for inference pipelines, not loading or versioning.

3
Multi-Selectmedium

A team is deploying a large language model on NVIDIA Triton Inference Server with TensorRT-LLM backend. They want to reduce GPU memory consumption to fit a larger model on a single GPU without significantly degrading output quality. Which two techniques should they use? (Choose two.)

Select 2 answers
A.Increase the maximum batch size
B.Use FP32 precision for all computations
C.Use tensor parallelism across multiple GPUs
D.Enable INT8 weight-only quantization
E.Enable paged KV cache
AnswersD, E

INT8 weight-only quantization reduces the precision of model weights from FP16 to INT8, cutting memory usage by roughly half. It often preserves output quality well because activations remain in higher precision. This allows larger models to fit on a single GPU with minimal accuracy loss, making it a suitable technique for the scenario.

Why this answer

INT8 weight-only quantization reduces weight memory footprint with minimal quality loss, and paged KV cache optimizes memory usage during generation by reducing fragmentation. Together, they enable larger models to fit on a single GPU. FP32 increases memory, larger batch sizes demand more memory, and tensor parallelism requires multiple GPUs, so they do not meet the goal.

Exam trap

The trap here is assuming that increasing batch size or using FP32 improves memory efficiency, when they actually increase memory usage.

4
MCQeasy

A healthcare startup is deploying a Mistral 7B model for internal clinical note summarization. They need to serve the model with NVIDIA Triton Inference Server and want to minimize GPU memory footprint during inference. The team plans to use TensorRT-LLM and is choosing a numerical precision for the engine. Which precision should they select to reduce memory usage while maintaining acceptable accuracy for summarization?

A.FP16, because it halves memory compared to FP32 and is widely supported on NVIDIA data center GPUs.
B.FP32, because it uses the least GPU memory and is the default for TensorRT-LLM engines.
C.TF32, because it is a Tensor Core mode that automatically compresses weights to one byte per parameter.
D.INT8 or FP8 quantization, because it reduces weight and activation precision further than FP16 while maintaining acceptable accuracy for text summarization.
AnswerD

INT8 and FP8 quantization reduce the number of bits per weight and activation compared to FP16, lowering GPU memory usage and often improving throughput. For summarization, a tolerant task, the accuracy loss is typically acceptable when calibration is done properly. This directly addresses the requirement to minimize memory footprint while keeping output quality suitable for clinical note summarization.

Why this answer

INT8 or FP8 quantization reduces the bits per weight and activation below FP16, cutting GPU memory usage while keeping accuracy acceptable for summarization. TensorRT-LLM supports quantized engines with calibration or scaling factors, making it the appropriate precision choice when memory footprint is the primary constraint and the task tolerates small accuracy loss.

Exam trap

The trap here is equating TF32 with a memory-compression format, when it is actually a Tensor Core compute mode that does not shrink stored weights.

5
MCQhard

A team is deploying a 70B-parameter LLM across four NVIDIA H100 GPUs using NVIDIA TensorRT-LLM with tensor parallelism. They observe that inference works but throughput is lower than expected, and profiling shows significant inter-GPU communication overhead. Which optimization should they apply first to reduce communication overhead?

A.Enable NVLink and ensure the GPUs are connected via NVSwitch for peer-to-peer communication.
B.Reduce the tensor parallel size to 2 and run two independent replicas.
C.Increase the number of attention heads to improve parallelism.
D.Switch from tensor parallelism to pipeline parallelism to eliminate inter-GPU communication.
AnswerA

Tensor parallelism requires frequent all-reduce operations between GPUs. NVLink with NVSwitch provides high-bandwidth, low-latency peer-to-peer communication, which is essential to reduce the overhead. Without NVLink, communication over PCIe becomes a bottleneck. Ensuring NVLink is enabled and the topology uses NVSwitch is the first and most impactful optimization for multi-GPU tensor parallelism.

Why this answer

Tensor parallelism relies on frequent all-reduce operations, so the interconnect bandwidth is critical. NVLink with NVSwitch provides the necessary high-speed peer-to-peer communication to minimize overhead. Without it, PCIe becomes the bottleneck.

Other options either do not address the communication pattern or are infeasible for a 70B model on four GPUs.

Exam trap

The trap here is thinking that pipeline parallelism eliminates inter-GPU communication, when it actually introduces different communication and pipeline bubbles.

6
MCQmedium

An enterprise is deploying a large language model on NVIDIA Triton Inference Server. Which deployment strategy minimizes latency for requests that require high-throughput batching while maintaining consistent hardware utilization?

A.Implement static batching with a fixed size of 1.
B.Enable Dynamic Batching in the Triton model configuration file.
C.Disable all batching features to process requests serially.
D.Offload all batching logic to the client-side application layer.
AnswerB

Dynamic Batching aggregates individual requests into batches based on defined delay windows, maximizing GPU compute cycles. By adjusting batching parameters, administrators can balance throughput and latency effectively. This is the industry-standard method for optimizing NVIDIA hardware utilization when serving LLMs in real-world, high-concurrency production environments.

Why this answer

Dynamic Batching is the optimal strategy for Triton Inference Server in production environments. It groups individual inference requests arriving within a short time window into a single batch, allowing the GPU to process them in parallel. This maximizes throughput by fully saturating CUDA cores, reducing the overhead of kernel launches, and ensuring that hardware utilization remains high even under variable traffic loads, effectively balancing latency and overall system capacity.

Exam trap

Candidates often confuse dynamic batching with model parallelism or caching, failing to recognize that Triton's dynamic batching is specifically designed to maximize GPU utilization by grouping requests at runtime.

7
MCQmedium

An enterprise deployment team needs to deploy a Large Language Model on NVIDIA Triton Inference Server. They require the lowest possible latency for real-time inference while maximizing GPU memory utilization. Which configuration strategy should the team implement?

A.Enable dynamic batching with a maximum delay of 500ms to gather requests.
B.Deploy the model as a standard PyTorch model using the default model repository.
C.Implement TensorRT-LLM with In-flight Batching and CUDA Graph capture.
D.Use the Triton Model Analyzer to force a single-request batch size of 1.
AnswerC

This configuration maximizes GPU efficiency by decoupling request scheduling from token generation. In-flight batching allows for concurrent processing of tokens, while CUDA graphs minimize driver launch overhead. This combination is the industry-standard path for achieving high-performance LLM serving on NVIDIA hardware.

Why this answer

To achieve optimal performance on Triton, using TensorRT-LLM backends with In-flight Batching (IFB) is essential. IFB allows the server to process new requests before existing ones finish, significantly increasing throughput and reducing latency. Configuring CUDA Graph capture further reduces CPU overhead by recording GPU operations.

This approach is critical for LLMs because traditional static batching often leads to underutilized GPU cycles during the autoregressive decoding phase.

Exam trap

Candidates often confuse static batching with In-flight Batching, incorrectly assuming that standard model ensemble configurations provide the same latency benefits as native IFB support offered by the TensorRT-LLM backend.

8
MCQhard

A team is deploying a 70B-parameter LLM using NVIDIA Triton Inference Server with TensorRT-LLM backend on a node with four A100 80GB GPUs. They observe that during inference, only one GPU is utilized while the others remain idle. They have configured the model with tensor parallelism set to 1. What is the most likely cause of this underutilization?

A.The Triton model repository is not configured with the correct instance group for multi-GPU execution.
B.Tensor parallelism is set to 1, which means the model is not partitioned across GPUs; it runs entirely on one GPU.
C.The model was compiled with a batch size that is too small to trigger multi-GPU execution.
D.The KV cache is not enabled, causing the model to fall back to single-GPU execution.
AnswerB

Tensor parallelism splits model layers across multiple GPUs. When set to 1, no partitioning occurs, so the entire model resides on a single GPU, leaving others idle. To utilize all four GPUs for a single model instance, tensor parallelism must be set to 4, matching the number of GPUs. This is the direct cause of underutilization.

Why this answer

Tensor parallelism determines how a model is sharded across multiple GPUs. Setting it to 1 means no sharding, so the entire model runs on a single GPU, leaving others idle. To utilize all four GPUs for a single model instance, tensor parallelism should be set to 4.

Other factors like instance groups or batch size do not override this fundamental partitioning setting.

Exam trap

The trap here is confusing instance groups (which manage multiple model instances) with tensor parallelism (which splits a single model across GPUs).

9
MCQhard

Refer to the exhibit. What is the implication of setting the memory_limit to 0.8 in the context of an LLM inference service?

A.It guarantees that the system will never crash due to memory.
B.It will evict old KV cache entries when 80% limit is reached.
C.It forces the GPU to run at 80% of its clock speed.
D.It expands the memory capacity by 20% using swap space.
AnswerB

The 0.8 setting acts as a cap on the memory footprint of the KV cache. When usage reaches 80% of the assigned memory, the LRU policy triggers the eviction of the least recently used entries, allowing the system to continue operation without crashing, albeit at the cost of losing older sequence data.

Why this answer

A memory limit of 0.8 indicates that the system will reserve up to 80% of the allocated memory for the KV cache. Once this limit is reached, the Least Recently Used (LRU) policy will begin evicting older sequences to make room for new ones. This helps prevent hard OOM crashes, but users of the evicted sequences will experience errors or forced re-computations when trying to continue their generation tasks.

Exam trap

Candidates often mistakenly believe the 0.8 memory limit triggers a system-wide shutdown or error, failing to recognize it as a threshold for the LRU eviction policy used in KV cache management.

10
MCQmedium

What is the primary benefit of deploying a model with a 'Model Ensemble' configuration in Triton Inference Server?

A.It increases the number of GPUs available to the model.
B.It enables multi-model pipelining with reduced network overhead.
C.It automatically converts models into TensorRT engines.
D.It provides a load-balancing mechanism across server clusters.
AnswerB

Ensembles allow multiple inference tasks, such as text pre-processing and LLM generation, to occur sequentially on the server. By chaining these models together, the client only needs to send one request, avoiding the latency and security risks associated with multiple network round-trips for each part of the pipeline.

Why this answer

Model ensembles allow developers to chain multiple models together into a single pipeline (e.g., preprocessing, inference, and post-processing). This keeps the logic within the server, reducing the overhead of multiple network round-trips between the client and the server. By managing the entire pipeline as a single unit, you ensure consistent data flow and lower overall latency for the end-to-end application.

Exam trap

Candidates frequently assume ensembles are primarily for predictive accuracy or model voting, missing the architectural benefit of reducing network latency through server-side pipelining of preprocessing and post-processing steps.

11
MCQhard

An engineer is deploying a LLM using NVIDIA Triton Inference Server with the TensorRT-LLM backend. They need to ensure that the model can handle a sudden surge in requests without increasing latency beyond a specified threshold. They have configured the model with a maximum batch size of 32 and dynamic batching with a preferred batch size of 16. However, during peak load, latency spikes are observed. Which Triton configuration parameter should they adjust to control the maximum time a request waits in the dynamic batching queue before being processed?

A.max_batch_size
B.instance_group
C.preferred_batch_size
D.max_queue_delay_microseconds
AnswerD

max_queue_delay_microseconds sets the maximum time a request can wait in the dynamic batching queue before Triton processes it, even if the preferred batch size is not reached. Reducing this value limits latency spikes during peak load by forcing earlier execution, though it may reduce batching efficiency. It directly controls the trade-off between latency and throughput.

Why this answer

The max_queue_delay_microseconds parameter in Triton's dynamic batching configuration specifies the maximum time a request can wait in the queue before being processed. Lowering it reduces latency spikes under surge conditions by preventing requests from waiting too long for a full batch. Other parameters like max_batch_size and preferred_batch_size affect batching but not the wait timeout.

Exam trap

The trap here is assuming that increasing batch size or instance count will solve latency spikes, when the root cause is the queue wait time controlled by max_queue_delay_microseconds.

12
MCQeasy

Which metric is most critical to monitor for identifying 'bottlenecks' in a high-throughput LLM deployment?

A.Total number of files in the model repository.
B.GPU Utilization percentage.
C.The name of the backend framework used for training.
D.The number of times the server was restarted.
AnswerB

GPU utilization is a key indicator of hardware efficiency. Low utilization during high-traffic periods indicates that the model is starving for data or that the batching strategy is inefficient. High utilization confirms that the compute resources are being used effectively to process tokens within the inference engine.

Why this answer

GPU utilization is the primary metric for understanding how well the model is saturating the compute resources. If GPU utilization is low while latency is high, it suggests a bottleneck elsewhere, such as CPU preprocessing, data transfer, or synchronization issues. Monitoring this metric allows engineers to determine if they need to increase batch sizes, optimize the pipeline, or scale the infrastructure to maintain performance.

Exam trap

Candidates often select 'latency' or 'throughput' as the primary bottleneck metric, failing to realize that GPU utilization is the fundamental indicator of whether the underlying hardware is actually being saturated.

13
MCQhard

A media-analytics firm serves a 13B-parameter summarization model on two A100 GPUs using NVIDIA TensorRT-LLM behind Triton Inference Server. Traffic is bursty: during live events concurrency triples for about ten minutes, then returns to baseline. Operators report that the first requests after each burst begin are slow and sometimes time out, although steady-state latency is acceptable. Which deployment change most directly addresses the cold-start penalty at the beginning of each burst?

A.Switch the Triton scheduler from dynamic batching to sequence batching so that in-flight conversations are tracked per client.
B.Configure Triton's instance groups to keep a resident model instance and enable the model's warmup configuration so activation buffers and CUDA graphs are exercised before live traffic arrives.
C.Enable TensorRT-LLM's in-flight batching and raise the KV cache fraction so more concurrent sequences can share the cache.
D.Increase the TensorRT-LLM engine's max_batch_size so larger batches can be formed during the burst peaks.
AnswerB

Resident instances plus a warmup configuration cause Triton to load the engine and run representative inference requests at startup, allocating workspace, compiling or replaying CUDA graphs, and paging in weights. When the burst begins, those resources are already hot, so the first real requests no longer pay the initialization penalty. This directly targets the cold-start behavior described.

Why this answer

The reported pattern, acceptable steady-state latency but slow and timing-out requests at the start of each burst, is a classic cold-start symptom. Keeping a resident Triton instance and supplying a warmup configuration forces engine loading, workspace allocation, and CUDA graph capture to happen at server startup rather than on the first live request. This removes the initialization stall precisely when traffic spikes.

Exam trap

The trap here is diagnosing burst-boundary latency as a batching or KV-cache capacity problem when it is actually resource initialization that occurs on first use.

14
MCQhard

Refer to the exhibit. The Triton Inference Server is failing to load a model. The configuration shows two instances assigned to GPU 0. What is the most likely cause and the correct remediation?

A.The model count exceeds the number of available CUDA cores, requiring a driver update.
B.The total VRAM required for two instances exceeds the GPU capacity; reduce instance count to 1.
C.The model weights are corrupted, preventing the inference engine from initializing the memory space.
D.The batch size is set to zero in the configuration, preventing memory allocation.
AnswerB

Each model instance requires a dedicated memory buffer for weights and activation tensors. When configured with 'count: 2', Triton attempts to load the model twice. If the sum exceeds the VRAM, an OOM occurs. Reducing the instance count is the most direct way to resolve the startup conflict.

Why this answer

The error indicates that the two instances of the model are collectively requesting more VRAM than is available on the physical GPU. By default, Triton attempts to allocate memory for every configured instance upon startup. The remediation requires either reducing the number of instances or implementing a memory-aware model partitioning strategy, such as using Model Analyzer to determine the safe memory footprint per instance before deployment.

Exam trap

Candidates frequently assume the error is due to a software version mismatch or driver issue, failing to calculate the cumulative VRAM consumption of multiple model instances relative to the hardware limit.

15
Multi-Selecthard

A team is deploying a 70B parameter LLM with NVIDIA Triton Inference Server across four NVIDIA H100 GPUs. They are using TensorRT-LLM and need to fit the model within the combined GPU memory while maintaining high throughput. Which two techniques should they use? (Choose two.)

Select 2 answers
A.Pipeline parallelism, which assigns entire transformer layers to different GPUs and passes activations between stages.
B.Increasing the maximum batch size to the largest value the backend accepts, so memory is fully utilized.
C.Data parallelism, which replicates the full model on every GPU and splits incoming requests across replicas.
D.Enabling FP32 precision for all weights and activations to maximize numerical stability across GPUs.
E.Tensor parallelism, which shards model layers and attention heads across the GPUs so each GPU holds a fraction of the weights.
AnswersA, E

Pipeline parallelism assigns groups of layers to different GPUs, so each GPU stores only a portion of the model. Combined with tensor parallelism, it allows very large models such as 70B parameters to fit across four H100 GPUs. It is commonly used together with tensor parallelism to balance memory and communication overhead in multi-GPU TensorRT-LLM deployments.

Why this answer

Tensor parallelism and pipeline parallelism are the two model-parallel techniques that split a 70B model across multiple GPUs, allowing the weights to fit in aggregate memory on four H100s. Tensor parallelism shards layers and attention heads, while pipeline parallelism assigns layer groups to stages; together they enable large-model deployment with high throughput.

Exam trap

The trap here is confusing data parallelism, which replicates the full model per GPU, with model parallelism, which actually splits weights to fit a large model across devices.

16
MCQmedium

A healthcare company is deploying an LLM for clinical note summarization using NVIDIA Triton Inference Server. They must ensure that only authorized users can access the model and that all inference requests are logged for audit. Which Triton feature should they configure to enforce authentication and authorization?

A.Triton's built-in model access control list (ACL) configured via the model repository.
B.Triton's HTTP/REST and gRPC endpoints with a reverse proxy that performs OAuth 2.0 token validation.
C.Triton's dynamic batching configuration with priority levels.
D.Triton's ensemble scheduler with a custom authentication model.
AnswerB

Triton itself does not provide built-in authentication or authorization. The recommended approach is to place a reverse proxy (such as NGINX or Envoy) in front of Triton to handle OAuth 2.0 token validation and access control. This satisfies the requirement for authorized access while Triton focuses on inference. Logging can be handled at the proxy or application level for audit.

Why this answer

Triton Inference Server does not include native authentication or authorization. The standard pattern is to deploy a reverse proxy that validates OAuth 2.0 tokens and forwards authorized requests to Triton. This separates security concerns from inference and allows audit logging at the proxy.

Other options describe scheduling or non-existent features that do not enforce access control.

Exam trap

The trap here is assuming Triton has built-in user authentication, when it actually relies on external components like reverse proxies for security.

17
MCQeasy

An engineer needs to expose an LLM served by NVIDIA Triton Inference Server to a web application over HTTP with token streaming. Which Triton feature should they enable?

A.Model ensembles combining preprocessing and postprocessing models.
B.Dynamic batching in the model configuration.
C.Instance groups with multiple GPU instances per model.
D.Decoupled mode with streaming responses in the model backend.
AnswerD

Decoupled mode allows a model backend to return multiple responses for a single request, which is essential for token streaming in LLMs. Triton's HTTP and gRPC endpoints support streaming when the model is configured for decoupled transactions. This lets the web application receive partial outputs as tokens are generated, improving perceived latency.

Why this answer

Triton's decoupled mode lets a backend emit multiple responses for one request, which is the mechanism used for LLM token streaming. Configuring the model for decoupled transactions and using a streaming-capable client over HTTP or gRPC delivers tokens incrementally. Other Triton features like batching, ensembles, or instance groups improve throughput or composition but do not stream partial outputs.

Exam trap

The trap here is confusing throughput optimizations such as dynamic batching with the response-streaming capability required for token-by-token delivery.

18
MCQhard

An inference team runs a TensorRT-LLM model behind NVIDIA Triton Inference Server with the tensorrtllm_backend. They notice that when clients send requests with widely varying prompt lengths, throughput drops and some requests wait much longer than others. They want Triton to group compatible requests dynamically to improve GPU utilization without changing the model build. Which Triton feature should they configure?

A.Model warmup with sample inputs to pre-allocate memory at load time
B.Dynamic batching with a preferred batch size and max queue delay in the model's config.pbtxt
C.Rate limiting with a max queue size to reject excess requests
D.Sequence batching with a sequence ID to maintain state across requests
AnswerB

Dynamic batching lets Triton combine independent requests arriving within a time window into a single batch, improving GPU utilization when request sizes vary. Setting preferred_batch_size and max_queue_delay controls the tradeoff between latency and throughput, and it requires no model rebuild.

Why this answer

Triton's dynamic batching groups independent requests arriving within max_queue_delay into a single batch, using preferred_batch_size to target efficient batch shapes. This improves GPU utilization under varying prompt lengths without rebuilding the TensorRT-LLM engine, directly addressing the throughput and latency variance described.

Exam trap

The trap here is confusing sequence batching, which maintains state for ordered requests, with dynamic batching, which merges independent concurrent requests.

19
MCQmedium

When deploying an LLM using NVIDIA TensorRT-LLM, what is the primary benefit of pre-compiling the model into a TensorRT engine?

A.It allows the model to run on any generic CPU architecture.
B.It enables layer fusion and kernel selection for the target GPU.
C.It eliminates the need for any GPU memory during inference.
D.It automatically scales the model across multiple distributed nodes.
AnswerB

TensorRT-LLM optimizes the computational graph by fusing layers and selecting the most efficient kernels for the specific GPU architecture. This significantly reduces memory bandwidth consumption and increases computational throughput, which is essential for the high-performance requirements of modern generative AI models in production environments.

Why this answer

Pre-compiling with TensorRT-LLM allows for layer fusion, kernel auto-tuning, and memory optimization tailored specifically to the target GPU architecture. By performing these heavy optimizations offline, the inference engine can execute at peak performance immediately upon loading. This eliminates the runtime overhead associated with graph compilation or dynamic graph execution, ensuring minimal latency and optimal resource utilization from the very first inference request.

Exam trap

Candidates often think TensorRT-LLM just compresses the model, missing the critical benefit of offline graph optimization and kernel fusion that significantly reduces runtime latency during inference.

20
MCQhard

Refer to the exhibit. An engineer observes that memory utilization spikes during inference, causing OOM errors. Given the configuration, what is the most likely cause of the failure?

A.The count of 2 causes the GPU to oversubscribe its thermal limits.
B.The instance group count of 2 causes redundant loading of weights, exceeding VRAM.
C.Triton requires KIND_CPU for concurrent instance execution.
D.The gpus index [0] is invalid for multi-instance deployment.
AnswerB

Setting the instance count to two instructs Triton to create two independent model runners. Each runner requires its own memory allocation for weights and activations. If the model occupies a large portion of the GPU memory, running two instances simultaneously will inevitably exhaust the total available VRAM.

Why this answer

The instance group configuration defines two concurrent instances on the same GPU. Each instance attempts to load a separate copy of the model weights into the GPU memory. If the model size is large, doubling the instances exceeds the available VRAM capacity.

This configuration is a common mistake when deploying LLMs where model footprint is significant relative to total available device memory.

Exam trap

Candidates often blame the model size or the GPU hardware itself, failing to notice that the configuration defines multiple instances, which causes a multiplicative effect on VRAM consumption.

21
MCQmedium

An organization is deploying a high-throughput LLM on NVIDIA Triton Inference Server. They observe significant tail latency spikes when serving multiple concurrent requests. Which strategy most effectively optimizes GPU utilization and reduces latency jitter for these concurrent model instances?

A.Increase the number of CPU threads per model instance in the configuration file.
B.Disable all model instances to ensure requests are processed in serial order.
C.Enable dynamic batching with an optimized max_queue_delay_microseconds setting.
D.Switch the model to a lower precision format like FP8 without model recalibration.
AnswerC

Dynamic batching allows Triton to accumulate requests over a specified window to create larger, more efficient tensor operations. Fine-tuning the queue delay ensures that the server waits just long enough to fill a batch without unnecessarily delaying individual requests, directly mitigating tail latency spikes during peak traffic.

Why this answer

Dynamic Batching is the primary mechanism in Triton to combine individual inference requests into a single batch, significantly improving throughput while minimizing latency. By configuring the 'max_queue_delay_microseconds' parameter, the system balances wait times with compute efficiency. This is critical for enterprise deployments where maximizing GPU hardware investment while maintaining strict service-level agreements is the standard requirement for production-grade generative AI applications.

Exam trap

Candidates often recommend simply scaling up GPU count or adjusting thread counts, overlooking Triton's native queue management mechanisms designed explicitly to control latency jitter.

22
MCQmedium

A team is deploying a 13B-parameter LLM with NVIDIA TensorRT-LLM on a single A100 80GB GPU. They want to reduce GPU memory usage during inference without retraining the model, while keeping acceptable output quality. Which technique should they apply?

A.Enable INT8 weight-only quantization using TensorRT-LLM's quantization toolkit.
B.Reduce the maximum sequence length to 512 tokens to lower activation memory.
C.Increase the KV cache block size to 128 tokens to reduce memory fragmentation.
D.Enable Tensor Parallelism across two GPUs to split the model.
AnswerA

INT8 weight-only quantization compresses model weights to 8-bit integers while keeping activations in higher precision, cutting weight memory roughly in half with minimal quality loss. TensorRT-LLM supports this via its quantization toolkit and calibration workflow, making it a practical post-training approach for a 13B model on a single 80GB GPU without retraining.

Why this answer

INT8 weight-only quantization is a post-training technique supported by TensorRT-LLM that reduces weight memory by roughly half while preserving output quality. It directly addresses the goal of lowering GPU memory usage for a large model on a single GPU without retraining. Other listed changes affect scheduling or activation memory rather than the core weight footprint.

Exam trap

The trap here is assuming that KV cache tuning or sequence-length reduction will solve weight-dominated memory pressure for a large model.

23
MCQeasy

An AI engineer is deploying a large language model using NVIDIA Triton Inference Server. They need to ensure that the server can handle multiple concurrent requests efficiently while maintaining low latency. Which Triton feature allows the server to dynamically batch incoming requests to maximize GPU utilization?

A.Dynamic batching
B.Instance groups
C.Model versioning
D.Ensemble scheduling
AnswerA

Dynamic batching in Triton automatically groups individual inference requests into batches on the server side, within a configurable time window. This increases throughput and GPU utilization while keeping latency low for high-concurrency scenarios. It is specifically designed to handle multiple concurrent requests efficiently without client-side batching.

Why this answer

Dynamic batching is a Triton feature that aggregates concurrent inference requests into larger batches on the server side, improving GPU utilization and throughput while managing latency. It is the standard mechanism for handling high concurrency efficiently. Model versioning, instance groups, and ensemble scheduling serve different purposes and do not provide dynamic request batching.

Exam trap

The trap here is confusing dynamic batching with instance groups, which also affect concurrency but through multiple model instances rather than request aggregation.

24
Multi-Selectmedium

Which TWO factors should be prioritized when selecting a quantization strategy for deploying a large language model on constrained edge hardware? (Select TWO)

Select 2 answers
A.Maximizing the model training loss to improve generalization.
B.Reducing the memory footprint of the model weights.
C.Increasing the number of neural network layers in the architecture.
D.Improving inference throughput via reduced bit-precision arithmetic.
E.Replacing the Transformer architecture with a linear regression model.
AnswersB, D

On edge devices, VRAM is severely limited. Quantization reduces the bit-depth of weights, directly lowering the memory requirement. This allows larger models to fit into the limited VRAM of edge hardware, which is critical for enabling complex LLM inference tasks that would otherwise fail to load.

Why this answer

Selecting a quantization strategy requires balancing precision loss against hardware performance gains. When deploying to edge devices, memory bandwidth and storage capacity are typically the primary bottlenecks. By reducing precision from FP16 or FP32 to INT8 or FP8, you directly decrease the model's memory footprint and increase the number of operations per clock cycle, which is essential for maintaining acceptable real-time inference speeds on resource-limited hardware.

Exam trap

Candidates often confuse hardware training constraints with edge deployment limitations, incorrectly prioritizing compute scaling factors instead of focusing strictly on memory footprint and memory bandwidth bottlenecks inherent to edge devices.

25
Multi-Selecthard

A platform team is deploying a 70B-parameter LLM with NVIDIA TensorRT-LLM across four 80 GB H100 GPUs and needs to serve long-context requests efficiently. They are deciding how to combine parallelism and memory techniques in the build and runtime configuration. (Choose two.)

Select 2 answers
A.Use tensor parallelism across the four GPUs so each GPU holds a shard of every layer's weights
B.Configure the paged KV cache with a block size and max sequence length sized for the target context window
C.Reduce the number of GPUs to one and enable FP8 quantization to fit the model on a single device
D.Rely on CUDA Unified Memory to transparently page weights and KV cache between host and device
E.Enable pipeline parallelism with a single micro-batch to minimize inter-GPU communication
AnswersA, B

Tensor parallelism splits each layer's weight matrices across GPUs, allowing a 70B model that would not fit on one 80 GB H100 to be served across four. With NVLink-connected H100s, the all-reduce communication overhead is manageable, and it is the standard way to serve very large models with low latency.

Why this answer

Tensor parallelism shards each layer across the four H100s so a 70B model fits and serves with low latency over NVLink, while a properly sized paged KV cache prevents memory over-reservation for long contexts. Together they address both weight distribution and the dominant dynamic memory consumer in long-context LLM serving.

Exam trap

The trap here is treating pipeline parallelism or Unified Memory as easy wins, when pipeline bubbles and host-device paging both hurt latency-sensitive long-context serving.

26
MCQeasy

Which component in the NVIDIA AI Enterprise stack is primarily responsible for serving multiple models, managing model versions, and providing metrics for monitoring model health?

A.CUDA Toolkit
B.NVIDIA Triton Inference Server
C.NVIDIA NeMo
D.NVIDIA DALI
AnswerB

Triton is the purpose-built inference server that manages model lifecycles, supports various frameworks, and exposes comprehensive metrics via endpoints like Prometheus. It is designed to optimize serving across different hardware configurations and ensures that models are served efficiently and reliably within large-scale enterprise production environments.

Why this answer

NVIDIA Triton Inference Server is the standard tool for model serving in production. It provides a unified API for various frameworks, supports concurrent model execution, and performs health checks. It is designed to handle model versioning and provide detailed telemetry data, which is essential for maintaining reliable and scalable AI deployments in enterprise production pipelines.

Exam trap

Candidates often confuse the model serving layer with the training framework or the orchestration layer, failing to identify Triton as the specific tool for serving and monitoring models.

27
MCQmedium

In the context of LLM deployment, why is it recommended to use a dedicated inference server like Triton rather than a basic Flask or FastAPI wrapper?

A.Flask is incompatible with the NVIDIA CUDA drivers.
B.Triton manages concurrent model execution and hardware utilization.
C.Custom wrappers require more memory than Triton servers.
D.Triton is the only way to support REST API endpoints.
AnswerB

Triton is optimized for GPU utilization, providing features such as dynamic batching, concurrent model execution, and shared memory management. These optimizations are difficult to implement correctly in standard web frameworks and are essential for maximizing the performance of expensive GPU resources in large-scale LLM deployments.

Why this answer

While basic web frameworks are easy to implement, they lack the sophisticated features required for production AI. Triton offers hardware-optimized concurrency, automatic dynamic batching, support for multiple backends (TensorRT, PyTorch, ONNX), and advanced metrics. These features are critical for handling the high computational demands and low-latency requirements of LLMs, providing better scalability and reliability compared to custom-built wrappers that do not natively manage GPU streams and queues.

Exam trap

Candidates assume that because Flask/FastAPI can serve models, they are sufficient for production. They overlook the critical need for dynamic batching and hardware-optimized concurrency provided by dedicated inference servers.

28
MCQmedium

When deploying a model, what is the benefit of using Triton's 'Model Versioning' feature?

A.It automatically cleans up old model files from the drive.
B.It allows seamless model updates without service restarts.
C.It provides built-in encryption for sensitive model weights.
D.It compiles models into different formats like ONNX and TorchScript.
AnswerB

Versioning enables the server to detect and load new versions of a model dynamically. This allows updates to be pushed to production without interrupting current inference requests, providing the high availability and zero-downtime deployment capabilities required by enterprise-grade AI production environments.

Why this answer

Model versioning allows multiple versions of a model to exist in the repository simultaneously. This enables A/B testing, gradual rollouts, and instant rollbacks. By simply changing the configuration or updating a symbolic link, administrators can shift traffic to a new model version without downtime, ensuring that the service remains available while testing new model iterations or applying hotfixes to production deployments.

Exam trap

Candidates often assume versioning is only for tracking experiments or storage management, failing to realize its critical role in enabling zero-downtime deployments through seamless traffic routing and instant rollbacks.

29
MCQmedium

A team is deploying a quantized LLM using NVIDIA NIM. To ensure the highest level of security and compliance, they need to verify that the container image has been scanned for vulnerabilities before production use. Which tool is the primary source for certified, production-ready NIM containers?

A.The public NVIDIA Docker Hub repository.
B.The official NVIDIA NGC catalog.
C.A custom build using the standard PyTorch base image.
D.The GitHub repository containing the source code for the model.
AnswerB

NVIDIA NGC is the authoritative source for enterprise-ready containers. Images in this catalog are built, tested, and scanned by NVIDIA, ensuring compatibility with NVIDIA GPUs and adherence to security standards required for deploying LLMs in production environments with strictly defined compliance needs.

Why this answer

NVIDIA NGC (NVIDIA GPU Cloud) is the central repository for certified, secure, and optimized container images. Using NIM containers from the NVIDIA NGC catalog ensures that the images have been scanned by NVIDIA security pipelines. This is critical for enterprise compliance, as it guarantees the software components, including CUDA libraries and inference runtimes, are patched and verified for production deployment.

Exam trap

Candidates often choose public registries like Docker Hub. While accessible, they lack the specific NVIDIA certification and security vetting provided by NGC, which is a requirement for enterprise-grade LLM deployments.

30
Multi-Selecthard

A team is deploying a large language model on NVIDIA Triton Inference Server with NVIDIA TensorRT-LLM backend. They need to reduce GPU memory usage to fit a larger model on the same hardware while maintaining acceptable latency. Which two techniques should they use? (Choose two.)

Select 2 answers
A.Set the Triton dynamic batching max_queue_delay_microseconds to 0 to process requests immediately.
B.Quantize the model weights to INT8 or FP8 using TensorRT-LLM's quantization toolkit.
C.Increase the tensor parallel size to 8 across eight GPUs.
D.Enable paged KV cache and in-flight batching in the TensorRT-LLM backend.
E.Use FP32 precision for all layers to avoid quantization errors.
AnswersB, D

Quantizing weights to INT8 or FP8 reduces the memory footprint by 2x to 4x compared to FP16, allowing larger models to fit. TensorRT-LLM supports post-training quantization and quantization-aware training. This directly addresses the memory constraint while maintaining acceptable latency, as quantized kernels are optimized for NVIDIA GPUs. It is a standard technique for memory-constrained deployments.

Why this answer

Quantizing weights to INT8 or FP8 reduces memory footprint significantly, and paged KV cache with in-flight batching optimizes runtime memory usage. Together they allow larger models or longer sequences on the same GPU. Other options either require more hardware, increase memory usage, or affect scheduling rather than memory footprint.

Exam trap

The trap here is thinking that increasing tensor parallel size reduces memory on the same hardware, when it actually requires additional GPUs.

31
MCQhard

Refer to the exhibit. What is the most likely cause of the failure based on the log entries?

A.The model weights are corrupted in the repository.
B.The KV cache size is exceeding the allocated GPU memory limit.
C.Network latency is causing the client-side timeout.
D.The GPU driver version is incompatible with TensorRT.
AnswerB

LLMs require a significant amount of memory for the KV cache to store key-value pairs of previous tokens. When the sequence length or batch size grows too large for the allocated memory, the system exhausts VRAM, leading to performance degradation, evictions, and eventual OOM termination during inference.

Why this answer

The logs indicate high processing latency followed by KV cache eviction and finally an OOM error. This sequence suggests that the system is running out of VRAM due to the growing KV cache during long-context generation. As the context length increases, the memory required for the KV cache exceeds the available capacity, forcing evictions, slowing down processing, and ultimately triggering an Out of Memory crash.

Exam trap

Candidates often misidentify the error as a general memory leak or a model weight loading issue, failing to correlate the specific symptom of KV cache eviction with long-context generation demands.

32
MCQmedium

A company needs to deploy a generative AI model that will serve prompts containing regulated customer data. Security policy requires that all inference stays on-premises, that the model be quantized to fit existing GPUs, and that no external network calls occur at runtime. Which deployment approach should the engineer choose?

A.Run the model locally but allow it to fetch updated weights from a vendor endpoint at startup
B.Deploy the model in a public cloud region with a private VPC and customer-managed encryption keys
C.Deploy the model with NVIDIA NIM microservices in an air-gapped environment using locally hosted containers from NVIDIA NGC
D.Use a public hosted API for the model and rely on TLS encryption for data in transit
AnswerC

NVIDIA NIM microservices can be pulled as containers and run fully on-premises, including in air-gapped environments, which satisfies the no-external-calls requirement. They support quantized model variants and expose standard APIs, so the team gets a deployable, optimized inference service without relying on any cloud endpoint.

Why this answer

NVIDIA NIM microservices are distributed as containers that can be deployed on-premises or in air-gapped environments, support quantized model variants, and do not require external network calls at runtime. This satisfies the on-premises, quantized, and no-external-calls requirements simultaneously, unlike hosted or cloud-based options.

Exam trap

The trap here is treating transport encryption or VPC isolation as equivalent to on-premises execution, when the policy specifically forbids external inference calls.

33
MCQmedium

What is the primary role of a Model Repository in the NVIDIA Triton Inference Server architecture?

A.To store training datasets for real-time model retraining.
B.To act as a centralized filesystem for serving multiple models.
C.To handle network traffic load balancing between server nodes.
D.To compile the model into an optimized executable format.
AnswerB

The repository is the source of truth for the server. It organizes models into a hierarchical structure, enabling versioning and easy configuration management. Triton periodically scans this path, allowing updates to be deployed simply by adding files, which is essential for high-availability production AI systems.

Why this answer

The Model Repository acts as a centralized storage location for all deployed models, their configurations, and their versions. Triton monitors this directory to detect when new versions are added or updated, allowing for seamless model updates without restarting the server. This design supports robust MLOps practices by decoupling the model storage from the inference engine runtime, ensuring that deployments remain manageable and version-controlled.

Exam trap

Candidates often think the Model Repository is a database or a training dataset storage, failing to recognize it as a simple filesystem-based interface for Triton to manage model versions.

34
Multi-Selectmedium

Which TWO of the following NVIDIA AI Enterprise tools are specifically designed to optimize and accelerate the deployment of LLMs in containerized environments?

Select 2 answers
A.NVIDIA Triton Inference Server
B.NVIDIA Modulus
C.TensorRT-LLM
D.NVIDIA NeMo Framework
E.NVIDIA DriveWorks
AnswersA, C

Triton is a production-grade inference serving software that supports multiple frameworks and provides advanced features like dynamic and in-flight batching. It is essential for managing model lifecycle, scaling inference workloads, and ensuring high availability for LLM services in enterprise production environments.

Why this answer

NVIDIA Triton Inference Server and TensorRT-LLM are the cornerstones of the NVIDIA AI Enterprise deployment stack. Triton provides a unified serving platform that abstracts the complexities of hardware management, while TensorRT-LLM provides the compiler technology to fuse kernels and quantize weights. Together, they enable developers to move from research models to production-ready services with minimal code changes, ensuring optimized performance across NVIDIA GPUs.

Exam trap

Candidates often confuse general management tools like NVIDIA AI Workbench with deployment-specific runtime engines like TensorRT-LLM, leading them to select development tools rather than serving technologies.

35
MCQhard

A media analytics company runs a TensorRT-LLM optimized GPT-J model on a single NVIDIA A100 80GB GPU using NVIDIA Triton Inference Server. During peak hours, request concurrency rises sharply and the team observes that the GPU is idle for long periods while waiting on host-side tokenization and detokenization. Profiling shows that CPU preprocessing and postprocessing dominate request latency. The team wants to reduce end-to-end latency without changing model weights or adding GPUs. Which Triton feature should they use?

A.Dynamic batching, because it groups requests on the server and reduces the number of model executions.
B.The TensorRT-LLM backend's in-flight batching and built-in tokenizer, which perform tokenization and detokenization inside the backend and keep the GPU busy with continuous batching.
C.Sequence batching with a custom scheduler, because it keeps tokenization on the host and overlaps it with GPU compute across requests.
D.A Triton ensemble that places a Python model for tokenization before the TensorRT-LLM model and a Python model for detokenization after it.
AnswerB

The TensorRT-LLM backend for Triton includes an integrated tokenizer and supports in-flight batching, which moves tokenization into the backend and schedules new requests into ongoing GPU work. This directly targets the host-side bottleneck described in the profile and improves latency by keeping the GPU continuously utilized without adding hardware or changing weights.

Why this answer

The TensorRT-LLM Triton backend provides an integrated tokenizer and in-flight batching, which eliminate the host-side tokenization and detokenization stalls observed in the profile. In-flight batching also continuously admits new requests into the running batch, improving GPU utilization and reducing end-to-end latency without changing model weights or adding GPUs.

Exam trap

The trap here is assuming that generic batching or ensemble orchestration will fix a host-side preprocessing bottleneck, when the actual fix is moving tokenization into the TensorRT-LLM backend and using in-flight batching.

36
Multi-Selectmedium

A team is optimizing an NVIDIA TensorRT-LLM deployment of a 70B model on multiple GPUs. They want to reduce inter-GPU communication overhead and improve throughput. Which two techniques should they consider? (Choose two.)

Select 2 answers
A.Enable pipeline parallelism to split layers into stages.
B.Disable in-flight batching to simplify scheduling.
C.Use FP32 precision for all weights to improve numerical stability.
D.Use tensor parallelism with NVLink-connected GPUs.
E.Increase the KV cache block size to 256 tokens.
AnswersA, D

Pipeline parallelism assigns different layer stages to different GPUs, reducing the frequency of cross-GPU communication compared with tensor parallelism. With micro-batching, stages can overlap and keep GPUs busy. This lowers communication overhead per token and can improve throughput when combined with appropriate batch scheduling.

Why this answer

Tensor parallelism benefits from NVLink's high bandwidth to reduce all-reduce overhead, while pipeline parallelism lowers communication frequency by staging layers across GPUs. Together they address the communication bottleneck for large multi-GPU models. KV cache block size, disabling in-flight batching, and FP32 precision do not reduce inter-GPU communication and can hurt throughput.

Exam trap

The trap here is treating memory-management knobs like KV cache block size as if they also reduce inter-GPU communication overhead.

37
MCQhard

An LLM service on NVIDIA Triton Inference Server experiences high time-to-first-token because the dynamic batcher waits for full batches. The team wants to reduce time-to-first-token while still benefiting from batching. Which adjustment is most appropriate?

A.Set preferred_batch_size to a very large value.
B.Disable dynamic batching entirely for the model.
C.Decrease max_queue_delay_microseconds to shorten batching wait time.
D.Increase max_queue_delay_microseconds to allow larger batches.
AnswerC

The dynamic batcher waits up to max_queue_delay_microseconds before dispatching a batch. Reducing this value shortens the wait, so requests begin processing sooner and time-to-first-token drops. The model still batches whatever requests are available, preserving some throughput benefit while improving responsiveness for interactive workloads.

Why this answer

Time-to-first-token is driven by how long the dynamic batcher holds requests before dispatch. Lowering max_queue_delay_microseconds shortens that wait, so requests start sooner while still being batched with any available peers. Raising the delay, inflating preferred batch size, or disabling batching either increases latency or sacrifices throughput unnecessarily.

Exam trap

The trap here is assuming that larger batches or longer queue delays always improve the user experience, when they actually raise time-to-first-token.

Ready to test yourself?

Try a timed practice session using only Model Deployment questions.