Courseiva

CCNA Production Monitoring and Reliability Questions

35 questions · Production Monitoring and Reliability · All types, answers revealed

1
MCQmedium

When auditing an NVIDIA Triton deployment for security, which action is most critical to protect sensitive inference data?

A.Disabling the GPU logging
B.Enabling TLS for client-server communication
C.Increasing the memory buffer
D.Switching to an unsecured port
AnswerB

TLS ensures that the data transmitted between the client application and the Triton server is encrypted. This is essential for preventing unauthorized eavesdropping on sensitive prompts or responses, fulfilling basic data security requirements for production LLM deployments.

Why this answer

Securing the communication channel is the first line of defense in protecting sensitive user input data sent to an LLM. Utilizing TLS/SSL encryption prevents man-in-the-middle attacks where data could be intercepted during transit. In professional production environments, this is a standard requirement for compliance and data privacy, ensuring that prompt data and generated completions remain confidential between the client and the inference server.

Exam trap

Candidates frequently select internal model weight encryption or file system permissions, forgetting that data in transit across network boundaries remains exposed without transport-layer security.

2
MCQmedium

A team is using NVIDIA Triton Inference Server to serve multiple LLMs. They want to automatically detect when a model's inference latency degrades beyond acceptable thresholds and trigger an alert. Which Triton feature should they configure?

A.Model Analyzer
B.Triton's built-in model warmup
C.Triton Metrics API with Prometheus and Alertmanager
D.Triton's dynamic batching configuration
AnswerC

Triton exposes a Metrics API endpoint that provides detailed inference metrics, including latency percentiles, request counts, and queue times. Integrating this with Prometheus allows scraping and storing these metrics, while Alertmanager can evaluate rules and trigger alerts when latency exceeds defined thresholds. This combination provides real-time monitoring and automated alerting, exactly what the team needs for production reliability.

Why this answer

Triton's Metrics API, when paired with Prometheus and Alertmanager, provides the necessary observability to monitor inference latency in real time and automatically alert when thresholds are exceeded. This integration is a standard practice for production monitoring, enabling proactive incident response. Other options are either static optimization tools or lack alerting functionality.

Exam trap

The trap here is confusing performance tuning tools like Model Analyzer or dynamic batching with runtime monitoring and alerting features.

3
MCQhard

An enterprise is running a mission-critical generative AI application on an NVIDIA DGX cluster. The MLOps team notices occasional silent GPU memory corruption during long-running inference jobs that do not trigger hard crashes. Which monitoring tool and strategy should be utilized for early detection?

A.Deploy NVIDIA Data Center GPU Manager (DCGM) with Prometheus exporter to monitor XID errors and hardware ECC events continuously.
B.Increase the frequency of client-side HTTP ping probes sent from the load balancer to the API gateway.
C.Write a custom bash script that runs nvidia-smi every five minutes and parses plain text output for warning strings.
D.Rely on standard Kubernetes node liveness probes to automatically restart pods when system memory usage exceeds 90%.
AnswerA

DCGM provides specialized diagnostic tests and continuous health monitoring, including ECC error tracking and XID error detection. Configuring DCGM to raise alerts on uncorrectable memory errors or specific XID failure codes allows operators to isolate failing GPUs before they impact production workloads.

Why this answer

DCGM (Data Center GPU Manager) provides specialized diagnostic tests and continuous health monitoring, including ECC error tracking and XID error detection. Configuring DCGM to raise alerts on uncorrectable memory errors or specific XID failure codes allows operators to isolate failing GPUs before they impact production workloads.

Exam trap

Engineers often rely solely on standard Kubernetes pod health checks, completely missing hardware-level silent errors that occur beneath the container runtime layer.

4
Multi-Selecthard

A production LLM inference service on NVIDIA Triton Inference Server is being monitored for reliability. The team wants to implement effective logging to diagnose issues such as high latency and errors. Which TWO logging practices are recommended for a production LLM environment? (Choose two.)

Select 2 answers
A.Log every inference request and response payload for full traceability.
B.Use structured logging (e.g., JSON) to enable efficient parsing and querying.
C.Disable logging in production to maximize inference throughput.
D.Log GPU temperature and power metrics at debug level for every inference.
E.Include a unique request ID in logs to correlate events across distributed components.
AnswersB, E

Structured logging formats like JSON make it easier to parse and query logs programmatically. This is crucial for automated monitoring and alerting systems. It allows filtering by fields such as request ID, model name, and latency, enabling quick identification of issues in production LLM deployments without manual log inspection.

Why this answer

Including a unique request ID enables end-to-end tracing, which is vital for diagnosing latency and errors in distributed LLM inference pipelines. Structured logging facilitates automated parsing and querying, allowing efficient monitoring and alerting. Together, they provide robust observability without excessive overhead or security risks.

Exam trap

The trap here is assuming that more logging is always better, leading to practices like logging full payloads or per-inference GPU metrics, which can harm performance and privacy.

5
MCQhard

Refer to the exhibit. What is the most likely reason for the high P99 latency despite low GPU utilization?

A.The GPU compute units are failing
B.Inappropriate concurrency settings
C.The model is too large for memory
D.Network bandwidth is saturated
AnswerB

Low GPU utilization combined with high queue depth confirms that the system is not processing enough concurrent requests to utilize the hardware fully. The bottleneck is the configuration limiting the number of parallel inferences, causing requests to queue unnecessarily.

Why this answer

When GPU utilization is low but P99 latency is high, the bottleneck is usually not the compute capability but rather the request queuing or orchestration layer. High queue depth indicates that requests are waiting for a slot, likely due to a restrictive concurrent request setting or inefficient backend task scheduling. This scenario highlights the importance of matching inference concurrency settings with the hardware's throughput capacity to avoid artificial bottlenecks.

Exam trap

Candidates often blame the model size or GPU compute capacity. If utilization is low, the hardware is idle, meaning the bottleneck is clearly in the request handling or queuing configuration.

6
MCQmedium

Which metric provides the best indication of 'inference queue saturation' in a Triton deployment?

A.Total GPU memory usage
B.Request queue duration
C.GPU temperature
D.System clock speed
AnswerB

Queue duration measures the time a request waits before being executed. An increasing trend in this metric is the definitive sign that the inference server is unable to process incoming requests as quickly as they arrive, indicating clear saturation.

Why this answer

Queue duration is the most direct indicator of saturation. It measures how long an inference request spends waiting in the server's input buffer before it is processed by the GPU. In a production environment, monitoring this metric is crucial to identify when the server's request capacity has been exceeded, allowing for auto-scaling triggers to provision more instances and maintain performance standards during traffic spikes.

Exam trap

Candidates often choose GPU utilization or throughput. These metrics indicate how well the hardware is working, but they do not reveal if requests are being delayed in the server's input buffer.

7
MCQmedium

An LLM inference service on NVIDIA Triton Inference Server is experiencing intermittent failures. The operations team wants to set up alerting to detect when the GPU memory utilization exceeds 90% for more than 5 minutes, as this could lead to out-of-memory errors. Which combination of tools should they use to achieve this?

A.NVIDIA DCGM-Exporter and Prometheus Alertmanager
B.nvidia-smi with custom scripting and email notifications
C.NVIDIA Nsight Systems and Prometheus
D.NVIDIA Triton Inference Server metrics and Grafana
AnswerA

DCGM-Exporter exposes GPU metrics, including memory utilization, to Prometheus. Prometheus can then evaluate alerting rules, such as memory utilization >90% for 5 minutes, and trigger alerts via Alertmanager. This combination is standard for GPU monitoring and alerting in production environments, providing reliable and scalable detection of potential OOM conditions.

Why this answer

DCGM-Exporter collects GPU metrics and exposes them to Prometheus. Prometheus evaluates alerting rules based on thresholds and durations, and Alertmanager sends notifications. This setup is ideal for detecting high GPU memory utilization over time, helping prevent OOM errors in LLM inference.

Exam trap

The trap here is assuming that Triton's built-in metrics include GPU memory utilization, when in fact they are inference-specific and require DCGM-Exporter for GPU telemetry.

8
MCQhard

A production LLM service on NVIDIA Triton Inference Server is deployed across multiple GPUs. The team notices that one GPU consistently shows higher latency for inference requests compared to others, despite similar utilization. Which NVIDIA tool should be used to investigate per-GPU performance discrepancies and identify bottlenecks?

A.NVIDIA Nsight Systems
B.NVIDIA Triton Inference Server metrics
C.NVIDIA Data Center GPU Manager (DCGM)
D.NVIDIA System Management Interface (nvidia-smi)
AnswerA

Nsight Systems is a system-wide profiling tool that can capture GPU activities, including kernel execution times and memory transfers, across multiple GPUs. It can highlight per-GPU performance differences by showing timelines and bottlenecks. In this scenario, it can identify why one GPU is slower, such as inefficient kernel usage or thermal throttling, enabling targeted optimization.

Why this answer

Nsight Systems provides detailed profiling of GPU activities, including kernel execution, memory transfers, and synchronization. By capturing a timeline across all GPUs, it can reveal why one GPU lags, such as longer kernel runtimes or thermal issues. This level of detail is essential for diagnosing per-GPU performance discrepancies and optimizing LLM inference in a multi-GPU deployment.

Exam trap

The trap here is relying on monitoring tools like DCGM or nvidia-smi that provide metrics but not the granular profiling needed to identify kernel-level bottlenecks on a specific GPU.

9
Multi-Selectmedium

You are deploying a large language model on NVIDIA Triton Inference Server in a Kubernetes cluster. To ensure high availability and reliability, which TWO practices should you implement? (Choose two.)

Select 2 answers
A.Enable Triton's model warmup to reduce cold-start latency.
B.Use a single large GPU instance to maximize performance and reduce complexity.
C.Set up health checks and readiness probes for the Triton server pods.
D.Configure multiple model instances per GPU to increase concurrency.
E.Deploy the Triton server as a Kubernetes Deployment with multiple replicas across nodes.
AnswersC, E

Health checks and readiness probes allow Kubernetes to detect unhealthy pods and stop routing traffic to them, and to restart failed pods. This ensures that only healthy replicas serve requests, improving reliability and availability. It is a standard practice for production deployments.

Why this answer

High availability in Kubernetes requires redundancy and health monitoring. Deploying multiple replicas across nodes ensures that a single node failure does not take down the service. Readiness probes ensure traffic is only sent to healthy pods.

Together, these practices minimize downtime and maintain reliable inference serving.

Exam trap

The trap here is focusing on performance optimizations like model warmup or multiple instances, which do not provide redundancy across failures.

10
MCQmedium

An LLM application deployed on NVIDIA Triton Inference Server is experiencing intermittent latency spikes. Which metric is most critical to monitor to determine if the issue stems from GPU compute saturation?

A.Total system CPU usage
B.Network interface throughput
C.SM (Streaming Multiprocessor) occupancy
D.Total GPU memory capacity
AnswerC

SM occupancy directly measures the percentage of active warps compared to the maximum supported by the GPU. High occupancy indicates that the streaming multiprocessors are busy executing instructions, serving as the primary metric for identifying compute-bound bottlenecks in neural network inference tasks.

Why this answer

Monitoring GPU utilization alone is insufficient because it does not distinguish between compute saturation and memory bandwidth bottlenecks. NVML-based metrics specifically tracking SM (Streaming Multiprocessor) occupancy provide the granular insight required to identify whether the execution cores are fully utilized. This metric is essential for capacity planning and ensuring that the inference throughput meets strict service level agreements under high concurrent request loads in production environments.

Exam trap

Candidates often monitor general host CPU utilization or basic GPU memory usage, failing to recognize that SM occupancy directly reflects compute execution core saturation.

11
MCQhard

An LLM inference service running on NVIDIA Triton Inference Server is configured with model ensembles. During production monitoring, the operations team notices that the end-to-end latency reported by the client is significantly higher than the sum of individual model latencies reported by Triton's metrics. Which Triton feature should be investigated to identify the source of the additional latency?

A.Model ensemble scheduler overhead and inter-model data transfer
B.GPU memory utilization causing queuing of inference requests
C.Network latency between the client and the Triton server
D.Dynamic batching configuration causing delays in request aggregation
AnswerA

Triton's ensemble scheduler manages the execution of multiple models and handles data transfer between them. This can introduce overhead not captured in individual model latency metrics. Investigating ensemble scheduling and data transfer times can reveal bottlenecks. The discrepancy between client-side and server-side latency suggests overhead in the ensemble pipeline, such as serialization or scheduling delays.

Why this answer

In Triton model ensembles, the ensemble scheduler coordinates execution and data transfer between models. This adds overhead that is not attributed to any single model's latency. Monitoring ensemble-specific metrics and tracing the pipeline can identify where the extra time is spent, such as in data serialization or scheduling delays.

Exam trap

The trap here is assuming that client-side latency should always match server-side model latency, overlooking the overhead introduced by ensemble scheduling and inter-model communication.

12
Multi-Selectmedium

Which TWO of the following telemetry types are essential for detecting 'model drift' in a production LLM deployment?

Select 2 answers
A.Output token distribution statistics
B.GPU power consumption levels
C.User feedback and semantic similarity scores
D.Hardware temperature monitoring
E.Network latency between nodes
AnswersA, C

Monitoring shifts in the distribution of generated tokens helps identify if the model is defaulting to repetitive or nonsensical outputs. Significant deviations from established baseline token patterns are often a leading indicator that the input distribution has shifted, triggering a drift condition.

Why this answer

Model drift refers to the degradation of model output quality over time as real-world data deviates from training distributions. Tracking output tokens and response semantic consistency allows engineers to identify when the model starts producing incoherent or irrelevant results. Proactive monitoring of these metrics is critical because drift often occurs silently, potentially damaging user trust before traditional error logs or system performance metrics indicate a failure.

Exam trap

Candidates often select 'latency' or 'throughput' as drift metrics. While these indicate system health, they do not measure the actual semantic quality or output distribution drift of the LLM's generated content.

13
MCQhard

Refer to the exhibit. The inference server reports an OOM error while running multiple LLMs. What configuration change most effectively improves reliability without upgrading the hardware?

A.Increase the batch size
B.Implement KV cache block paging
C.Enable FP32 precision mode
D.Increase the number of threads
AnswerB

KV cache paging divides the cache into smaller, manageable blocks, significantly reducing external fragmentation. This allows the system to utilize memory more effectively, enabling higher concurrency without requiring a physical hardware upgrade, which is the standard strategy for resolving VRAM saturation issues.

Why this answer

The exhibit shows that the VRAM is almost completely exhausted, likely due to excessive KV cache allocation for concurrent models. Implementing PagedAttention or optimizing the KV cache block size allows for more efficient memory management. This approach improves reliability by preventing memory fragmentation and allowing the system to handle higher request concurrency within existing hardware constraints, which is vital for production systems requiring high uptime and predictable performance.

Exam trap

Candidates often suggest reducing the batch size or model precision as the primary fix. While these help, KV cache paging is the specific architectural solution for memory fragmentation in high-concurrency LLM serving.

14
MCQhard

Refer to the exhibit. Why did the system fail to trigger an alert despite high latency?

A.The error rate threshold was too low
B.The alert is suppressed by a duration requirement
C.The Prometheus backend is down
D.The logging level is set too high
AnswerB

Monitoring systems usually require a threshold to be exceeded for a specific time duration to prevent noise from transient spikes. Since the alert did not fire despite exceeding the 500ms limit, a duration-based trigger condition is the most probable cause for the suppression.

Why this answer

Alerting systems often require sustained violation of thresholds to avoid 'flapping' or false positives. The exhibit shows the current latency is 850ms, which exceeds the threshold, but the alerting logic likely requires consecutive samples or a window-based average to trigger. This is a common reliability feature in monitoring stacks to ensure that transient spikes do not disrupt operational teams with unnecessary alerts during non-critical fluctuations.

Exam trap

Candidates often assume the system failed due to a misconfigured threshold or a bug in the monitoring agent, ignoring the common practice of duration-based alert suppression to prevent false positives.

15
MCQmedium

What is the primary benefit of using NVIDIA DCGM (Data Center GPU Manager) for monitoring production LLMs?

A.To train LLMs faster
B.To detect hardware-level issues
C.To optimize model weights
D.To generate natural language output
AnswerB

DCGM is specifically designed to provide deep hardware insights, including temperature, power, and memory errors. Identifying these problems early is critical for infrastructure reliability, as it allows for maintenance to be scheduled before a component causes a hard failure.

Why this answer

DCGM provides hardware-level telemetry that is essential for proactive maintenance and reliability. Unlike higher-level metrics that only show software performance, DCGM can identify physical issues such as ECC memory errors, thermal throttling, or failing power supplies. This level of visibility is crucial for anticipating hardware failure before it results in a service outage, allowing operations teams to migrate workloads gracefully and maintain high system reliability.

Exam trap

Candidates often confuse DCGM with software-level monitoring tools like Prometheus or Grafana. They incorrectly assume it monitors model accuracy or inference throughput rather than physical hardware health and GPU-specific telemetry.

16
MCQeasy

A production LLM inference service on NVIDIA Triton Inference Server is experiencing intermittent slowdowns. The team wants a single metric that best indicates whether the GPU is the bottleneck during these periods. Which metric should they monitor?

A.GPU utilization (SM utilization)
B.Disk I/O operations per second on the model repository
C.Network throughput between clients and the server
D.CPU utilization of the inference server host
AnswerA

GPU utilization, specifically SM utilization, directly measures the percentage of time the GPU's streaming multiprocessors are active. In LLM inference, high SM utilization sustained near 100% strongly suggests the GPU is the bottleneck. Monitoring this metric during slowdowns helps determine if the GPU is saturated by compute or memory-bound operations, guiding optimization such as batching or model quantization.

Why this answer

GPU utilization (SM utilization) is the most direct indicator of whether the GPU's compute units are saturated. During LLM inference, if slowdowns coincide with near-100% SM utilization, the GPU is the bottleneck. Other metrics like CPU, network, or disk I/O are less likely to explain intermittent latency spikes when the GPU is the primary compute resource.

Exam trap

The trap here is assuming that host CPU or network metrics are sufficient to diagnose GPU-bound LLM inference slowdowns.

17
MCQmedium

An LLM inference service on NVIDIA Triton Inference Server uses dynamic batching with a max_queue_delay of 500 microseconds. During a load test, p99 latency exceeds the SLA while GPU utilization remains below 40%. Which change should you make first to improve latency without reducing throughput?

A.Reduce max_queue_delay to a smaller value to allow quicker dispatch of requests.
B.Increase the max_batch_size in the model configuration to allow larger batches.
C.Enable model warmup to reduce first-inference latency.
D.Switch to a larger GPU with more memory to increase batch capacity.
AnswerA

The p99 latency is caused by requests waiting in the dynamic batching queue for up to 500 microseconds. Lowering max_queue_delay reduces this wait, improving tail latency. Because GPU utilization is low, throughput is not constrained by batch size, so reducing the delay will not significantly reduce throughput and directly addresses the latency symptom.

Why this answer

The high p99 latency with low GPU utilization indicates that requests are spending too long in the dynamic batching queue. Reducing max_queue_delay allows requests to be dispatched sooner, cutting tail latency. Since GPU utilization is low, throughput is not limited by batch size, so this change improves latency without harming throughput.

Exam trap

The trap here is assuming that increasing batch size or GPU resources will fix latency, when the real issue is the batching delay.

18
Multi-Selecthard

An LLM inference service deployed on NVIDIA Triton Inference Server is experiencing occasional failures under high load. The team wants to implement proactive monitoring to predict and prevent these failures. Which two metrics should be prioritized for early detection of potential issues? (Choose two.)

Select 2 answers
A.Disk I/O throughput
B.CPU utilization of the Triton server process
C.Number of active connections
D.GPU memory utilization per instance
E.Request queue time
AnswersD, E

GPU memory utilization is critical because LLMs are memory-intensive. Rising memory usage can indicate memory leaks or increased batch sizes, leading to out-of-memory errors. Monitoring this metric allows proactive scaling or optimization before failures occur. It directly relates to resource exhaustion, a common cause of inference failures under load.

Why this answer

Under high load, LLM inference failures often stem from resource exhaustion or overload. GPU memory utilization is critical because insufficient memory leads to out-of-memory errors. Request queue time indicates if the system is falling behind, predicting timeouts.

Together, these metrics provide early warning of impending failures, enabling proactive measures.

Exam trap

The trap here is focusing on general system metrics like CPU or disk I/O, which are less relevant for GPU-bound LLM inference under load.

19
MCQeasy

A team is deploying an LLM for real-time inference using NVIDIA Triton Inference Server. They need to monitor the health and performance of the GPU to ensure reliability. Which NVIDIA tool provides comprehensive GPU telemetry, including utilization, memory, temperature, and power, and integrates with Prometheus for monitoring?

A.NVIDIA Triton Inference Server metrics endpoint
B.NVIDIA Nsight Systems
C.NVIDIA System Management Interface (nvidia-smi)
D.NVIDIA Data Center GPU Manager (DCGM)
AnswerD

DCGM is designed for data center GPU monitoring and management. It provides detailed telemetry on GPU utilization, memory, temperature, power, and more. It integrates with Prometheus via DCGM-Exporter, enabling robust monitoring and alerting for production LLM deployments. This makes it the ideal tool for ensuring GPU health and performance.

Why this answer

DCGM is the comprehensive GPU telemetry tool from NVIDIA, offering metrics on utilization, memory, temperature, power, and more. It integrates with Prometheus through DCGM-Exporter, enabling scalable monitoring and alerting. For production LLM deployments, DCGM is essential for ensuring GPU reliability and performance.

Exam trap

The trap here is confusing basic GPU monitoring utilities like nvidia-smi with full-featured telemetry tools like DCGM, which are required for production-grade monitoring.

20
MCQeasy

Which of the following is considered a best practice for logging in a production LLM environment?

A.Logging only full chat transcripts
B.Using structured JSON logs
C.Logging all data to standard output only
D.Disabling logging to save disk space
AnswerB

JSON-formatted logs are machine-readable and easily parsed by log aggregators like ELK or Splunk. This structure allows for reliable querying of specific fields such as model version, latency, and error codes, which is critical for maintaining production reliability.

Why this answer

Structured logging (typically in JSON format) is the gold standard for production environments because it allows for efficient searching, indexing, and automated analysis. In LLM applications, keeping track of request metadata, model versions, and latency in a structured way is essential for debugging and auditing. This approach enables operations teams to quickly correlate specific issues with model outputs and performance metrics across thousands of requests.

Exam trap

Engineers often rely on unstructured plain-text print statements, assuming human readability is sufficient for large-scale automated pipeline monitoring and analysis.

21
MCQmedium

You are responsible for the reliability of an LLM inference service running on NVIDIA Triton Inference Server across a fleet of A100 GPUs. The service is deployed with dynamic batching enabled, but during peak hours you observe that end-to-end latency for some requests exceeds the SLO while GPU utilization remains moderate. You suspect that the dynamic batching configuration is causing the issue. Which Triton configuration parameter should you adjust to directly limit the maximum time a request waits in the scheduler queue before being batched?

A.Increase the `instance_group` count to add more model instances.
B.Set the `max_queue_delay_microseconds` parameter to a lower value.
C.Set the `max_batch_size` parameter to a lower value.
D.Enable the `priority_levels` parameter to prioritize certain requests.
AnswerB

The `max_queue_delay_microseconds` parameter in Triton's dynamic batching configuration specifies the maximum time a request can wait in the queue before the scheduler dispatches the batch. By lowering this value, you reduce the maximum latency contributed by batching, which directly addresses the observed SLO violations. This is the intended parameter for controlling batching-induced latency while still allowing some batching for efficiency.

Why this answer

The `max_queue_delay_microseconds` parameter in Triton's dynamic batching configuration sets the maximum time a request can wait in the scheduler queue before being dispatched for inference. Lowering this value reduces the worst-case latency introduced by batching, directly addressing the SLO violations observed during peak hours. While other parameters like `max_batch_size` and `instance_group` affect performance, only `max_queue_delay_microseconds` explicitly bounds the queue wait time, making it the correct adjustment for this scenario.

Exam trap

The trap here is assuming that reducing `max_batch_size` will directly reduce latency, when the real control for queue wait time is `max_queue_delay_microseconds`.

22
MCQeasy

A team is deploying a large language model on NVIDIA Triton Inference Server in a production environment. They need to ensure that the model server can automatically recover from GPU failures without manual intervention. Which feature of Triton should they configure to achieve this?

A.Triton's health check endpoints and Kubernetes liveness probes
B.Instance groups with multiple GPU instances
C.Model repository polling
D.Model ensembles
AnswerA

Triton exposes health check endpoints that Kubernetes can use with liveness probes to detect when the server is unhealthy due to GPU failure. Kubernetes can then restart the pod, potentially on a healthy GPU node, providing automatic recovery. This combination is a standard practice for achieving high availability in containerized Triton deployments.

Why this answer

To automatically recover from GPU failures, Triton should be deployed in an orchestrated environment like Kubernetes. Triton provides health check endpoints that, when integrated with Kubernetes liveness probes, allow the orchestrator to detect an unhealthy server and restart it, possibly on a different node with a working GPU. This approach ensures minimal downtime without manual intervention.

Exam trap

The trap here is assuming that Triton's internal features like instance groups or ensembles provide automatic failover, when actually orchestration-level health checks are required for recovery.

23
MCQhard

A generative AI application built on NVIDIA Triton Inference Server is deployed in a Kubernetes cluster with GPU nodes. The operations team wants to detect silent data corruption in model outputs, which could occur due to GPU memory errors. They plan to implement a monitoring solution using NVIDIA Data Center GPU Manager (DCGM). Which DCGM feature should they enable to detect and alert on GPU memory errors that could lead to silent data corruption?

A.DCGM configuration with `EnableAutoBoost` set to true.
B.DCGM health checks with the `memory` watch.
C.DCGM profiling metrics with `DCGM_FI_DEV_GPU_UTIL`.
D.DCGM diagnostics with the `-r` option for a full run.
AnswerB

DCGM health checks include a `memory` watch that monitors GPU memory for errors such as ECC errors (correctable and uncorrectable). Enabling this watch allows DCGM to detect memory errors that could cause silent data corruption. It can trigger alerts or take corrective actions based on policy. This is the appropriate feature to monitor for GPU memory issues that may affect model output integrity, making it the correct choice for detecting silent data corruption.

Why this answer

DCGM health checks with the `memory` watch continuously monitor GPU memory for ECC errors and other faults that can cause silent data corruption. This feature can be configured to raise alerts when errors exceed thresholds, enabling proactive remediation. In contrast, diagnostics are run on-demand, utilization metrics reflect workload, and AutoBoost affects clocks.

For detecting memory errors that could corrupt LLM outputs, the memory watch is the correct DCGM feature to enable.

Exam trap

The trap here is confusing DCGM diagnostics (run manually) with health checks (continuous monitoring) and assuming that any memory-related test will provide real-time alerting.

24
MCQmedium

What is the primary function of the 'Triton Model Control' API in a production environment?

A.To optimize GPU memory frequency
B.To load and unload models dynamically
C.To adjust the model's architecture
D.To capture raw network packets
AnswerB

The primary purpose of this API is to enable runtime control over model loading and unloading. This capability is critical for production systems that need to update models or deploy new versions without causing service interruptions for active end users.

Why this answer

The Model Control API allows for dynamic model management without requiring a full server restart. This is essential for zero-downtime deployments, where new model versions must be loaded, tested, or swapped out while the server continues to handle existing inference traffic. By decoupling model management from server availability, the API significantly increases the flexibility and operational reliability of large-scale, continuously evolving AI model deployments.

Exam trap

Candidates often assume the API is for monitoring model performance or accuracy. They overlook the word 'Control' and fail to associate it with the dynamic lifecycle management of model instances.

25
MCQhard

An LLM inference service on NVIDIA Triton Inference Server is experiencing a gradual increase in P99 latency over several hours without a corresponding increase in request rate. GPU utilization remains stable at around 60%. Which action should the team take first to diagnose the root cause?

A.Increase the number of model instances to handle the load.
B.Check for memory leaks in the model or Triton server process.
C.Profile the inference pipeline with NVIDIA Nsight Systems to identify bottlenecks.
D.Review the model's input data for changes in sequence length.
AnswerB

A gradual increase in latency without increased load or GPU utilization suggests a resource leak, such as memory fragmentation or a memory leak in the model or server. Over time, this can cause more frequent garbage collection or swapping, increasing latency. Checking for memory leaks is a logical first step to identify the root cause before applying fixes.

Why this answer

The gradual increase in P99 latency with stable GPU utilization and request rate suggests a resource leak, such as memory fragmentation or a leak in the model or Triton process. Checking for memory leaks is the most direct first step to diagnose the issue. Scaling or profiling may be premature and could mask the root cause.

Reviewing input data is unlikely given the stable utilization.

Exam trap

The trap here is assuming that latency increases always require scaling, without considering resource leaks that don't affect utilization.

26
MCQhard

An LLM inference service on NVIDIA Triton Inference Server uses TensorRT-LLM as the backend. During peak load, you notice that the first token latency (time to first token) is high, but subsequent tokens are generated quickly. Which metric should you monitor to diagnose this issue?

A.Time to first token (TTFT)
B.GPU utilization
C.Requests per second (RPS)
D.Inter-token latency (ITL)
AnswerA

TTFT directly measures the latency from request submission to the first generated token. In LLM inference, the prefill phase (processing the input prompt) dominates TTFT. Monitoring TTFT helps identify if the prefill is slow due to large batch sizes, long prompts, or insufficient compute, allowing targeted optimization.

Why this answer

Time to first token (TTFT) is the key metric for diagnosing prefill latency in LLM inference. Since the issue is high TTFT but fast subsequent tokens, monitoring TTFT allows you to isolate the prefill phase and optimize it, for example by reducing batch size or using more efficient attention kernels.

Exam trap

The trap here is focusing on GPU utilization or throughput instead of the specific latency metric that reflects the prefill phase.

27
MCQmedium

An LLM inference service on NVIDIA Triton Inference Server is experiencing occasional out-of-memory (OOM) errors on the GPU. The team wants to monitor GPU memory usage to predict and prevent OOM conditions. Which NVIDIA tool provides real-time GPU memory metrics that can be integrated with Prometheus for alerting?

A.NVIDIA Nsight Systems
B.NVIDIA Triton Inference Server metrics endpoint
C.NVIDIA System Management Interface (nvidia-smi)
D.NVIDIA Data Center GPU Manager (DCGM)
AnswerD

DCGM is a suite of tools for managing and monitoring NVIDIA GPUs in clusters. It provides detailed metrics including GPU memory usage, and DCGM Exporter can expose these metrics in Prometheus format. This integration enables real-time monitoring and alerting on memory usage, helping predict and prevent OOM conditions in production LLM inference.

Why this answer

DCGM is designed for data center GPU monitoring and provides comprehensive metrics, including memory usage. With DCGM Exporter, these metrics can be scraped by Prometheus, enabling real-time alerts when memory usage approaches limits. This proactive monitoring allows the team to adjust batch sizes or resource allocation before OOM errors occur, ensuring reliable LLM inference.

Exam trap

The trap here is confusing profiling tools or basic command-line utilities with a production-grade monitoring solution that integrates with Prometheus for alerting.

28
Multi-Selectmedium

Which TWO actions should be part of a robust incident response plan for an LLM deployment failing in production?

Select 2 answers
A.Rolling back to a previous stable model version
B.Deploying the model to more GPUs immediately
C.Reviewing recent logs and telemetry metrics
D.Turning off all security monitoring systems
E.Forcing a reboot of all servers
AnswersA, C

If a new model deployment causes issues, a rollback is the fastest way to restore service stability. This action minimizes downtime and gives the engineering team the necessary time to debug the problematic version in a non-production environment.

Why this answer

A robust incident response plan focuses on rapid mitigation and root cause analysis. Immediately reverting to a known good version reduces the impact on users, while analyzing logs and telemetry provides the necessary data to understand the failure. These steps ensure that service is restored quickly, which is the primary objective of reliability engineering in generative AI systems serving critical user traffic.

Exam trap

Candidates often select 'retraining the model' as a primary incident response action. Retraining is a long-term fix, not an immediate incident response step for a failing production deployment.

29
Multi-Selecthard

A generative AI application deployed on NVIDIA Triton Inference Server with NVIDIA AI Enterprise is experiencing silent data corruption in model outputs. The MLOps team needs to implement monitoring to detect such issues early. Which two actions should be taken to enhance observability for silent data corruption? (Choose two.)

Select 2 answers
A.Enable NVIDIA Data Center GPU Manager (DCGM) to monitor GPU ECC error counts and XID errors.
B.Configure Triton's model warmup to run dummy inferences at startup to stabilize performance.
C.Increase the batch size for all inference requests to improve throughput and reduce per-request overhead.
D.Set up logging of all inference inputs and outputs to a centralized system for later analysis.
E.Implement output validation by comparing inference results against a known-good baseline for a sample of requests.
AnswersA, E

DCGM tracks GPU hardware errors such as ECC memory errors and XID errors, which can cause silent data corruption. By monitoring these metrics, the team can detect hardware-level issues that might lead to incorrect model outputs. This is a proactive measure to identify and alert on potential corruption sources before they affect production results.

Why this answer

Silent data corruption can stem from hardware faults or software bugs. Monitoring GPU ECC and XID errors via DCGM catches hardware-induced corruption, while output validation against a baseline detects incorrect results regardless of cause. Together, they provide both infrastructure and application-level observability, enabling early detection and mitigation of silent data corruption in production LLM inference.

Exam trap

The trap here is focusing on performance tuning or logging instead of active detection methods like hardware error monitoring and output validation, which are essential for catching silent data corruption.

30
MCQhard

A production LLM service on NVIDIA Triton Inference Server uses dynamic batching. During peak load, the 99th percentile latency increases significantly, but GPU utilization remains at 60%. Which configuration change is most likely to improve latency while maintaining throughput?

A.Decrease the maximum batch size and adjust the batch timeout to reduce waiting
B.Enable model instances to increase parallelism on the same GPU
C.Switch to a larger GPU with more memory bandwidth
D.Increase the maximum batch size in the dynamic batching configuration
AnswerA

With GPU utilization at 60%, the GPU has spare capacity. High P99 latency during peak load suggests requests are waiting too long for batches to fill. Reducing the maximum batch size and lowering the batch timeout (e.g., from 100 microseconds to 50) allows smaller batches to be processed more quickly, reducing queue time. This can improve latency while still maintaining acceptable throughput because the GPU is not fully utilized.

Why this answer

When GPU utilization is low but latency is high, the bottleneck is often the batching mechanism waiting to accumulate requests. Reducing the maximum batch size and batch timeout allows requests to be processed sooner, lowering queue time and P99 latency. Since the GPU has spare capacity, this change can improve latency without sacrificing throughput.

Exam trap

The trap here is assuming that increasing batch size always improves performance, overlooking the impact of batch timeout on latency.

31
MCQmedium

An enterprise deploying a large language model on NVIDIA Triton Inference Server needs to track GPU utilization, memory allocation, and custom inference latency histograms in Prometheus. Which approach should the MLOps engineer implement to ensure robust production observability without overloading the inference execution threads?

A.Write a custom Python wrapper around every generate call to log execution duration directly to a shared text file.
B.Query the NVIDIA Management Library via a cron job every second to poll current GPU statistics and write them to a database.
C.Enable the native Prometheus metrics endpoint in Triton Inference Server configuration and configure Prometheus to scrape it.
D.Disable all internal instrumentation layers to maximize raw throughput, relying solely on external black-box API health probes.
AnswerC

Enabling the native Prometheus metrics endpoint allows the server to expose internal hardware and performance counters directly over an HTTP port. This scraping mechanism operates asynchronously from the main inference pipeline, preventing metric collection overhead from degrading GPU compute efficiency.

Why this answer

Enabling Triton's native Prometheus metrics endpoint allows the server to expose internal hardware and performance counters directly over an HTTP port. This scraping mechanism operates asynchronously from the main inference pipeline, preventing metric collection overhead from degrading GPU compute efficiency or increasing inference request latency.

Exam trap

Candidates often assume custom Python application loggers must manually wrap every inference call, mistakenly believing built-in server metrics lack the required granularity for enterprise LLM latency tracking.

32
MCQeasy

In the context of NVIDIA AI Enterprise, what is the primary purpose of using 'NVIDIA Triton Model Analyzer'?

A.To perform real-time model training
B.To automate performance optimization and configuration
C.To detect GPU hardware failures
D.To encrypt model weights at rest
AnswerB

The tool systematically benchmarks different combinations of runtime parameters to identify the most efficient setup. This automated approach ensures the model meets performance targets reliably without the human error inherent in manual configuration of complex inference engines.

Why this answer

The Model Analyzer is a critical tool for determining the optimal configuration for models deployed on Triton Inference Server. By automating the testing of various batch sizes and concurrency levels, it ensures that models are tuned for maximum throughput and minimum latency. This tuning is essential for maintaining reliability and cost-efficiency in large-scale production, preventing performance degradation caused by suboptimal configuration settings that often go unnoticed in manual deployments.

Exam trap

Test-takers often confuse Model Analyzer with profiling tools used for training or code execution time, missing its specific role in searching optimal Triton configurations.

33
MCQhard

A production LLM inference service on NVIDIA Triton Inference Server runs on a multi-GPU node. You observe that one GPU reports ECC XID errors and the model's throughput gradually degrades. Which NVIDIA tool should you use to monitor GPU health and set up alerts for these errors?

A.NVIDIA Data Center GPU Manager (DCGM)
B.NVIDIA TensorRT
C.NVIDIA Nsight Systems
D.NVIDIA Triton Model Analyzer
AnswerA

DCGM is designed for monitoring and managing NVIDIA data center GPUs. It can detect ECC errors, XID errors, and other health metrics, and can be integrated with alerting systems. In this scenario, DCGM would provide the necessary visibility into GPU health and enable proactive alerts when errors occur, helping maintain reliability.

Why this answer

NVIDIA DCGM is the standard tool for monitoring GPU health in data centers. It tracks ECC and XID errors, temperature, power, and other metrics, and can trigger alerts. For a production LLM service experiencing hardware errors, DCGM provides the necessary observability to detect and respond to GPU failures.

Exam trap

The trap here is confusing profiling tools like Nsight Systems or optimization tools like TensorRT with health monitoring tools.

34
MCQhard

When evaluating LLM reliability under stress, what is the primary goal of conducting 'Chaos Engineering' on a Triton inference cluster?

A.To increase the total throughput of the model
B.To verify system recovery and failover logic
C.To reduce the physical power usage of GPUs
D.To train the model on noisy data
AnswerB

Validating that the system automatically handles failures is the cornerstone of chaos engineering. By simulating real-world failures, engineers can confirm that the failover mechanisms, such as health checks and load rebalancing, function correctly in an automated and reliable manner.

Why this answer

Chaos Engineering involves deliberately injecting failures, such as network latency or node crashes, to test system resilience. The goal is to verify that the system can automatically recover and maintain service levels despite these disruptions. This is vital for production-grade LLM reliability, as it identifies hidden weaknesses in the failover logic or load balancing configurations before a real-world outage occurs, ensuring the cluster behaves predictably under adverse conditions.

Exam trap

Candidates often confuse Chaos Engineering with performance testing or load testing. The primary distinction is that Chaos Engineering is specifically about verifying system resilience and recovery during failures.

35
MCQmedium

A production LLM inference service runs on NVIDIA Triton Inference Server across multiple GPUs. The SRE team wants to detect when the service starts returning incorrect or degraded responses compared to a baseline, even when latency and throughput remain normal. Which monitoring approach is most appropriate?

A.Monitor GPU utilization and memory usage via NVIDIA DCGM to detect anomalies.
B.Implement model output quality checks by comparing responses against a baseline using statistical or embedding-based metrics.
C.Use NVIDIA Nsight Systems to profile the inference pipeline and identify bottlenecks.
D.Set up alerting on Triton Inference Server's error rate and request latency percentiles.
AnswerB

Comparing live model outputs to a baseline using metrics such as BLEU, ROUGE, or embedding similarity can detect semantic drift or degradation that doesn't affect latency or throughput. This directly addresses the need to monitor response correctness. It is the most appropriate because it focuses on output quality, which is the core concern here.

Why this answer

To detect degraded or incorrect responses while latency and throughput are normal, monitoring must focus on output quality. Comparing model outputs to a baseline using statistical or embedding-based metrics can reveal semantic drift or accuracy drops. Resource and performance metrics like GPU utilization, error rates, and latency are insufficient because they do not reflect the correctness of generated text.

Exam trap

The trap here is assuming that normal latency and throughput imply correct model behavior, overlooking the need for output quality monitoring.

Ready to test yourself?

Try a timed practice session using only Production Monitoring and Reliability questions.