NCP-GENL · domain
Production Monitoring and Reliability
This domain covers keeping LLM inference healthy on NVIDIA Triton Inference Server once it is live: watching GPU memory and utilization, collecting metrics into Prometheus, tuning dynamic batching, and responding to production incidents. Questions are scenario-based, asking you to pick the monitoring tool, configuration change, or incident-response action that fixes a described failure.
Focused practice
Practice Production Monitoring and Reliability questions
Scored sessions drawing only from this domain — pick a length below.
Start 20-question practice test →What this domain covers
What to know about Production Monitoring and Reliability
Be able to select the right NVIDIA monitoring path (DCGM Exporter plus Triton's Prometheus metrics), read GPU memory and latency signals, and adjust Triton dynamic batching. The key is matching the observed symptom to the correct metric and configuration fix, not guessing.
Using NVIDIA DCGM and DCGM Exporter to scrape GPU memory and utilization into Prometheus
Exposing Triton metrics via its Prometheus endpoint for inference latency and request counts
Diagnosing OOM risk by tracking GPU memory allocation against model and batch size
Tuning Triton dynamic batching parameters when p99 latency rises but GPU utilization is low
Watch out for
Common Production Monitoring and Reliability exam traps
- ▸Assuming high GPU utilization always means the bottleneck, when low utilization with high p99 latency points to batching or queueing misconfiguration.
- ▸Confusing DCGM GPU-level metrics with Triton model-level metrics; OOM prediction needs GPU memory data, not just inference counts.
- ▸Treating incident response as only restarting the service, ignoring rollback, alerting, and root-cause steps the question expects.
Question index
All Production Monitoring and Reliability questions (35)
Click any question to see the full explanation, or start a practice session above.
When auditing an NVIDIA Triton deployment for security, which action is most critical to protect sensitive inference data?
Medium2A team is using NVIDIA Triton Inference Server to serve multiple LLMs. They want to automatically detect when a model's inference latency degrades beyond acceptable thresholds and trigger an alert. Which Triton feature should they configure?
Medium3An enterprise is running a mission-critical generative AI application on an NVIDIA DGX cluster. The MLOps team notices occasional silent GPU memory corruption during long-running inference jobs that do not trigger hard crashes. Which monitoring tool and strategy should be utilized for early detection?
Hard4A production LLM inference service on NVIDIA Triton Inference Server is being monitored for reliability. The team wants to implement effective logging to diagnose issues such as high latency and errors. Which TWO logging practices are recommended for a production LLM environment? (Choose two.)
Hard5Refer to the exhibit. What is the most likely reason for the high P99 latency despite low GPU utilization?
Hard6Which metric provides the best indication of 'inference queue saturation' in a Triton deployment?
Medium7An LLM inference service on NVIDIA Triton Inference Server is experiencing intermittent failures. The operations team wants to set up alerting to detect when the GPU memory utilization exceeds 90% for more than 5 minutes, as this could lead to out-of-memory errors. Which combination of tools should they use to achieve this?
Medium8A production LLM service on NVIDIA Triton Inference Server is deployed across multiple GPUs. The team notices that one GPU consistently shows higher latency for inference requests compared to others, despite similar utilization. Which NVIDIA tool should be used to investigate per-GPU performance discrepancies and identify bottlenecks?
Hard9You are deploying a large language model on NVIDIA Triton Inference Server in a Kubernetes cluster. To ensure high availability and reliability, which TWO practices should you implement? (Choose two.)
Medium10An LLM application deployed on NVIDIA Triton Inference Server is experiencing intermittent latency spikes. Which metric is most critical to monitor to determine if the issue stems from GPU compute saturation?
Medium11An LLM inference service running on NVIDIA Triton Inference Server is configured with model ensembles. During production monitoring, the operations team notices that the end-to-end latency reported by the client is significantly higher than the sum of individual model latencies reported by Triton's metrics. Which Triton feature should be investigated to identify the source of the additional latency?
Hard12Which TWO of the following telemetry types are essential for detecting 'model drift' in a production LLM deployment?
Medium13Refer to the exhibit. The inference server reports an OOM error while running multiple LLMs. What configuration change most effectively improves reliability without upgrading the hardware?
Hard14Refer to the exhibit. Why did the system fail to trigger an alert despite high latency?
Hard15What is the primary benefit of using NVIDIA DCGM (Data Center GPU Manager) for monitoring production LLMs?
Medium16A production LLM inference service on NVIDIA Triton Inference Server is experiencing intermittent slowdowns. The team wants a single metric that best indicates whether the GPU is the bottleneck during these periods. Which metric should they monitor?
Easy17An LLM inference service on NVIDIA Triton Inference Server uses dynamic batching with a max_queue_delay of 500 microseconds. During a load test, p99 latency exceeds the SLA while GPU utilization remains below 40%. Which change should you make first to improve latency without reducing throughput?
Medium18An LLM inference service deployed on NVIDIA Triton Inference Server is experiencing occasional failures under high load. The team wants to implement proactive monitoring to predict and prevent these failures. Which two metrics should be prioritized for early detection of potential issues? (Choose two.)
Hard19A team is deploying an LLM for real-time inference using NVIDIA Triton Inference Server. They need to monitor the health and performance of the GPU to ensure reliability. Which NVIDIA tool provides comprehensive GPU telemetry, including utilization, memory, temperature, and power, and integrates with Prometheus for monitoring?
Easy20Which of the following is considered a best practice for logging in a production LLM environment?
Easy21You are responsible for the reliability of an LLM inference service running on NVIDIA Triton Inference Server across a fleet of A100 GPUs. The service is deployed with dynamic batching enabled, but during peak hours you observe that end-to-end latency for some requests exceeds the SLO while GPU utilization remains moderate. You suspect that the dynamic batching configuration is causing the issue. Which Triton configuration parameter should you adjust to directly limit the maximum time a request waits in the scheduler queue before being batched?
Medium22A team is deploying a large language model on NVIDIA Triton Inference Server in a production environment. They need to ensure that the model server can automatically recover from GPU failures without manual intervention. Which feature of Triton should they configure to achieve this?
Easy23A generative AI application built on NVIDIA Triton Inference Server is deployed in a Kubernetes cluster with GPU nodes. The operations team wants to detect silent data corruption in model outputs, which could occur due to GPU memory errors. They plan to implement a monitoring solution using NVIDIA Data Center GPU Manager (DCGM). Which DCGM feature should they enable to detect and alert on GPU memory errors that could lead to silent data corruption?
Hard24What is the primary function of the 'Triton Model Control' API in a production environment?
Medium25An LLM inference service on NVIDIA Triton Inference Server is experiencing a gradual increase in P99 latency over several hours without a corresponding increase in request rate. GPU utilization remains stable at around 60%. Which action should the team take first to diagnose the root cause?
Hard26An LLM inference service on NVIDIA Triton Inference Server uses TensorRT-LLM as the backend. During peak load, you notice that the first token latency (time to first token) is high, but subsequent tokens are generated quickly. Which metric should you monitor to diagnose this issue?
Hard27An LLM inference service on NVIDIA Triton Inference Server is experiencing occasional out-of-memory (OOM) errors on the GPU. The team wants to monitor GPU memory usage to predict and prevent OOM conditions. Which NVIDIA tool provides real-time GPU memory metrics that can be integrated with Prometheus for alerting?
Medium28Which TWO actions should be part of a robust incident response plan for an LLM deployment failing in production?
Medium29A generative AI application deployed on NVIDIA Triton Inference Server with NVIDIA AI Enterprise is experiencing silent data corruption in model outputs. The MLOps team needs to implement monitoring to detect such issues early. Which two actions should be taken to enhance observability for silent data corruption? (Choose two.)
Hard30A production LLM service on NVIDIA Triton Inference Server uses dynamic batching. During peak load, the 99th percentile latency increases significantly, but GPU utilization remains at 60%. Which configuration change is most likely to improve latency while maintaining throughput?
Hard31An enterprise deploying a large language model on NVIDIA Triton Inference Server needs to track GPU utilization, memory allocation, and custom inference latency histograms in Prometheus. Which approach should the MLOps engineer implement to ensure robust production observability without overloading the inference execution threads?
Medium32In the context of NVIDIA AI Enterprise, what is the primary purpose of using 'NVIDIA Triton Model Analyzer'?
Easy33A production LLM inference service on NVIDIA Triton Inference Server runs on a multi-GPU node. You observe that one GPU reports ECC XID errors and the model's throughput gradually degrades. Which NVIDIA tool should you use to monitor GPU health and set up alerts for these errors?
Hard34When evaluating LLM reliability under stress, what is the primary goal of conducting 'Chaos Engineering' on a Triton inference cluster?
Hard35A production LLM inference service runs on NVIDIA Triton Inference Server across multiple GPUs. The SRE team wants to detect when the service starts returning incorrect or degraded responses compared to a baseline, even when latency and throughput remain normal. Which monitoring approach is most appropriate?
MediumOther domains
All NCP-GENL exam domains
Frequently asked questions
- What does the Production Monitoring and Reliability domain cover on the NCP-GENL exam?
- Be able to select the right NVIDIA monitoring path (DCGM Exporter plus Triton's Prometheus metrics), read GPU memory and latency signals, and adjust Triton dynamic batching. The key is matching the observed symptom to the correct metric and configuration fix, not guessing.
- How many questions are in this domain?
- This page lists all 35 Production Monitoring and Reliability questions in the NCP-GENL question bank. The actual exam draws from this domain proportionally to its weighting in the official exam blueprint.
- What is the best way to practise this domain?
- Start with a short focused session (10 questions) to identify gaps, then work through explanations. Repeat with a longer session once the weak areas feel solid.
- Can I practise only Production Monitoring and Reliability questions?
- Yes — the session launcher on this page filters questions to this domain only. Choose any session length for inline explanations and scoring.