NCP-GENL Production Monitoring and Reliability Practice Question
An LLM inference service on NVIDIA Triton Inference Server is experiencing intermittent failures. The operations team wants to set up alerting to detect when the GPU memory utilization exceeds 90% for more than 5 minutes, as this could lead to out-of-memory errors. Which combination of tools should they use to achieve this?
⚠ Common exam trap
The trap here is assuming that Triton's built-in metrics include GPU memory utilization, when in fact they are inference-specific and require DCGM-Exporter for GPU telemetry.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
NVIDIA DCGM-Exporter and Prometheus Alertmanager
DCGM-Exporter collects GPU metrics and exposes them to Prometheus. Prometheus evaluates alerting rules based on thresholds and durations, and Alertmanager sends notifications. This setup is ideal for detecting high GPU memory utilization over time, helping prevent OOM errors in LLM inference.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✓
NVIDIA DCGM-Exporter and Prometheus Alertmanager
Why this is correct
DCGM-Exporter exposes GPU metrics, including memory utilization, to Prometheus. Prometheus can then evaluate alerting rules, such as memory utilization >90% for 5 minutes, and trigger alerts via Alertmanager. This combination is standard for GPU monitoring and alerting in production environments, providing reliable and scalable detection of potential OOM conditions.
- ✗
nvidia-smi with custom scripting and email notifications
Why it's wrong here
nvidia-smi can be scripted to poll GPU memory, but it is not designed for continuous monitoring or alerting at scale. Custom scripts are brittle and lack integration with modern monitoring stacks. This approach is not recommended for production-grade alerting due to maintenance overhead and potential reliability issues.
- ✗
NVIDIA Nsight Systems and Prometheus
Why it's wrong here
Nsight Systems is a profiling tool for performance analysis, not a monitoring tool. It does not export metrics to Prometheus or provide continuous telemetry. Using it for alerting is impractical and not its intended purpose. Prometheus requires a metrics exporter like DCGM-Exporter to collect GPU data.
- ✗
NVIDIA Triton Inference Server metrics and Grafana
Why it's wrong here
Triton metrics provide inference-related data but not GPU memory utilization. Grafana is a visualization tool and does not natively generate alerts without a data source that includes GPU metrics. While Grafana can alert on Triton metrics, it cannot directly monitor GPU memory without an exporter like DCGM-Exporter.
About these practice questions
Courseiva writes every NCP-GENL question from scratch — 352 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official NVIDIA exam blueprint
This NCP-GENL practice question is part of Courseiva's free NVIDIA certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the NCP-GENL exam.