A media company is deploying an inference service on a Kubernetes cluster with the NVIDIA GPU Operator installed. The service pods remain in Pending with the message that no nodes have the requested nvidia.com/gpu resource, even though the GPUs are healthy and the driver loads correctly on every node. Which troubleshooting step should the engineer perform first?
The scheduler only sees nvidia.com/gpu capacity after the device plugin registers each GPU with kubelet through the plugin socket. If the device plugin pods are crashing or not scheduled, nodes show no GPU resource and pods stay Pending. Verifying plugin pod status and logs is the most direct first step when the driver itself is healthy.
Why this answer
Kubernetes learns about GPUs through the device plugin framework: the NVIDIA device plugin advertises nvidia.com/gpu for each visible GPU by registering with kubelet. When the driver is healthy but pods report no such resource, the plugin is the missing link. Checking its pod status and logs quickly reveals scheduling failures, crashes, or socket registration errors before deeper investigation.
Exam trap
The trap here is jumping to driver or hardware remediation when the symptom, a healthy driver with no advertised nvidia.com/gpu resource, points squarely at the device plugin registration path.