A platform team runs a Kubernetes cluster where the NVIDIA GPU Operator is installed and time-slicing is configured with a ConfigMap that advertises four replicas per physical GPU. A data scientist submits a PyTorch training pod requesting nvidia.com/gpu: 1. The pod stays Pending indefinitely, and the scheduler event reads 'Insufficient nvidia.com/gpu'. The node's GPUs are otherwise idle and healthy. Which action most directly resolves the pending state?
Time-slicing only takes effect when the ConfigMap carries the label the device plugin watches and the plugin successfully reloads it. Confirming the node's allocatable nvidia.com/gpu reflects replicas times physical GPUs proves the configuration propagated. If capacity is still one per GPU, the ConfigMap is unlabeled or malformed, and correcting that plus requeuing the pod restores schedulability.
Why this answer
Time-slicing is enabled by a ConfigMap that the NVIDIA device plugin watches; without the expected label the plugin ignores it and keeps advertising one GPU per device. Confirming the ConfigMap is labeled and that the node's allocatable nvidia.com/gpu equals replicas multiplied by physical GPUs verifies the change propagated. Once capacity is correctly republished, the pending pod can be scheduled without altering its resource request.
Exam trap
The trap here is assuming time-slicing creates a distinct resource name such as nvidia.com/gpu.shared, when it actually multiplies the existing nvidia.com/gpu count.