NCP-GENL • Practice Test 21
Free NCP-GENL practice test — 15 questions with explanations. Set 21. No signup required.
A team is deploying a 70B-parameter LLM using NVIDIA Triton Inference Server with TensorRT-LLM backend on a node with four A100 80GB GPUs. They observe that during inference, only one GPU is utilized while the others remain idle. They have configured the model with tensor parallelism set to 1. What is the most likely cause of this underutilization?
Choose an answer to begin — your selection is scored in the full session.
15 questions · instant feedback and full explanations after every question.