NCP-GENL GPU Acceleration and Optimization • 30 Questions
30 NCP-GENL GPU Acceleration and Optimization practice questions with answers and explanations. Free, no signup.
An AI engineer is deploying a large language model on an NVIDIA A100 GPU using TensorRT-LLM. During inference profiling, they notice that token generation latency is higher than expected due to memory bandwidth bottlenecks during the autoregressive decoding phase. Which optimization technique should be applied first to mitigate this bandwidth limitation?
Choose an answer to begin — your selection is scored in the full session.
30 questions · instant feedback and full explanations after every question.