NCP-GENL › Model Optimization
This domain covers how NVIDIA TensorRT and TensorRT-LLM turn trained generative models into fast, low-precision inference engines. Questions target quantization (FP16/INT8), kernel selection, KV cache management, layer fusion, and graph surgery, testing whether you know which optimization runs at build time versus runtime and why each exists.
NCP-GENL Model Optimization — All 60 Questions
Every question in this domain with answers and detailed explanations.