NCP-GENL › LLM Architecture
This domain covers the transformer internals that determine how NVIDIA LLMs are built, positioned, and served: attention variants, positional encoding schemes such as RoPE and YaRN, KV-cache behavior, quantization, and parallelism. Questions are scenario-based, asking you to diagnose OOM or scaling failures and pick the architectural feature that fixes them.
NCP-GENL LLM Architecture — All 37 Questions
Every question in this domain with answers and detailed explanations.