Courseiva
Software Development →mediumMultiple Choice

NCA-GENL Software Development Practice Question

Which NVIDIA technology enables efficient cross-GPU communication during Tensor Parallelism for large-scale model inference?

⚠ Common exam trap

Candidates often select general networking protocols like Ethernet or standard PCIe, ignoring that NVLink is the specific NVIDIA interconnect required for high-bandwidth communication between GPUs in parallel settings.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

NVLink.

Tensor Parallelism involves splitting model layers across multiple GPUs, which requires high-bandwidth, low-latency communication to synchronize activation values. NVIDIA's NVLink and NVSwitch technologies are designed specifically for this purpose, overcoming the limitations of standard PCIe buses. This is critical for maintaining performance in models that are too large to fit into a single GPU, enabling seamless, high-speed distributed computation.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    CUDA Stream Multi-Processor (SM) scheduling.

    Why it's wrong here

    SM scheduling manages compute tasks within a single GPU. It does not handle communication between multiple distinct GPUs. While efficient for local processing, it does not provide the interconnect fabric necessary for distributed tensor operations across multiple physical devices in a high-performance cluster or server.

  • ✓

    NVLink.

    Why this is correct

    NVLink provides a high-speed direct interconnect between GPUs, allowing them to share data much faster than standard PCIe. This low-latency communication is essential for Tensor Parallelism, where GPU layers must constantly exchange activations to perform synchronized matrix multiplications, directly impacting inference speed in multi-GPU configurations.

  • ✗

    cuDNN Convolutional Layers.

    Why it's wrong here

    cuDNN is a library for deep learning primitives, focusing on optimized mathematical operations like convolutions and activations. It is not an interconnect technology. While it is essential for the underlying math, it does not manage the physical data transfers between GPUs in a distributed system.

  • ✗

    TensorRT-LLM PagedAttention.

    Why it's wrong here

    PagedAttention is an optimization for KV cache memory management, designed to reduce fragmentation. It is purely a memory and compute optimization technique and has no relation to the high-speed data interconnects required for moving data between physical GPUs during distributed tensor parallel model execution.

About these practice questions

One of 367 original NCA-GENL practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official NVIDIA exam blueprint

This NCA-GENL practice question is part of Courseiva's free NVIDIA certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the NCA-GENL exam.