Courseiva
Model Optimization →mediumMultiple Choice

NCP-GENL Model Optimization Practice Question

Which optimization method should be prioritized when the model's inference performance is bottlenecked by the CPU-to-GPU data transfer overhead?

⚠ Common exam trap

Candidates often suggest optimizing model weights or using faster interconnects like NVLink, ignoring that moving pre-processing to the GPU is the most direct way to eliminate PCIe bus stalling.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

GPU-based Pre-processing

When data transfer between the host and GPU is the bottleneck, the most effective strategy is to move the pre-processing logic onto the GPU itself. Using CUDA-accelerated kernels to perform operations like tokenization or normalization directly on the GPU avoids moving data across the PCIe bus, which is a slow operation that stalls the GPU during the inference pipeline.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    Quantization

    Why it's wrong here

    Quantization reduces the precision of the weights, which helps with compute or memory bandwidth bottlenecks. It does not reduce the data transfer overhead between the CPU and GPU. If the model itself is being sent over the PCIe bus repeatedly, quantization won't solve the latency of that transfer.

  • ✓

    GPU-based Pre-processing

    Why this is correct

    Moving pre-processing to the GPU eliminates the need to transfer intermediate results over the slow PCIe bus. By executing the full pipeline on the GPU device, the system avoids the overhead of host-to-device transfers, which is critical for achieving low-latency inference in real-time generative applications.

  • ✗

    Model Pruning

    Why it's wrong here

    Pruning reduces the number of operations inside the neural network. While this speeds up the computation, it has no impact on the initial data transfer time required to get input data from the host system to the GPU device, which is the specific bottleneck described in the scenario.

  • ✗

    Operator Fusion

    Why it's wrong here

    Operator fusion optimizes the execution of layers within the GPU. It does not have any effect on the data movement between the host CPU and the GPU. If the bottleneck is the PCIe transfer, fusion will not prevent the GPU from waiting for the input data to arrive.

About these practice questions

This NCP-GENL question is part of Courseiva's 352-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official NVIDIA exam blueprint

This NCP-GENL practice question is part of Courseiva's free NVIDIA certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the NCP-GENL exam.