Courseiva

NCP-GENL GPU Acceleration and Optimization Practice Question

A developer is optimizing a BERT-based model for inference on an NVIDIA T4 GPU using TensorRT. The model has a fixed input sequence length of 128. Profiling shows that the kernel execution time is high due to many small operations. Which TensorRT feature should they use to reduce kernel launch overhead and improve latency?

⚠ Common exam trap

The trap here is focusing on precision or fusion when the bottleneck is CPU-side kernel launch overhead, which CUDA graphs specifically target.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Apply CUDA graphs to capture the entire inference graph and replay it with a single launch.

CUDA graphs are designed to reduce kernel launch overhead by capturing a static sequence of operations and replaying it as a single unit. For a fixed-shape model like the BERT model with sequence length 128, the graph can be captured once and reused across inferences. This eliminates the CPU-side launch latency for each kernel, directly improving latency.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    Increase the batch size to amortize kernel launch overhead across more samples.

    Why it's wrong here

    Increasing batch size improves GPU utilization and amortizes overhead per sample, but it does not reduce the absolute launch overhead per kernel. The latency per inference may still be high if the batch is not fully utilized, and it may exceed memory limits. This is a workaround, not a direct fix for launch overhead.

  • ✗

    Use TensorRT's builder optimization level 5 to enable aggressive layer fusion.

    Why it's wrong here

    Builder optimization levels control tactics like kernel selection and fusion, but they do not directly eliminate kernel launch overhead. Layer fusion can reduce the number of kernels, but the effect is limited if the graph is already well-fused. This setting is more about tuning performance than addressing launch overhead specifically.

  • ✓

    Apply CUDA graphs to capture the entire inference graph and replay it with a single launch.

    Why this is correct

    CUDA graphs capture a sequence of kernels and their dependencies into a single graph, then replay it with one launch, drastically reducing CPU launch overhead. For a fixed-shape model like this BERT with sequence length 128, CUDA graphs are ideal because the graph can be captured once and reused. This directly addresses the many small operations causing high kernel execution time.

  • ✗

    Enable FP16 precision and calibrate with a representative dataset.

    Why it's wrong here

    FP16 precision reduces memory bandwidth and increases compute throughput, but it does not address the overhead of launching many small kernels. The model may still suffer from launch latency if the operations remain fragmented. While FP16 is beneficial, it is not the primary solution for reducing kernel launch overhead in this scenario.

About these practice questions

Courseiva writes every NCP-GENL question from scratch — 352 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official NVIDIA exam blueprint

This NCP-GENL practice question is part of Courseiva's free NVIDIA certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the NCP-GENL exam.