NCP-GENL Model Optimization Practice Question
Exhibit
TRT_LOG: [E] Error: Failed to find valid tactics for layer 'Attention_Softmax_0'. TRT_LOG: [E] Error: Workspace memory limit exceeded for layer 'MatMul_QKV_1'. TRT_LOG: [I] INFO: Attempting to optimize graph with reduced precision...
Refer to the exhibit. An engineer receives these logs while converting a Transformer model to a TensorRT engine. What is the most appropriate action to resolve this build failure?
⚠ Common exam trap
Candidates often assume the error is due to a lack of overall GPU VRAM, attempting to reduce model precision or batch size instead of specifically increasing the builder's workspace memory allocation.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Increase the builder workspace size
The error indicates that the workspace memory allocated to the TensorRT builder is insufficient to explore the search space of kernels for specific operations. Increasing the workspace size allows the builder to allocate larger temporary memory buffers, which is necessary for complex Transformer operations that require high-memory intermediate calculations to optimize effectively on NVIDIA hardware.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Reduce the input batch size
Why it's wrong here
Reducing the batch size might mitigate the memory bottleneck, but it negatively impacts the total throughput of the inference engine. The engineer should first attempt to increase the allocated workspace size to allow TensorRT to find optimal tactics without sacrificing the desired production batch size requirements.
- ✓
Increase the builder workspace size
Why this is correct
The error specifically mentions workspace memory limit exceeded. By increasing the memory budget provided to the TensorRT builder, the engine can allocate sufficient scratch space to test various optimized kernels for the Attention and MatMul operations, thereby successfully completing the build process for the engine.
- ✗
Switch to a smaller base model
Why it's wrong here
Switching models is an unnecessary escalation. The error is a configuration limitation within the TensorRT build environment, not a fundamental issue with the model architecture itself. Increasing the resource allocation for the builder is the standard procedure for resolving kernel search failures in TensorRT.
- ✗
Disable INT8 calibration
Why it's wrong here
INT8 calibration is an optimization step and is likely unrelated to the workspace memory limit encountered during kernel tactical search. Disabling it would not address the memory allocation error for the MatMul layers; rather, it would likely result in an inefficient FP32 engine build.
Visual reference
About these practice questions
Courseiva writes every NCP-GENL question from scratch — 352 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official NVIDIA exam blueprint
This NCP-GENL practice question is part of Courseiva's free NVIDIA certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the NCP-GENL exam.