Courseiva
Model Optimization →mediumMultiple Choice

NCP-GENL Model Optimization Practice Question

Exhibit

TRT_LOG: [E] Error: Failed to find valid tactics for layer 'Attention_Softmax_0'.
TRT_LOG: [E] Error: Workspace memory limit exceeded for layer 'MatMul_QKV_1'.
TRT_LOG: [I] INFO: Attempting to optimize graph with reduced precision...

Refer to the exhibit. An engineer receives these logs while converting a Transformer model to a TensorRT engine. What is the most appropriate action to resolve this build failure?

⚠ Common exam trap

Candidates often assume the error is due to a lack of overall GPU VRAM, attempting to reduce model precision or batch size instead of specifically increasing the builder's workspace memory allocation.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Increase the builder workspace size

The error indicates that the workspace memory allocated to the TensorRT builder is insufficient to explore the search space of kernels for specific operations. Increasing the workspace size allows the builder to allocate larger temporary memory buffers, which is necessary for complex Transformer operations that require high-memory intermediate calculations to optimize effectively on NVIDIA hardware.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    Reduce the input batch size

    Why it's wrong here

    Reducing the batch size might mitigate the memory bottleneck, but it negatively impacts the total throughput of the inference engine. The engineer should first attempt to increase the allocated workspace size to allow TensorRT to find optimal tactics without sacrificing the desired production batch size requirements.

  • ✓

    Increase the builder workspace size

    Why this is correct

    The error specifically mentions workspace memory limit exceeded. By increasing the memory budget provided to the TensorRT builder, the engine can allocate sufficient scratch space to test various optimized kernels for the Attention and MatMul operations, thereby successfully completing the build process for the engine.

  • ✗

    Switch to a smaller base model

    Why it's wrong here

    Switching models is an unnecessary escalation. The error is a configuration limitation within the TensorRT build environment, not a fundamental issue with the model architecture itself. Increasing the resource allocation for the builder is the standard procedure for resolving kernel search failures in TensorRT.

  • ✗

    Disable INT8 calibration

    Why it's wrong here

    INT8 calibration is an optimization step and is likely unrelated to the workspace memory limit encountered during kernel tactical search. Disabling it would not address the memory allocation error for the MatMul layers; rather, it would likely result in an inefficient FP32 engine build.

Visual reference

Client Recursive Resolver Root DNS (13 root servers) TLD DNS (.com, .org, …) Authoritative example.com query IP addr answer

About these practice questions

Courseiva writes every NCP-GENL question from scratch — 352 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official NVIDIA exam blueprint

This NCP-GENL practice question is part of Courseiva's free NVIDIA certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the NCP-GENL exam.