Courseiva
Software Development →hardMultiple Choice

NCA-GENL Software Development Practice Question

A team is serving a 70B-parameter model with NVIDIA Triton Inference Server and TensorRT-LLM. Under concurrent load, GPU memory is exhausted because each request reserves its own large KV cache. Which Triton feature should the team enable to share KV cache blocks across requests that have common prompt prefixes?

⚠ Common exam trap

Many exam-takers confuse batching or instance tuning, which improve utilization, with prefix caching, which is what actually shares KV cache memory across requests.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Paged KV cache with prefix caching enabled in the TensorRT-LLM backend.

Prefix caching combined with paged KV cache in the TensorRT-LLM backend lets Triton reuse cache blocks for requests sharing a prompt prefix, directly reducing memory consumed by duplicate caches. Batching, sequence scheduling, and instance groups tune throughput or state handling but do not deduplicate KV cache blocks.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    Dynamic batching with a larger maximum batch size.

    Why it's wrong here

    Dynamic batching groups requests to improve GPU utilization, but it does not deduplicate or share KV cache memory between sequences. Each sequence in the batch still allocates its own cache blocks, so memory pressure from long or repeated prefixes persists. Batching changes scheduling, not cache reuse.

  • ✗

    Model instance groups with multiple instances per GPU.

    Why it's wrong here

    Instance groups control how many copies of the model run and on which devices. Adding instances on the same GPU divides memory further rather than sharing cache, worsening exhaustion. This setting addresses parallelism and isolation, not KV cache deduplication across requests with shared prefixes.

  • ✗

    Sequence batching with a higher maximum queue delay.

    Why it's wrong here

    Sequence batching manages stateful conversations and controls how long requests wait, but it does not provide cross-request KV cache sharing. Increasing queue delay only changes latency characteristics. Memory consumed by duplicate prefix caches remains, so this does not resolve exhaustion under concurrent load.

  • ✓

    Paged KV cache with prefix caching enabled in the TensorRT-LLM backend.

    Why this is correct

    TensorRT-LLM's paged KV cache stores cache in fixed-size blocks, and prefix caching reuses blocks whose token prefixes match earlier requests. In Triton, enabling this in the backend configuration lets many concurrent requests share common system prompts, drastically cutting GPU memory and raising throughput for the 70B deployment.

About these practice questions

Courseiva writes every NCA-GENL question from scratch — 367 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official NVIDIA exam blueprint

This NCA-GENL practice question is part of Courseiva's free NVIDIA certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the NCA-GENL exam.