Courseiva
Software Development →hardMultiple Choice

NCA-GENL Software Development Practice Question

Exhibit

{
  "policy": "strict",
  "quantization": "fp8",
  "max_concurrent_requests": 128,
  "scheduling": "fcfs"
}

Refer to the exhibit. Given the current configuration, what is the primary risk during high-traffic bursts?

⚠ Common exam trap

Students often assume that simple queueing algorithms like First-Come-First-Served work efficiently for LLMs, forgetting that varying token generation lengths cause severe head-of-line blocking under heavy load.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Head-of-line blocking due to the FCFS scheduling policy.

The 'fcfs' (First-Come, First-Served) scheduling policy is suboptimal for LLM inference because it lacks priority handling. In high-traffic scenarios, long requests block shorter ones, leading to 'head-of-line blocking'. This causes latency spikes for all users. For robust production systems, implementing continuous batching or priority-aware scheduling is necessary to ensure that the GPU utilization remains high while maintaining acceptable response times for individual requests.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    FP8 quantization will cause excessive model hallucinations.

    Why it's wrong here

    FP8 quantization is a standard optimization that preserves accuracy while reducing VRAM usage. It does not introduce hallucinations. The risks associated with FP8 are usually related to precision loss if the model hasn't been calibrated, but it does not inherently impact the logic or factual output.

  • ✓

    Head-of-line blocking due to the FCFS scheduling policy.

    Why this is correct

    FCFS processes requests in the order they arrive regardless of their compute requirements. A single long-generation request will occupy GPU resources, causing all subsequent, potentially short requests to wait. This leads to poor overall system latency and creates a bottleneck during high-traffic bursts, significantly degrading user experience.

  • ✗

    The request limit of 128 is too low to saturate the GPU.

    Why it's wrong here

    A limit of 128 concurrent requests is actually quite high for many models. Setting it higher could lead to memory exhaustion (OOM errors) rather than failing to saturate the GPU. The bottleneck is the scheduling policy, not the number of requests the system is configured to permit.

  • ✗

    The strict policy setting prevents dynamic batching.

    Why it's wrong here

    The 'strict' policy usually refers to security or compliance, not the batching algorithm. While potentially restrictive, it is not the cause of performance degradation under load. Performance issues are driven by the FCFS scheduling, which is the incorrect choice for managing diverse, asynchronous inference requests.

About these practice questions

Courseiva writes every NCA-GENL question from scratch — 367 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official NVIDIA exam blueprint

This NCA-GENL practice question is part of Courseiva's free NVIDIA certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the NCA-GENL exam.