NCP-GENL Model Deployment Practice Question
A startup is deploying a small LLM for a chatbot on a single NVIDIA L4 GPU using NVIDIA Triton Inference Server. They want to ensure the model is automatically loaded when Triton starts and can be updated without restarting the server. Which Triton feature should they configure?
⚠ Common exam trap
Many exam-takers confuse the ensemble scheduler with model management, when ensembles are for inference pipelines, not loading or versioning.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Model repository with version policy and model control API.
The model repository with version policy allows Triton to automatically load models at startup, and the model control API enables loading and unloading models at runtime. This combination provides both automatic initialization and dynamic updates without server restarts. Other options address inference scheduling or parallelism, not model lifecycle management.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Instance groups with multiple model instances per GPU.
Why it's wrong here
Instance groups control how many instances of a model run on each GPU to increase parallelism. They do not manage model loading at startup or version updates. Configuring instance groups would not satisfy the requirement and could increase memory usage unnecessarily for a small LLM.
- ✓
Model repository with version policy and model control API.
Why this is correct
Triton's model repository stores model files and versions. By configuring the version policy (e.g., latest), Triton can automatically load the latest version at startup. The model control API allows loading and unloading models at runtime without restarting the server. This meets both requirements: automatic loading and dynamic updates.
- ✗
Dynamic batching with preferred batch sizes.
Why it's wrong here
Dynamic batching groups requests to improve throughput but has no role in model loading or version management. It does not enable automatic startup loading or runtime updates. While useful for performance, it is not the feature needed for the described deployment management.
- ✗
Ensemble scheduler with a custom preprocessing model.
Why it's wrong here
The ensemble scheduler chains models for inference but does not manage model loading or versioning. It is used for multi-step pipelines, not for automatic startup loading or dynamic updates. Configuring an ensemble would not address the requirement and adds unnecessary complexity.
About these practice questions
One of 352 original NCP-GENL practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official NVIDIA exam blueprint
This NCP-GENL practice question is part of Courseiva's free NVIDIA certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the NCP-GENL exam.