NCP-GENL Model Deployment Practice Question
A healthcare company is deploying an LLM for clinical note summarization using NVIDIA Triton Inference Server. They must ensure that only authorized users can access the model and that all inference requests are logged for audit. Which Triton feature should they configure to enforce authentication and authorization?
⚠ Common exam trap
The trap here is assuming Triton has built-in user authentication, when it actually relies on external components like reverse proxies for security.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Triton's HTTP/REST and gRPC endpoints with a reverse proxy that performs OAuth 2.0 token validation.
Triton Inference Server does not include native authentication or authorization. The standard pattern is to deploy a reverse proxy that validates OAuth 2.0 tokens and forwards authorized requests to Triton. This separates security concerns from inference and allows audit logging at the proxy. Other options describe scheduling or non-existent features that do not enforce access control.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Triton's built-in model access control list (ACL) configured via the model repository.
Why it's wrong here
Triton does not have a built-in model ACL feature in the model repository. Access control is not managed through repository configuration. While Triton supports model control APIs, they are for loading and unloading models, not for user authentication. Relying on a non-existent feature would fail to meet the security requirement.
- ✓
Triton's HTTP/REST and gRPC endpoints with a reverse proxy that performs OAuth 2.0 token validation.
Why this is correct
Triton itself does not provide built-in authentication or authorization. The recommended approach is to place a reverse proxy (such as NGINX or Envoy) in front of Triton to handle OAuth 2.0 token validation and access control. This satisfies the requirement for authorized access while Triton focuses on inference. Logging can be handled at the proxy or application level for audit.
- ✗
Triton's dynamic batching configuration with priority levels.
Why it's wrong here
Dynamic batching and priority levels control request scheduling and performance, not security. They do not authenticate users or enforce authorization. Configuring them would not meet the healthcare compliance requirement. Security must be handled at the network or application layer, not through batching parameters.
- ✗
Triton's ensemble scheduler with a custom authentication model.
Why it's wrong here
The ensemble scheduler is designed to chain multiple models for inference, not to perform user authentication. A custom authentication model would still require integration with Triton's request path and is not a standard security feature. It adds complexity and does not provide the required OAuth 2.0 validation or audit logging out of the box.
About these practice questions
This NCP-GENL question is part of Courseiva's 352-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official NVIDIA exam blueprint
This NCP-GENL practice question is part of Courseiva's free NVIDIA certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the NCP-GENL exam.