NCA-GENL Software Development Practice Question
A developer is writing a Python service that calls a locally hosted NVIDIA NIM microservice for a Llama model. They want to keep the client code portable so the same class can later target NVIDIA's hosted API endpoints without rewrites. Which client approach fits this goal?
⚠ Common exam trap
The trap here is assuming that NVIDIA NIM requires a proprietary SDK, when its chat surface is intentionally OpenAI-compatible and configurable by base URL.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Use the OpenAI Python client pointed at the NIM base_url, since NIM exposes OpenAI-compatible /v1/chat/completions routes.
Portability comes from targeting the OpenAI-compatible HTTP interface that NIM exposes, letting a single client class switch between a local container URL and NVIDIA hosted endpoints by configuration alone. Direct TensorRT-LLM bindings and the Triton gRPC client both couple code to a specific runtime or protocol, while GPU telemetry tools cannot serve inference requests.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Invoke nvidia-smi from a subprocess and parse GPU utilization to infer model outputs.
Why it's wrong here
nvidia-smi reports hardware telemetry such as utilization and memory, not model inference results. It cannot generate completions and has no relationship to the chat completions API. This approach would never produce usable LLM output and does not address portability between local and hosted endpoints at all.
- ✗
Use the Triton client library with the gRPC protocol and hardcode the model name and version in every call site.
Why it's wrong here
The Triton client speaks Triton's inference protocol, not the OpenAI chat schema, so moving to NVIDIA hosted API endpoints would require rewriting request construction and response parsing. Hardcoding model names and versions further couples the code to one deployment. Portability across local NIM and hosted endpoints is not achieved this way.
- ✗
Call the TensorRT-LLM C++ runtime directly from Python with ctypes and load the engine file inside the service.
Why it's wrong here
The TensorRT-LLM runtime binds to a specific compiled engine and GPU architecture, so the code is tied to one deployment artifact and cannot target remote HTTP endpoints. It also bypasses the NIM server entirely. This maximizes lock-in and complexity, the opposite of the portability requirement, and would not work against hosted API endpoints.
- ✓
Use the OpenAI Python client pointed at the NIM base_url, since NIM exposes OpenAI-compatible /v1/chat/completions routes.
Why this is correct
NIM microservices deliberately expose OpenAI-compatible endpoints, so an OpenAI-style client with a configurable base_url works against the local container and can be repointed at hosted endpoints by changing only the URL and key. This preserves portability and avoids vendor-specific SDK code, which is exactly what the developer wants for later migration.
About these practice questions
One of 367 original NCA-GENL practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official NVIDIA exam blueprint
This NCA-GENL practice question is part of Courseiva's free NVIDIA certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the NCA-GENL exam.