NCA-GENL Software Development Practice Question
A developer is writing a Python service that calls an NVIDIA-hosted NIM endpoint for a Llama model. The service must recover gracefully when the endpoint returns HTTP 429 responses during peak traffic, without dropping user requests. Which implementation approach best satisfies this requirement?
⚠ Common exam trap
The trap here is treating a 429 as a transient network fault that should be retried instantly rather than as an explicit signal to reduce request rate.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Implement exponential backoff with jitter and a bounded retry count, then surface a fallback error to the caller.
Rate-limit responses require the client to slow down and retry deliberately. Exponential backoff with jitter spreads retry attempts so multiple clients do not collide, and a bounded retry count keeps latency predictable. Because some requests may still fail after all retries, returning a controlled fallback error preserves a defined service contract for callers rather than hanging or crashing.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Retry the request in a tight loop immediately after each 429 response until it succeeds.
Why it's wrong here
Immediate retries in a tight loop will amplify load on an already saturated endpoint and typically keep receiving 429 responses, wasting client CPU and worsening congestion. The service also risks never returning a response to the caller. Rate limiting responses require backing off, not hammering the same endpoint without delay or a retry ceiling.
- ✗
Increase the client's request timeout value so the 429 responses stop occurring.
Why it's wrong here
A timeout controls how long the client waits for a response, not how many requests the server accepts. Extending it cannot change a rate-limit rejection and may make failures slower to detect. The endpoint would still return 429 until request pressure subsides, so this does not meet the requirement of recovering without dropping requests.
- ✓
Implement exponential backoff with jitter and a bounded retry count, then surface a fallback error to the caller.
Why this is correct
HTTP 429 signals rate limiting, so the client should wait progressively longer between attempts and randomize the delay with jitter to avoid synchronized retry storms from many clients. A bounded retry count prevents indefinite blocking, and a defined fallback keeps the service responsive when retries are exhausted. This is standard resilient client behavior for hosted inference endpoints.
- ✗
Switch the endpoint URL to a different NIM model and continue sending the same request payload unchanged.
Why it's wrong here
Changing to a different model does not address rate limiting and would likely break the application contract because the request payload and expected response schema differ per model. 429 responses are about request volume against a route, not model choice. This approach silently changes application behavior instead of implementing backoff and retry logic.
About these practice questions
This NCA-GENL question is part of Courseiva's 367-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official NVIDIA exam blueprint
This NCA-GENL practice question is part of Courseiva's free NVIDIA certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the NCA-GENL exam.