Courseiva
← Back to NVIDIA Certified Associate: Generative AI LLMs questions

Scenario-based practice

Refer to the Exhibit Practice Questions

Practise NVIDIA Certified Associate: Generative AI LLMs practice questions — original exam-style scenarios covering every exam domain, with detailed explanations, wrong-answer analysis, and common exam traps.

15
scenario questions
NCA-GENL
exam code
NVIDIA
vendor

Scenario guide

How to approach refer to the exhibit practice questions

Practise exhibit-style questions that ask you to read a topology, table, command output or diagram before choosing the best answer.

Quick answer

Exhibit-style questions test whether you can read a topology, command output, diagram or table before choosing the best answer.

How to extract the relevant detail from an exhibit.

How topology, command output or routing information affects the answer.

How to avoid answering from memory before reading the evidence.

How to map the exhibit back to the exam objective.

Related practice questions

Related NCA-GENL topic practice pages

Scenario questions usually connect to one or more exam topics. Use these links to review the underlying concepts behind the scenario.

Practice set

Practice scenarios

Question 1mediummultiple choice
Full question →

Refer to the exhibit. Which concept of Trustworthy AI is primarily demonstrated by the actions shown in the CLI output?

Exhibit

2023-10-27 10:15:02 [WARNING] Toxicity score detected: 0.85
2023-10-27 10:15:02 [ACTION] Request blocked by policy: 'safety_strict'
2023-10-27 10:15:03 [DEBUG] Response: 'I cannot fulfill this request.'
Question 2hardmultiple choice
Full question →

Refer to the exhibit. How should a data scientist interpret this evaluation result regarding the recent model update?

Exhibit

{"model": "llama-3-8b", "eval_set": "mmlu", "score_delta": -0.04, "confidence_interval": "[-0.06, -0.02]"}
Question 3mediummultiple choice
Full question →

Refer to the exhibit. An engineer is tuning a deployment config. Why is 'enable_cuda_graph' set to true in this JSON configuration?

Exhibit

{
  "model_name": "llama-3-8b",
  "max_batch_size": 128,
  "precision": "fp16",
  "enable_cuda_graph": true
}
Question 4hardmultiple choice
Full question →

Refer to the exhibit. A machine learning engineer reviews the monitoring output from a two-GPU distributed training job running on an NVIDIA DGX system. GPU 0 shows low utilization despite high memory consumption, while GPU 1 shows high utilization and high memory consumption. What is the most likely root cause of this performance imbalance?

Question 5mediummultiple choice
Full question →

Refer to the exhibit. If a developer increases the 'max_batch_size' in the JSON configuration, what is the primary expected trade-off in the system's performance metrics?

Exhibit

{
  "model_name": "llama-3",
  "precision": "fp16",
  "max_batch_size": 8
}
Question 6hardmultiple choice
Full question →

Refer to the exhibit. A developer implements this NVIDIA NeMo Guardrails configuration. A user submits a query about financial advice. What is the expected behavior of the LLM?

Exhibit

{
  "model_guard": {
    "pii_redaction": true,
    "toxicity_filter": {
      "threshold": 0.85,
      "action": "block"
    },
    "topic_whitelisting": ["tech", "science"]
  }
}
Question 7hardmultiple choice
Full question →

Refer to the exhibit. Which hyperparameter configuration in the provided JSON is directly responsible for preventing overfitting through weight penalty?

Exhibit

{
  "policy_name": "model_access_control",
  "enable_gradient_check": true,
  "optimizer_type": "adam",
  "learning_rate_scheduler": "cosine",
  "weight_decay": 0.05,
  "dropout_rate": 0.1
}
Question 8hardmultiple choice
Full question →

Refer to the exhibit. A monitoring script outputs this JSON for an LLM inference service. What does the 'p99' metric represent in this context?

Exhibit

{"task": "eval_latency", "metric": "p99", "value": 145.2, "unit": "ms", "threshold": 150.0, "status": "PASS"}
Question 9hardmultiple choice
Full question →

Refer to the exhibit. The experiment shows the model is failing to converge and exhibits loss spikes. Which adjustment to the configuration is most likely to stabilize the training process?

Exhibit

config_json: { "optimizer": "adamw", "mixed_precision": "fp16", "grad_clip": 1.0, "warmup_steps": 500, "weight_decay": 0.1 }
Question 10mediummultiple choice
Full question →

Refer to the exhibit. This error occurs during the training of an LLM. What is the most likely cause for this 'device-side assert' error?

Exhibit

Error: CUDA error: device-side assert triggered. Traceback: ... in forward_pass ... loss = criterion(logits, targets) ...
Question 11mediummultiple choice
Full question →

Refer to the exhibit. Which strategy is most effective for resolving this memory error without changing the hardware?

Exhibit

Error: CUDA out of memory. Tried to allocate 2.00 GiB (GPU 0; 24.00 GiB total capacity; 21.50 GiB already allocated)
Question 12mediummultiple choice
Full question →

Refer to the exhibit. Which technique is most appropriate to prevent this specific training failure?

Exhibit

2023-10-27 10:00:01 [INFO] Training Epoch 5/10 completed
2023-10-27 10:00:05 [WARNING] Gradient norm 452.1 exceeds threshold 1.0
2023-10-27 10:00:06 [ERROR] Divergence detected: Loss is NaN
Question 13hardmultiple choice
Full question →

Refer to the exhibit. A developer wants to make the model's output more deterministic and focused on highly probable tokens. Which change should be made to the configuration policy?

Exhibit

config_policy: { 'max_tokens': 1024, 'temperature': 0.7, 'top_p': 0.9, 'presence_penalty': 0.0 }
Question 14hardmultiple choice
Full question →

Refer to the exhibit. What is the primary risk indicated by the provided logs for this training job?

Exhibit

LOG: [INFO] Epoch 10: Step 5000 | LR: 0.0001 | Loss: 1.24 | GPU_MEM: 38.5GB/40GB | UTIL: 98.2%
LOG: [INFO] Epoch 10: Step 5010 | LR: 0.0001 | Loss: 1.25 | GPU_MEM: 39.9GB/40GB | UTIL: 99.1%
LOG: [WARNING] Epoch 10: Step 5020 | LR: 0.0001 | Loss: 1.24 | GPU_MEM: 40.0GB/40GB | UTIL: 99.5%
Question 15hardmultiple choice
Full question →

Refer to the exhibit. The model is failing with an OOM at layer 42 during training. What visualization would most likely point to the cause of the memory fragmentation?

Exhibit

Error: CUDA OOM at layer 42. 
GPU Utilization: 99% 
Memory Fragmentation: 85% 
Attention pattern: Dense-Attention 
Sequence Length: 32k

These NCA-GENL practice questions are part of Courseiva's free NVIDIA certification practice question bank. Courseiva provides original exam-style NCA-GENL questions with detailed explanations, topic-based practice, mock exams, readiness tracking, and study analytics.