Sample questions
NVIDIA Certified Professional: Generative AI LLMs practice questions
When profiling an application with NVIDIA Nsight Systems, which TWO metrics are most critical to identify if an application is limited by the PCIe bus?
Refer to the exhibit. Given this NeMo configuration, which prompt modification would best improve the reliability of technical support queries?
Refer to the exhibit. The Triton Inference Server is failing to load a model. The configuration shows two instances assigned to GPU 0. What is the most likely cause and the correct…
Which NVIDIA library is primarily used for optimizing and deploying deep learning inference models?
Which technique is most effective for mitigating data leakage during the training of an LLM on time-series-related document data?
Which hardware architecture feature is specifically leveraged by TensorRT to accelerate FP16 and INT8 matrix multiplications?
Refer to the exhibit. An engineer is fine-tuning an LLM using the provided configuration. What is the primary purpose of applying 'yarn' scaling in this architecture?
When optimizing a Generative AI model using NVIDIA TensorRT-LLM, which component is primarily responsible for managing the KV cache to minimize memory fragmentation?
What is the primary function of the 'rank' parameter in LoRA?
A team is training a 13B-parameter LLM on 8 NVIDIA A100 GPUs using NVIDIA NeMo. They observe that the all-reduce communication during data-parallel training consumes nearly 40% of…
When fine-tuning a Large Language Model using Low-Rank Adaptation (LoRA), which architectural component is primarily modified to reduce computational overhead while maintaining per…
When optimizing a model using NVIDIA TensorRT, what is the primary benefit of enabling 'layer fusion' during the optimization process?
A team is deploying a 13B-parameter decoder-only LLM on a single NVIDIA A100 40GB GPU for a real-time chatbot. During load, the process runs out of memory even though the model wei…
An LLM inference service on NVIDIA Triton Inference Server is experiencing occasional out-of-memory (OOM) errors on the GPU. The team wants to monitor GPU memory usage to predict a…
What is the primary role of a Model Repository in the NVIDIA Triton Inference Server architecture?
Which TWO of the following techniques are best suited for reducing the latency of LLM inference on NVIDIA GPUs?
When preparing a proprietary technical manual dataset for a RAG pipeline, which data preprocessing step is most critical to ensure the LLM avoids hallucinations regarding specific…
What is the primary purpose of 'Few-Shot Prompting' in the context of LLM optimization?
Which TWO actions should be part of a robust incident response plan for an LLM deployment failing in production?
A research team wants to train a large decoder-only LLM where each token's representation is computed independently of token order, then inject order information afterward. They ar…
A financial institution uses NVIDIA NeMo Guardrails to enforce ethical guidelines in its customer-facing LLM. During testing, the model occasionally generates responses that violat…
Refer to the exhibit. Which prompt engineering technique would best force the model to prioritize technical detail over marketing language?
When building an NVIDIA NeMo LLM application for automated document review, which THREE of the following prompt design choices are critical for ensuring high-quality output? (Selec…
In the context of model optimization, why is 'graph surgery' sometimes required before building a TensorRT engine?