Sample questions
NVIDIA Certified Associate: Generative AI LLMs practice questions
During an experiment, the researcher decides to increase the model's sequence length. What is the most significant side effect they must manage?
When integrating an LLM into a production application, you must protect against prompt injection. Which software engineering pattern is most effective for this purpose?
Which THREE factors should a developer consider when choosing between FP16 and INT8 quantization for a production LLM deployment?
A team fine-tuning a NeMo large language model runs the same training configuration three times and obtains validation loss values of 2.14, 2.31, and 2.09 at the end of the same nu…
When evaluating LLMs for bias, what is the primary purpose of conducting a 'red teaming' exercise?
An enterprise machine learning team is training a large-scale transformer model on a cluster of NVIDIA A100 GPUs using mixed precision (FP16). During the initial training phase, th…
Refer to the exhibit. Which concept of Trustworthy AI is primarily demonstrated by the actions shown in the CLI output?
When implementing Retrieval-Augmented Generation (RAG), why is the choice of 'Chunk Size' critical for model retrieval performance?
A developer is optimizing a retrieval-augmented generation (RAG) pipeline using NVIDIA TensorRT-LLM. They notice excessive latency during the document retrieval phase before the ge…
Refer to the exhibit. How should a data scientist interpret this evaluation result regarding the recent model update?
What is the primary function of the 'Attention' mechanism in Transformer models?
What is the role of 'Temperature' in the context of LLM text generation?
An organization is concerned about 'Model Drift' affecting the trustworthiness of their customer-facing chatbot. What is the most effective way to monitor and address this issue?
You are performing a comparative analysis of two different LLM architectures by visualizing their performance on a RAG (Retrieval-Augmented Generation) benchmark. Which visualizati…
Refer to the exhibit. An engineer is tuning a deployment config. Why is 'enable_cuda_graph' set to true in this JSON configuration?
A bank uses an NVIDIA NIM microservice to host an LLM for loan pre-screening. Before go-live, the risk team must confirm that the model's outputs are reproducible and that any chan…
What is the primary role of 'Loss Scaling' when training deep learning models in FP16 precision?
A team wants to load and run an optimized quantized LLM entirely inside a Python application with minimal dependencies, using a single high-level API that handles engine building a…
A researcher is fine-tuning a large language model using PEFT (Parameter-Efficient Fine-Tuning) techniques. Which method is specifically designed to inject trainable low-rank matri…
A team is fine-tuning an LLM and wants to detect whether individual training examples are causing unusually large gradient updates. They plan to visualize per-example gradient norm…
A developer is using a pretrained large language model for a text summarization task. They want to adapt the model to a domain-specific corpus of legal documents but have limited G…
Why is 'Warmup' used for the learning rate schedule during the initial phase of training large language models?
Refer to the exhibit. A machine learning engineer reviews the monitoring output from a two-GPU distributed training job running on an NVIDIA DGX system. GPU 0 shows low utilization…
During LLM experimentation, what is the primary purpose of maintaining a consistent 'seed' value across different runs?