NCA-GENL Core Machine Learning and AI Knowledge Practice Question
Exhibit
config.yaml: optimizer: adamw precision: bf16 sharding: fsdp fsdp_config: backward_prefetch: backward_pre mixed_precision: true
Refer to the exhibit. The configuration shows the use of FSDP with mixed precision. What is the main benefit of using 'bf16' (Bfloat16) over 'fp16' in this context?
⚠ Common exam trap
Candidates often believe BF16 is chosen for speed gains alone, missing that its primary architectural advantage over FP16 is the larger dynamic range that prevents gradient underflow without complex scaling.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
It provides a larger dynamic range, preventing gradient underflow.
Bfloat16 provides the same dynamic range as FP32, preventing the underflow issues common with FP16 when calculating gradients during deep learning training. This stability allows for training without loss-scaling techniques, simplifying the pipeline and improving convergence consistency. In high-performance training, using BF16 is the standard for modern GPUs like the NVIDIA A100/H100, as it ensures stability without sacrificing speed or requiring complex hyperparameter tuning for numeric precision.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
It reduces the size of the model weights on the disk.
Why it's wrong here
BF16, like FP16, is a 16-bit format, meaning it occupies the same amount of space as FP16. Neither format changes the size of the model on the disk compared to the other; their primary differences are in numerical range, exponent handling, and the resulting training stability of the model.
- ✓
It provides a larger dynamic range, preventing gradient underflow.
Why this is correct
BF16 uses an 8-bit exponent field identical to FP32, which allows it to represent a much wider range of values than FP16. This prevents the underflow of small gradient values that often occur in deep learning, significantly simplifying training stability by removing the need for explicit dynamic loss scaling.
- ✗
It is natively supported on older Pascal-based architectures.
Why it's wrong here
BF16 support is a hardware-level feature introduced in newer NVIDIA architectures like Ampere (A100) and Hopper (H100). Pascal-based architectures do not have native hardware acceleration for BF16, making it unusable or extremely inefficient on older hardware compared to FP16, which is supported across a wider range of legacy GPUs.
- ✗
It doubles the throughput compared to FP32.
Why it's wrong here
While BF16 is faster than FP32, it does not provide a guaranteed 'double' throughput benefit in every scenario. Its primary advantage is numerical stability compared to FP16, not just raw speed. Speed gains depend on the specific hardware generation and kernel implementation being used in the training loop.
About these practice questions
Courseiva writes every NCA-GENL question from scratch — 367 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official NVIDIA exam blueprint
This NCA-GENL practice question is part of Courseiva's free NVIDIA certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the NCA-GENL exam.