AI0-001 AI Infrastructure and Technologies Practice Question
A healthcare analytics team is preparing to fine-tune a 7-billion-parameter open-weight language model on a single server with four NVIDIA A100 40 GB GPUs. Full fine-tuning runs out of memory, and the team wants to train on their clinical notes dataset while keeping GPU memory within the available budget. Which TWO techniques should they apply to reduce memory consumption during fine-tuning? (Choose two.)
⚠ Common exam trap
The trap here is assuming any low-precision or format-conversion step reduces training memory, when only the optimizer-state and activation reductions actually free GPU memory during fine-tuning.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Gradient checkpointing
Parameter-efficient tuning with LoRA removes the need to store optimizer state for all base parameters, and gradient checkpointing cuts activation memory by recomputing intermediates during the backward pass. Together they target the two dominant memory consumers in fine-tuning, making a 7B model trainable on four 40 GB GPUs without changing the base weights.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Converting the model to ONNX format before training
Why it's wrong here
ONNX is an interchange format aimed at portable inference across runtimes and hardware. It does not provide a training-time memory optimization, and exporting a model to ONNX is a post-training step. Using it here would not reduce the optimizer and activation memory consumed during fine-tuning, so it cannot address the out-of-memory failure the team is facing.
- ✗
Quantization-aware training of the base weights at 4-bit
Why it's wrong here
Quantization-aware training simulates low-precision arithmetic so a model can later run efficiently in inference, and it is typically applied to the forward path. It does not by itself eliminate the optimizer state and gradient memory that full fine-tuning of a 7B model requires, so it would not resolve the out-of-memory condition described during training on this hardware.
- ✓
Gradient checkpointing
Why this is correct
Gradient checkpointing discards most intermediate activations during the forward pass and recomputes them during backpropagation. This trades additional compute for a large reduction in activation memory, which is a major consumer at long sequence lengths. Combined with parameter-efficient tuning, it helps fit the fine-tuning job into the four A100 40 GB GPUs available to this team.
- ✓
LoRA (Low-Rank Adaptation)
Why this is correct
LoRA freezes the pretrained weights and injects small trainable low-rank matrices into selected layers, so gradients and optimizer states are only maintained for those adapters. That removes the largest memory consumers of full fine-tuning, the optimizer moments for billions of parameters, and lets a 7B model train within the memory of four 40 GB GPUs while preserving base model quality.
- ✗
Increasing the global batch size
Why it's wrong here
Raising the global batch size increases the number of activations and gradients held at once, which raises memory pressure rather than lowering it. The team is already exceeding GPU memory, so a larger batch would make the out-of-memory condition worse. Batch size is a throughput and convergence knob here, not a memory-reduction technique.
About these practice questions
This AI0-001 question is part of Courseiva's 962-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official CompTIA exam blueprint
This AI0-001 practice question is part of Courseiva's free CompTIA certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the AI0-001 exam.