AI0-001 AI Infrastructure and Technologies Practice Question
A computer vision team trains a convolutional neural network for manufacturing defect detection on a workstation with an NVIDIA RTX A6000 GPU. They want to reduce training time by increasing throughput without changing model architecture or batch size. Which action should they take?
⚠ Common exam trap
The trap here is assuming that any change to precision degrades model accuracy, when mixed-precision keeps FP32 accumulation specifically to preserve numerical stability.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Enable mixed-precision training using NVIDIA Tensor Cores with FP16 compute and FP32 accumulation.
Mixed-precision training exploits Tensor Cores to compute in FP16 while accumulating in FP32, delivering substantial throughput gains on modern NVIDIA GPUs with minimal accuracy impact. Because the architecture and batch size remain unchanged, it fits the constraint precisely. Other listed actions either alter training semantics or move work to slower hardware, so they do not satisfy the goal.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Increase the batch size to the maximum that fits in GPU memory and keep the learning rate unchanged.
Why it's wrong here
Increasing batch size changes the optimization dynamics and often requires learning-rate scaling to preserve convergence quality; the scenario explicitly says batch size must not change. It also does not inherently increase hardware throughput per sample. This action alters training behavior rather than accelerating the existing configuration.
- ✗
Shard the dataset across multiple CPU cores using NumPy and disable GPU acceleration for the convolution layers.
Why it's wrong here
Disabling GPU acceleration for convolutions moves the most compute-intensive operations to CPUs, which are far slower for dense matrix math than the RTX A6000. CPU-side data sharding may improve input pipeline throughput but cannot compensate for losing GPU convolution performance. This would increase, not decrease, training time.
- ✓
Enable mixed-precision training using NVIDIA Tensor Cores with FP16 compute and FP32 accumulation.
Why this is correct
Mixed-precision training uses Tensor Cores to perform matrix multiplications in FP16 while accumulating in FP32, roughly doubling throughput on Ampere-class GPUs without altering the model architecture or batch size. It preserves numerical stability because accumulation stays in FP32. This directly reduces training time for the existing CNN on the RTX A6000.
- ✗
Convert the trained model to TensorFlow Lite and redeploy it for training on the GPU.
Why it's wrong here
TensorFlow Lite is an inference runtime optimized for mobile and edge deployment, not a training accelerator for workstation GPUs. Converting to TFLite would not speed up training and could strip operations needed for backpropagation. It addresses a deployment concern, not the training-throughput problem described.
About these practice questions
This AI0-001 question is part of Courseiva's 962-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official CompTIA exam blueprint
This AI0-001 practice question is part of Courseiva's free CompTIA certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the AI0-001 exam.