AI0-001 AI Infrastructure and Technologies Practice Question
An ML team deploys a model on edge devices using INT8 quantization. They notice a significant drop in accuracy on a subset of classes. Which technique should they apply to recover accuracy without increasing model size?
⚠ Common exam trap
CompTIA AI often tests the misconception that post-training quantization is always lossless, leading candidates to overlook the need for QAT when accuracy drops on specific classes due to uneven weight distributions.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Apply quantization-aware training (QAT)
Quantization-aware training (QAT) simulates INT8 quantization effects during the forward pass of training, allowing the model to learn weights and activations that are more robust to the lower precision. This recovers accuracy lost during post-training quantization without increasing the model's size, as the architecture and number of parameters remain unchanged.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Use pruning to remove less important weights
Why it's wrong here
Pruning removes weights, reducing capacity further; it does not restore the precision lost on the affected classes, so accuracy on those subsets stays degraded. It is tempting because pruning is a real compression technique, and it would be correct when the goal is shrinking an already-accurate model rather than recovering quantisation accuracy.
- ✗
Increase the model architecture size
Why it's wrong here
Enlarging the architecture increases parameter count and model size, contradicting the constraint, and does not target quantisation error. Bigger models help when underfitting. Quantisation-aware training simulates INT8 rounding during training so weights adapt, recovering accuracy at the same size.
- ✗
Switch to FP16 quantization
Why it's wrong here
FP16 doubles the bit-width of INT8, so the model's memory footprint and size grow, breaching the no-increase constraint. It is tempting because FP16 is a genuine post-training quantization format that recovers accuracy, and it would be the right choice when the target hardware supports half-precision and size is not restricted.
- ✓
Apply quantization-aware training (QAT)
Why this is correct
Quantization-aware training simulates INT8 rounding during the forward pass while keeping weights in higher precision for gradient updates, letting the model learn to compensate for that error. This recovers accuracy on the affected classes while the deployed artefact stays INT8, so model size is unchanged.
About these practice questions
Courseiva writes every AI0-001 question from scratch — 962 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →
JA
Written by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
This AI0-001 practice question is part of Courseiva's free CompTIA certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the AI0-001 exam.