AI0-001 AI Infrastructure and Technologies Practice Question
An AI team is optimizing a convolutional neural network (CNN) for inference on a mobile device. The model has many layers and uses 32-bit floating-point weights. They need to reduce the model size and latency without significant accuracy loss. Which technique should they apply?
⚠ Common exam trap
The trap here is assuming that any model compression technique will equally reduce latency, when in fact quantization specifically addresses precision and hardware acceleration.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Quantization
Quantization is the most effective technique to reduce model size and latency on mobile devices by converting 32-bit floating-point weights to lower precision, such as 8-bit integers. This leverages hardware acceleration for integer operations and reduces memory bandwidth. Pruning and knowledge distillation can also help but are not as directly targeted at precision reduction.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Pruning
Why it's wrong here
Pruning removes redundant weights or neurons from the network, which can reduce model size and computation. However, it often requires fine-tuning to recover accuracy and may not yield the same level of size reduction as quantization. It is effective but not the most direct method for reducing precision and latency on mobile devices.
- ✗
Knowledge distillation
Why it's wrong here
Knowledge distillation trains a smaller student model to mimic a larger teacher model. While it can reduce model size, it requires training a new model and may not directly leverage the existing architecture. It is more complex and time-consuming than quantization for the goal of reducing precision and latency.
- ✓
Quantization
Why this is correct
Quantization reduces the precision of the model's weights and activations, typically from 32-bit floating-point to 8-bit integers. This significantly decreases model size and speeds up inference on mobile hardware that supports integer operations. It can be done post-training or with quantization-aware training to minimize accuracy loss.
- ✗
Data augmentation
Why it's wrong here
Data augmentation increases the diversity of the training set by applying transformations to input data. It helps improve model generalization but does not reduce model size or inference latency. It is irrelevant to the deployment constraints described in the scenario.
About these practice questions
Courseiva writes every AI0-001 question from scratch — 962 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official CompTIA exam blueprint
This AI0-001 practice question is part of Courseiva's free CompTIA certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the AI0-001 exam.