Courseiva

AI0-001 AI Infrastructure and Technologies Practice Question

An AI team is optimizing a convolutional neural network (CNN) for inference on a mobile device. The model has many layers and uses 32-bit floating-point weights. They need to reduce the model size and latency without significant accuracy loss. Which technique should they apply?

⚠ Common exam trap

The trap here is assuming that any model compression technique will equally reduce latency, when in fact quantization specifically addresses precision and hardware acceleration.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Quantization

Quantization is the most effective technique to reduce model size and latency on mobile devices by converting 32-bit floating-point weights to lower precision, such as 8-bit integers. This leverages hardware acceleration for integer operations and reduces memory bandwidth. Pruning and knowledge distillation can also help but are not as directly targeted at precision reduction.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    Pruning

    Why it's wrong here

    Pruning removes redundant weights or neurons from the network, which can reduce model size and computation. However, it often requires fine-tuning to recover accuracy and may not yield the same level of size reduction as quantization. It is effective but not the most direct method for reducing precision and latency on mobile devices.

  • ✗

    Knowledge distillation

    Why it's wrong here

    Knowledge distillation trains a smaller student model to mimic a larger teacher model. While it can reduce model size, it requires training a new model and may not directly leverage the existing architecture. It is more complex and time-consuming than quantization for the goal of reducing precision and latency.

  • ✓

    Quantization

    Why this is correct

    Quantization reduces the precision of the model's weights and activations, typically from 32-bit floating-point to 8-bit integers. This significantly decreases model size and speeds up inference on mobile hardware that supports integer operations. It can be done post-training or with quantization-aware training to minimize accuracy loss.

  • ✗

    Data augmentation

    Why it's wrong here

    Data augmentation increases the diversity of the training set by applying transformations to input data. It helps improve model generalization but does not reduce model size or inference latency. It is irrelevant to the deployment constraints described in the scenario.

About these practice questions

Courseiva writes every AI0-001 question from scratch — 962 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official CompTIA exam blueprint

This AI0-001 practice question is part of Courseiva's free CompTIA certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the AI0-001 exam.