AI0-001 AI Infrastructure and Technologies Practice Question
A team wants to deploy a large language model on edge devices with limited memory and compute. They need to reduce model size by at least 50% while preserving accuracy. Which combination of techniques is most effective?
⚠ Common exam trap
The exam often tests the misconception that a single technique (like distillation or FP16) is sufficient for aggressive size reduction, when in reality, combining complementary compression methods (quantization and pruning) is necessary to meet both the 50% size reduction and accuracy preservation requirements on edge devices.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Apply INT8 quantization and weight pruning
INT8 quantization reduces the precision of weights and activations from 32-bit to 8-bit, cutting memory usage by approximately 75% for those tensors, while weight pruning removes redundant connections, often achieving over 50% size reduction with minimal accuracy loss when combined. Together, they directly address the constraints of edge devices by shrinking the model footprint and computational requirements without requiring a complete architecture redesign.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✓
Apply INT8 quantization and weight pruning
Why this is correct
INT8 quantization shrinks weights from 32-bit to 8-bit, cutting size roughly 75%, while weight pruning removes redundant connections. Combined, they exceed the 50% reduction target with minimal accuracy loss, fitting edge memory and compute limits.
- ✗
Distill the model into a smaller architecture without quantization or pruning
Why it's wrong here
Distillation alone yields a smaller architecture but typically falls short of a guaranteed 50% reduction, and omitting pruning and quantization forfeits further compression. Distillation suits producing a compact student model when accuracy retention matters more than a hard size target.
- ✗
Use FP32 precision and increase batch size
Why it's wrong here
FP32 is full precision, so weights stay at four bytes each and no size reduction occurs; larger batches also raise activation memory. FP32 training suits servers with ample memory, where numerical fidelity matters more than footprint, not edge deployment.
- ✗
Use FP16 quantization and add more layers
Why it's wrong here
FP16 halves weight storage but adding layers increases parameter count and memory, so the 50% reduction is not achieved. FP16 alone suits GPU inference where memory is adequate; adding layers targets accuracy, not compression, and works against the stated constraint.
About these practice questions
Courseiva writes every AI0-001 question from scratch — 962 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →
JA
Written by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
This AI0-001 practice question is part of Courseiva's free CompTIA certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the AI0-001 exam.