AI0-001 AI Infrastructure and Technologies Practice Question
A media company uses a large language model (LLM) to generate article summaries. They want to reduce inference costs and latency without significantly degrading summary quality. The LLM is currently served at full precision. Which optimization technique is most appropriate?
⚠ Common exam trap
Many exam-takers confuse throughput optimizations like batching with latency and cost reductions, or assuming that caching will solve the problem when inputs are largely unique.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Apply quantization to convert the model weights to lower precision (e.g., INT8).
Quantization converts model weights to lower precision, reducing memory bandwidth and compute requirements, which lowers latency and cost. For LLMs, INT8 quantization often preserves summary quality well. Increasing batch size helps throughput but not per-request latency, a larger model worsens cost, and caching is ineffective for unique articles. Thus, quantization is the most appropriate optimization.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Cache all generated summaries to avoid repeated inference.
Why it's wrong here
Caching can help for repeated identical requests, but article summaries are often unique per article, so cache hit rates would be low. This does not reduce the cost of generating new summaries. Caching is a complementary optimization but not a primary technique for reducing model inference costs and latency across diverse inputs. It does not address the computational expense of the LLM itself.
- ✓
Apply quantization to convert the model weights to lower precision (e.g., INT8).
Why this is correct
Quantization reduces the precision of model weights and activations, typically from FP32 to INT8, which decreases memory usage and speeds up inference on compatible hardware. For LLMs, post-training quantization can significantly lower latency and cost with minimal quality loss, especially when using techniques like GPTQ or AWQ. This directly addresses the media company's need to optimize inference without major degradation.
- ✗
Use a larger model with more parameters to improve summary quality.
Why it's wrong here
A larger model would increase inference costs and latency, opposite to the stated goal. While it might improve quality, the company wants to reduce costs and latency without significant quality degradation. Scaling up is not an optimization technique for efficiency; it is a trade-off that prioritizes quality at the expense of resources. This option fails to meet the primary requirements.
- ✗
Increase the batch size for inference requests.
Why it's wrong here
Increasing batch size can improve throughput but does not reduce the computational cost per token or the memory footprint of the model. It may also increase latency for individual requests if not managed properly. While batching is useful for throughput-oriented workloads, it does not address the goal of reducing inference costs and latency for a summarization service where response time matters.
Visual reference
About these practice questions
Courseiva writes every AI0-001 question from scratch — 962 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official CompTIA exam blueprint
This AI0-001 practice question is part of Courseiva's free CompTIA certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the AI0-001 exam.