AI0-001 AI Infrastructure and Technologies Practice Question
An organization is deploying a large language model on-premises for compliance reasons. They need to serve inference requests with low latency. Which architecture should they use?
⚠ Common exam trap
CompTIA often tests the misconception that a single powerful server is sufficient for low-latency inference, but the trap is that it ignores the need for horizontal scalability and fault tolerance, which are critical for production workloads.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Containerize the model and deploy it on a Kubernetes cluster with autoscaling
Containerizing the model and deploying it on a Kubernetes cluster with autoscaling is the correct architecture because it provides horizontal scaling, low-latency inference through load-balanced pods, and supports on-premises deployment for compliance. Kubernetes can automatically scale replicas based on CPU/memory utilization or custom metrics (e.g., request queue depth), ensuring consistent response times under varying load.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Use a batch processing system like Apache Spark
Why it's wrong here
Spark processes data in scheduled batch jobs, so results arrive after job completion rather than per request, failing the low-latency requirement. Batch pipelines suit offline scoring of large datasets. On-premises inference needs a serving runtime such as NVIDIA Triton or vLLM.
- ✓
Containerize the model and deploy it on a Kubernetes cluster with autoscaling
Why this is correct
Containerising the model on Kubernetes keeps inference inside the organisation's own infrastructure, satisfying the on-premises compliance constraint. Autoscaling horizontally scales replicas to match request load, and local GPU nodes serve requests without network round trips to a public endpoint, delivering the low latency required.
- ✗
Use a serverless function like AWS Lambda
Why it's wrong here
Serverless functions impose cold-start delays and execution time limits, and cannot host a large model's weights on-premises for compliance; the correct architecture keeps the model resident on dedicated GPU hardware. It is tempting because serverless suits short, event-driven tasks, not sustained low-latency inference.
- ✗
Deploy the model as a REST API on a single powerful server
Why it's wrong here
A single server cannot scale horizontally, so concurrent inference requests queue and latency rises under load; the correct architecture distributes inference across multiple GPU nodes behind a load balancer. It is tempting because a REST API on one host is the simplest way to expose a model for low-volume or development use.
About these practice questions
One of 962 original AI0-001 practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →
JA
Written by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
This AI0-001 practice question is part of Courseiva's free CompTIA certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the AI0-001 exam.