An organization is deploying a large language model on-premises for compliance reasons. They need to serve inference requests with low latency. Which architecture should they use?
Containerising the model on Kubernetes keeps inference inside the organisation's own infrastructure, satisfying the on-premises compliance constraint. Autoscaling horizontally scales replicas to match request load, and local GPU nodes serve requests without network round trips to a public endpoint, delivering the low latency required.
Why this answer
Containerizing the model and deploying it on a Kubernetes cluster with autoscaling is the correct architecture because it provides horizontal scaling, low-latency inference through load-balanced pods, and supports on-premises deployment for compliance. Kubernetes can automatically scale replicas based on CPU/memory utilization or custom metrics (e.g., request queue depth), ensuring consistent response times under varying load.
Exam trap
CompTIA often tests the misconception that a single powerful server is sufficient for low-latency inference, but the trap is that it ignores the need for horizontal scalability and fault tolerance, which are critical for production workloads.
How to eliminate wrong answers
Option A is wrong because batch processing systems like Apache Spark are designed for large-scale data processing jobs, not real-time inference; they introduce high latency due to job scheduling and data shuffling, making them unsuitable for serving low-latency requests. Option C is wrong because serverless functions like AWS Lambda are typically cloud-only and may not support on-premises deployment; they also have cold-start latency and execution time limits that conflict with low-latency inference requirements. Option D is wrong because deploying on a single powerful server creates a single point of failure and cannot scale horizontally to handle traffic spikes, leading to increased latency under load.