Refer to the exhibit. An engineer is configuring a model serving endpoint. Based on the configuration, which outcome is expected during a sudden surge in traffic?
Exhibit
{
"model_name": "llama-3-8b",
"task": "chat",
"serving_endpoint": {
"min_cpu_cores": 4,
"max_cpu_cores": 16,
"auto_scaling": true
}
}Trap 1: The endpoint will crash due to lack of static resource allocation.
Auto-scaling is specifically designed to prevent crashes by dynamically adding capacity when load increases. The system monitors utilization and adjusts the number of cores based on the defined range, ensuring that the service remains available even during periods of high demand without requiring permanent over-provisioning of resources.
Trap 2: The endpoint will force requests to queue indefinitely until…
The primary purpose of auto-scaling is to increase throughput capacity, not to throttle requests. By scaling up resources, the system accommodates increased demand in real-time, preventing the need for request queuing that would otherwise lead to poor user experiences and latency issues in production generative AI deployments.
Trap 3: The configuration is invalid because it lacks a GPU definition.
While GPUs are often used for LLM inference, CPU-based serving is a valid configuration for many models depending on size and performance requirements. The provided JSON is syntactically valid for a CPU-backed endpoint, and the absence of a GPU does not make the configuration inherently invalid or dysfunctional.
- A
The endpoint will crash due to lack of static resource allocation.
Why it fails: Auto-scaling is specifically designed to prevent crashes by dynamically adding capacity when load increases. The system monitors utilization and adjusts the number of cores based on the defined range, ensuring that the service remains available even during periods of high demand without requiring permanent over-provisioning of resources.
- B
The endpoint will scale up to 16 CPU cores to handle the increased load.
With auto-scaling enabled, the model serving infrastructure monitors the incoming traffic and scales the compute resources up to the defined maximum of 16 CPU cores. This ensures that the application maintains low latency for user requests during traffic surges while staying within the predefined infrastructure cost boundaries.
- C
The endpoint will force requests to queue indefinitely until traffic subsides.
Why it fails: The primary purpose of auto-scaling is to increase throughput capacity, not to throttle requests. By scaling up resources, the system accommodates increased demand in real-time, preventing the need for request queuing that would otherwise lead to poor user experiences and latency issues in production generative AI deployments.
- D
The configuration is invalid because it lacks a GPU definition.
Why it fails: While GPUs are often used for LLM inference, CPU-based serving is a valid configuration for many models depending on size and performance requirements. The provided JSON is syntactically valid for a CPU-backed endpoint, and the absence of a GPU does not make the configuration inherently invalid or dysfunctional.