A developer wants to deploy a custom generative AI model using Azure Machine Learning. Which compute target should they choose for low-latency real-time inference?
Trap 1: Local deployment
Local deployment is not scalable for production.
Trap 2: Azure Batch
Azure Batch is for batch processing, not real-time.
Trap 3: Azure Functions
Azure Functions may have cold start latency.
- A
Local deployment
Why it fails: Local deployment is not scalable for production.
- B
Azure Batch
Why it fails: Azure Batch is for batch processing, not real-time.
- C
Azure Functions
Why it fails: Azure Functions may have cold start latency.
- D
Azure Kubernetes Service (AKS)
Azure Kubernetes Service provides managed Kubernetes clusters that Azure Machine Learning can attach as an inference compute target, supporting real-time endpoints with autoscaling and GPU node pools. This satisfies the low-latency requirement because AKS keeps model replicas continuously loaded and ready, unlike batch or serverless compute that incurs cold-start delays.