A team is deploying a BERT-based question-answering model using a REST API endpoint with gRPC for internal microservices. They notice high latency for small payloads. Which optimization is MOST likely to reduce latency?
Trap 1: Convert the model to ONNX and use ONNX Runtime
ONNX Runtime can speed up inference, but the latency issue is likely due to request overhead, not model compute.
Trap 2: Switch from gRPC to REST with HTTP/2
Both support HTTP/2; gRPC is binary and typically faster. The issue is likely overhead from stream establishment.
Trap 3: Use a larger instance type with more CPU
Compute power may not be the bottleneck; more CPU adds cost without addressing protocol overhead.
- A
Enable batching of multiple queries into a single request
Batching increases payload size per request, reducing per-query overhead and improving throughput/latency.
- B
Convert the model to ONNX and use ONNX Runtime
Why it fails: ONNX Runtime can speed up inference, but the latency issue is likely due to request overhead, not model compute.
- C
Switch from gRPC to REST with HTTP/2
Why it fails: Both support HTTP/2; gRPC is binary and typically faster. The issue is likely overhead from stream establishment.
- D
Use a larger instance type with more CPU
Why it fails: Compute power may not be the bottleneck; more CPU adds cost without addressing protocol overhead.