A company is using Amazon SageMaker JumpStart to deploy a pre-trained text generation model. After deployment, the model produces slow inference responses. Which action is most likely to improve inference latency?
Trap 1: Quantize the model weights to FP16 or INT8.
Quantization can reduce latency but may also reduce accuracy. It is not always the most straightforward fix.
Trap 2: Fine-tune the model on a smaller dataset.
Fine-tuning does not affect inference speed.
Trap 3: Increase the batch size for inference requests.
Larger batch sizes improve throughput but may increase latency for individual requests.
- A
Quantize the model weights to FP16 or INT8.
Why it fails: Quantization can reduce latency but may also reduce accuracy. It is not always the most straightforward fix.
- B
Deploy the model on a more powerful instance type with higher GPU memory.
Inference latency for a deployed text generation model is bound by compute throughput and GPU memory bandwidth. Moving to an instance type with more GPU memory and compute raises tokens-per-second, directly reducing the slow response times described.
- C
Fine-tune the model on a smaller dataset.
Why it fails: Fine-tuning does not affect inference speed.
- D
Increase the batch size for inference requests.
Why it fails: Larger batch sizes improve throughput but may increase latency for individual requests.