A data scientist notices that a text generation model deployed on Vertex AI returns repetitive outputs after a few turns in a chat application. What is the most likely cause and the best parameter adjustment?
Trap 1: The max_output_tokens is too low; increase it to allow more diverse…
Max tokens controls length, not repetition.
Trap 2: The model is overfitted; switch to a smaller model.
Overfitting is unlikely in pre-trained models; repetition is a decoding issue.
Trap 3: The temperature is too low; increase temperature to add randomness.
Low temperature makes output more deterministic, increasing repetition.
- A
The max_output_tokens is too low; increase it to allow more diverse output.
Why it fails: Max tokens controls length, not repetition.
- B
The top_p value is too high; reduce top_p to limit token sampling.
A high top_p widens the nucleus of sampled tokens, so low-probability tokens get selected, producing loops and repetition. Reducing top_p restricts sampling to the most probable tokens, tightening output distribution and breaking the repetitive cycle described.
- C
The model is overfitted; switch to a smaller model.
Why it fails: Overfitting is unlikely in pre-trained models; repetition is a decoding issue.
- D
The temperature is too low; increase temperature to add randomness.
Why it fails: Low temperature makes output more deterministic, increasing repetition.