A healthcare company is using a fine-tuned version of PaLM 2 on Vertex AI to generate clinical notes from doctor-patient conversations. The model was fine-tuned on a dataset of 10,000 de-identified transcripts and corresponding notes. During testing, the generated notes are grammatically correct and well-structured, but they often contain subtle inaccuracies: for example, they might mention a medication that was not discussed, or omit a key symptom. The team has already tried increasing the training epochs and adjusting learning rates, with minimal improvement. They need a solution that can be implemented quickly to improve factual accuracy without retraining the entire model. The team has access to a large archive of verified clinical notes and a small set of recent conversation-to-note pairs that have been manually reviewed and corrected. The inference pipeline currently uses a single call to the model with the conversation transcript as input. What should the team do?
Trap 1: Decrease the temperature to 0.1 to reduce randomness and force the…
Low temperature reduces randomness but does not add factual grounding; the model may still omit or invent details.
Trap 2: Use prompt engineering to instruct the model to only include…
Prompt engineering helps but may not overcome the model's training biases; it can still hallucinate.
Trap 3: Add a human-in-the-loop step to review and correct every generated…
Human review is slow and doesn't reduce the model's tendency to produce inaccuracies; it's a workaround not a fix.
- A
Implement retrieval-augmented generation (RAG) by retrieving similar verified notes from the archive and providing them as context in the prompt.
RAG injects retrieved verified notes as grounding context, so the model conditions on factual exemplars rather than relying solely on fine-tuned weights. This directly targets the hallucinated medications and omitted symptoms, needs no retraining, and can be deployed quickly in the existing single-call pipeline.
- B
Decrease the temperature to 0.1 to reduce randomness and force the model to stick to the input.
Why it fails: Low temperature reduces randomness but does not add factual grounding; the model may still omit or invent details.
- C
Use prompt engineering to instruct the model to only include information explicitly mentioned in the conversation.
Why it fails: Prompt engineering helps but may not overcome the model's training biases; it can still hallucinate.
- D
Add a human-in-the-loop step to review and correct every generated note before use.
Why it fails: Human review is slow and doesn't reduce the model's tendency to produce inaccuracies; it's a workaround not a fix.