A company wants to build a customer service chatbot that answers questions about their internal policy documents. The documents are updated monthly, and the team cannot afford to retrain a model each time. Which approach is MOST appropriate?
Trap 1: Train a custom model from scratch on the policy documents each month
Training from scratch each month demands labelled data, compute and ML expertise, and cannot keep pace with monthly policy edits. It is tempting because bespoke training suits stable domains with abundant proprietary data, but here retrieval-augmented generation over an updated index avoids retraining entirely.
Trap 2: Fine-tune a base LLM on the policy documents monthly
Fine-tuning bakes policy text into weights, so each monthly update still requires a fresh training run and redeployment, exactly the recurring cost the team cannot afford. It is tempting because fine-tuning adapts tone and task behaviour, which suits stable formats rather than frequently changing factual content.
Trap 3: Use a larger foundation model with a longer context window and…
Pasting every policy document into each prompt exhausts the context window, inflates latency and cost per query, and cannot scale as the corpus grows. It is tempting because long-context models handle small, static document sets well, but retrieval of only relevant passages satisfies the monthly-update requirement.
- A
Train a custom model from scratch on the policy documents each month
Why it fails: Training from scratch each month demands labelled data, compute and ML expertise, and cannot keep pace with monthly policy edits. It is tempting because bespoke training suits stable domains with abundant proprietary data, but here retrieval-augmented generation over an updated index avoids retraining entirely.
- B
Use Retrieval-Augmented Generation (RAG) with the policy documents indexed in a vector store
RAG retrieves relevant passages from the indexed policy documents at query time and supplies them as context to the language model, so answers reflect current content. Monthly document updates only require re-indexing the vector store, avoiding the cost and delay of retraining the model.
- C
Fine-tune a base LLM on the policy documents monthly
Why it fails: Fine-tuning bakes policy text into weights, so each monthly update still requires a fresh training run and redeployment, exactly the recurring cost the team cannot afford. It is tempting because fine-tuning adapts tone and task behaviour, which suits stable formats rather than frequently changing factual content.
- D
Use a larger foundation model with a longer context window and paste all documents into each prompt
Why it fails: Pasting every policy document into each prompt exhausts the context window, inflates latency and cost per query, and cannot scale as the corpus grows. It is tempting because long-context models handle small, static document sets well, but retrieval of only relevant passages satisfies the monthly-update requirement.