E-time | the software company
What is RAFT (Retrieval-Augmented Fine-Tuning) and how does it work
Retrieval-Augmented Fine-Tuning (RAFT) is a training technique designed to adapt Large Language Models (LLMs) to highly specialized domains. The method involves training the model through questions to which both the original question and a set of retrieved documents are associated.
Compared to traditional training methods, RAFT teaches the model to interpret the context, identify and extract relevant information from the relevant documents, defined as oracle, and to actively recognize and ignore irrelevant or potentially misleading ones, called distractors.
How RAFT combines RAG and Fine-Tuning
RAFT combines the flexibility of Retrieval-Augmented Generation (RAG) with the specific adaptation offered by Supervised Fine-Tuning. In traditional RAG, external documents are made available to the model exclusively during the inference phase, without the model having previously been trained to evaluate and select them.
Traditional Fine-Tuning (SFT), on the other hand, focuses mainly on learning information and format, without preparing the model to use a real document context.
RAFT overcomes this distinction by directly inserting the retrieved documents into the Fine-Tuning dataset, including both the correct ones and the distractor documents. In this way, the LLM learns not only the structure and language characteristic of the reference domain, but also the ability to dynamically evaluate the context and identify and extract the correct information in real time.
What are the benefits of RAFT for LLMs
The main advantages of RAFT concern greater accuracy of responses and significant robustness towards inaccurate or noisy retrievals. Through training with distractor documents and Chain-of-Thought reasoning, the model learns not to be influenced by irrelevant content retrieved by the retrieval system.
Furthermore, RAFT allows relatively smaller models, such as LLaMA-7B, to achieve higher performance compared to zero-shot RAG and traditional domain Fine-Tuning, both with and without RAG. In different application contexts, these models can achieve better results even compared to larger models, such as GPT-3.5 combined with RAG.
Limitations and critical aspects of Retrieval-Augmented Fine-Tuning
Despite its advantages, Retrieval-Augmented Fine-Tuning presents some limitations both from an operational and methodological point of view. First of all, the effectiveness of the technique may vary depending on the nature of the task.
In tests conducted on datasets characterized by synthetic or binary answers, RAFT does not show particularly significant improvements compared to the combination of traditional domain Fine-Tuning and RAG. This highlights how the benefits of the method may be more limited in cases where an articulated reasoning process is not required.
A second aspect concerns the data preparation phase. The construction of the dataset in fact requires the accurate generation of answers explained step by step through Chain-of-Thought, as well as careful selection of oracle and distractor documents. This involves a significant investment of time and resources in dataset preparation.
Finally, the model’s performance may be sensitive to the configuration of the dataset hyperparameters, including the percentage P and the number of distractor documents used during training.
How to Optimize RAG with Margot AI
In enterprise customer service, RAG technology can be integrated into solutions such as Margot, E-time’s AI Agent, integrated with Rexpondo, the ticketing and IT Service Management platform.
Thanks to the dynamic retrieval of information from the company knowledge base, Margot can provide accurate and up-to-date answers, automatically classify requests, and support operators and users across different channels.





