Blog · Technology
What is RAG and how an AI knowledge base works
If you ask a general-purpose language model about your office opening hours, your shop's return policy or your company's complaints procedure, you will get an answer that sounds convincing – and is most likely made up. The model has never seen your documents, so it fills the gaps with what is "usually" true. RAG is the most popular way to change that. Below we explain what it is, how it works and when it makes sense in a small or mid-size company.
What is RAG
RAG stands for retrieval augmented generation: generating answers supported by search. The principle is simple: before the model answers a question, the system searches your documents for passages related to that question and hands them to the model together with an instruction such as "answer only on the basis of these materials".
So the model does not have to "know" your company. It only needs to get the right pages from the binder at the right moment. It is a bit like the difference between an employee who answers from memory and one who checks the terms before answering and tells you which clause says so.
What problem it solves
- The model does not know your documents. Large models were trained on public data. Your internal price list, framework agreements, work instructions or the history of arrangements with a client never made it in.
- Hallucinations. When a model does not know, it often does not say "I don't know" but generates a plausible-sounding answer. In customer service this is the worst possible error, because it is hard to catch.
- Outdated knowledge. A model's knowledge ends on the day its training ended. Your price list changed last week – the model does not know that, but RAG does, as long as the document has been replaced.
- No way to verify. An answer without a source has to be taken on faith. An answer with a link to a specific document can be verified in a few seconds.
How RAG works – step by step
For a business owner it is enough to know that the system "searches, then answers". For a technical person, let's lay it out in more detail. The process has two phases: preparing the knowledge base (once, and again whenever documents change) and handling a question (every time someone asks).
- Collecting documents. PDFs, Word files, spreadsheets, web pages, exports from systems, a Q&A database. At this stage duplicates and outdated versions are removed.
- Splitting into chunks. Documents are divided into smaller parts: paragraphs, sections, clauses of the terms. The model then receives a few relevant chunks rather than a whole hundred-page file.
- Embeddings. Each chunk is converted into a vector of numbers that describes its meaning. Chunks with similar content have similar vectors even if they use different words – "returning goods" and "sending back an order" sit close together.
- Vector database. The vectors, together with the chunk text and source information (document name, page, date), go into a database that can quickly find the closest items by meaning.
- Retrieval. When a user asks a question, it is also converted into a vector, and the system picks the few best-matching chunks. This is often combined with classic keyword search, because contract numbers, product codes or surnames are caught less well by semantic search.
- An answer with a source citation. The model gets the question and the retrieved chunks, formulates an answer and points to the document the information comes from. If the materials do not contain the answer, a well-configured system says so plainly and suggests contacting a person.
RAG, fine-tuning or pasting into ChatGPT
These are three different approaches that are often confused.
Pasting documents into ChatGPT. Works for one file and one question. It does not scale to hundreds of documents, every employee does it their own way, there is no control over who pasted what, and company data ends up in a tool the company has no agreement or processing rules with. From a GDPR point of view this is often the weakest link.
Fine-tuning. Further training of a model on your own data. It is good at teaching style, answer format or specific industry vocabulary, but poorly suited to storing facts. Every price list change means retraining, and the model still will not show where it got the information. It is also more expensive and slower to implement.
RAG. The knowledge stays in the documents and the model only uses it. Updating means replacing a file, answers have a source, and access can be restricted by permissions. For most business uses this is a sensible starting point; fine-tuning is sometimes a complement, rarely a replacement.
Typical uses in a small or mid-size company
- Knowledge base for the team. A new employee asks "how do I issue a corrective invoice for a foreign client" and gets an answer from the internal instruction instead of pulling a senior colleague away from work.
- Chatbot built on your price list and terms. A customer on your site asks about delivery cost, return deadline or warranty conditions, and the bot answers from current documents. We write more about such implementations on the page about chatbots for business.
- Customer service. A first line of support that answers repetitive questions and hands unusual cases to a person together with the context of the conversation.
- Procedures and instructions. Health and safety, quality procedures, equipment manuals – documents that exist but nobody reads, because it is hard to find anything in them.
- Contracts. Finding provisions in contracts with clients and suppliers: notice periods, contractual penalties, payment terms. Source citations and permissions matter especially here.
RAG is also often the foundation of more complex systems. An agent that is meant not only to answer but to act – create a ticket, check an order status – usually relies on a knowledge base built exactly this way. We describe this on the page about AI agents for business.
What answer quality depends on
The model itself matters less here than it seems. Whether the system answers accurately is decided mainly by what happens before it.
- Document quality. If three versions of your terms circulate in the company, the bot will quote all three. Contradictory, outdated and incomplete materials are the most common cause of bad answers – and the most common reason why the first part of an implementation is putting your knowledge in order.
- Chunking. Chunks that are too large blur relevance, chunks that are too small lose context. A price table cut in the middle of a row will give a wrong price. Splitting has to be matched to the structure of the documents, not set once for all of them.
- Updating. The knowledge base has to keep up with the company. Best of all, indexing starts automatically when a file changes in a designated folder or system, rather than depending on someone remembering to upload it.
- Permissions. An assistant for the sales team should not answer from HR documents. Permissions have to be carried over to the retrieval layer – the system should search only what the person asking has access to.
- Tests on real questions. A set of a few dozen real questions from customers or employees, checked after every change, says more about quality than any demo.
Local RAG on your own hardware
In many industries the documents that would go into the knowledge base are sensitive: medical records, case files at a law firm, clients' financial data, trade secrets. Sending them to an external cloud is then sometimes unacceptable – for legal or contractual reasons, or simply out of caution.
The solution is local RAG: the vector database and the language model run on hardware in your company, in a closed network. Documents, questions and answers never leave the building. At Axisway we install such systems on Mac Mini computers or GPU servers, with local LLMs chosen to match the scale and type of data. Local models are somewhat weaker than the largest cloud models, but for the task "answer on the basis of these passages" the difference is usually small – because the knowledge is supplied by retrieval, not by the model's memory.
If you do not process sensitive data, a cloud model with a proper data processing agreement and GDPR-compliant settings is usually cheaper and sufficient. The decision is worth making based on the type of documents, not on fashion.
How to get started
- Choose one area. Best one where questions repeat and the answers are already written down somewhere: customer service, onboarding new employees, support for sales reps. If you do not know where to start, an AI audit will help, as it ranks processes by potential gain.
- Collect documents and questions. A few dozen source documents and a list of real questions are enough to check whether the approach works.
- Build a first version and test. The first working version usually takes us 7–14 days. Then comes tuning: document splitting, answers for missing topics, permissions.
- Expand gradually. More departments, channels and integrations only once the first area works well.
We write more about the whole process in the guide to AI implementation in a company, and about the budget in the article how much does AI implementation cost. For qualifying projects, a PARP (Polish Agency for Enterprise Development) grant can cover part of the costs, up to 75%.
Summary
RAG is a way to make a language model answer from your documents rather than its own guesses – and show where it got the answer. Technically it is indexing, embeddings, a vector database and retrieval before every answer. In practice, success depends on order in your documents, sensible chunking, updates and permissions. With sensitive data the whole thing can run locally, without sending anything outside.
FAQ
Questions about RAG
Does RAG completely eliminate hallucinations?
Not completely, but it reduces them significantly. The model can still misread a passage or combine two that do not fit together. That is why source citations, an instruction to answer only from the materials and regular tests on real questions matter.
What document formats can be connected?
Most often PDF, Word, spreadsheets, web pages and exports from company systems. Scans require text recognition, and complex tables and diagrams need extra processing. The better organized the document, the better the answers.
Will my documents be used to train the model?
In RAG documents are not used for training – the model only reads selected passages at the moment of answering. With cloud models the processing terms are set by the agreement with the provider, and with a local model the data never leaves your network.
How many documents do I need for this to make sense?
There is no lower limit – a useful chatbot can run on a few documents if they describe your offer and rules well. Completeness and freshness matter more than the number. With just a few files a simpler solution is often enough, and full RAG shows its advantage with dozens and hundreds of documents.
AI knowledge base
Let's check whether your documents are ready
In a free audit we go through your documents and the questions that come up most often, and tell you plainly: whether RAG makes sense here, in the cloud or locally, and which area to start with.
- info@axisway.com
- +48 516 068 354
- ul. Szewska 8, 50-122 Wrocław