RAG chatbot for company documents, with sources
Knowledge Hub is a RAG chatbot over the documents a company already keeps in Drive, Confluence, SharePoint or Notion. Each answer links to the passage it came from.
Question
Notice period in the standard supplier contract
Answer
Thirty days' written notice, sent to the address in clause 2.[1] Contracts on the older template allow sixty days.[2]
Sources
- 1Supplier contract template 2024Clause 14.2 · Legal / Contracts
- 2Supplier contract template 2021Clause 11 · Legal / Archive
Both readable by the asker in SharePoint
How an answer is built
Retrieval first; the model writes only from passages the asker may read.
Sources and what gets indexed
| Source | What gets indexed |
|---|---|
| Google Drive | Docs, Sheets and PDFs, with the folder permissions |
| Confluence | Spaces and pages, with page restrictions |
| SharePoint and OneDrive | Sites and libraries, with Microsoft 365 permissions |
| Notion | Pages and databases shared with the integration |
| PDFs and scans | Text extracted with OCR on the server, including scanned contracts |
RAG vs fine-tuning for a company assistant
Fine-tuning changes how a model writes. RAG changes what it knows when the question arrives.
| Criterion | RAG | Fine-tuning |
|---|---|---|
| A document changes | The next sync picks it up | The model needs retraining |
| Citations | One per claim | No reliable way |
| Permissions | Filtered per asker | Visible to everyone |
| Fits | Policies, contracts, manuals, wikis | A fixed style or format |
Project stages: import, tune, launch
- Import
- Sources connected read-only and indexed, permissions included.
- Tune
- Real questions with known answers. It ships when it passes them.
- Launch
- In Slack, Teams or a web page. Unanswered questions show the gaps.
Answers only from what the asker may read
A document the asker cannot open never shapes the answer.
How the permission filter works
- Index
- Every indexed passage keeps the access list of its source document.
- Filter
- Each question carries the asker's identity from the company login, Google or Microsoft, and retrieval filters by it.
- Answer
- The model receives only the passages left after the filter. Access removed in the source leaves the index at the next sync.
Our own knowledge base, where agents file every meeting
The closest shipped work: every meeting filed as a note under its project, with the transcript as its source. LifeOS case study
Every meeting, filed under its project
Meetings filed under their project
Every action item, in writing
Action items captured in writing
One knowledge base, versioned in git
Notes in the knowledge base
Data handling: where the data goes and how ELASTO handles security.
Company knowledge: common questions
RAG or fine-tuning for a company assistant
RAG, in nearly every case. Company knowledge changes every week and has to be cited and filtered by permission, which fine-tuning cannot do. Fine-tuning helps only when the assistant must write in a fixed format, and it can sit on top of RAG.
Answers from Confluence, Google Drive or PDFs, with the source
Confluence, Google Drive, SharePoint, Notion and PDFs are read where they live, and scanned PDFs go through OCR on the server. Every answer cites the document and the section, with a link that opens it.
Documents some people are not allowed to see
The assistant filters every search by the asker's permissions in the source system before the model sees any text. A document someone cannot open in SharePoint or Drive cannot appear in their answer, and removed access leaves the index at the next sync.
Keeping answers current when documents change
Sources sync on a schedule or through each system's change notifications, so an edited policy replaces the old passages. Every answer shows when its document was last edited, and deleted documents drop out of the index.
Company data in ChatGPT vs a private assistant
Consumer chat apps can keep and use conversations under their own terms. A private assistant calls the model through a business API, where Anthropic does not train on the data without explicit permission, and the index sits on infrastructure agreed with the company. The full data path is on the AI agents page.
When the documents hold no answer
The assistant says the documents do not cover the question and lists the closest passages it found. It does not fill the gap from the model's general knowledge. Unanswered questions are collected, and they show which documents are missing.
One team's documents, mapped in the AI workshop
Which sources, which weekly questions, and the cost to build and run it.
- Studio
- Cluj-Napoca, Romania, EU