Main content

RAG chatbot for company documents, with sources

Knowledge Hub is a RAG chatbot over the documents a company already keeps in Drive, Confluence, SharePoint or Notion. Each answer links to the passage it came from.

Question

Notice period in the standard supplier contract

Answer

Thirty days' written notice, sent to the address in clause 2.[1] Contracts on the older template allow sixty days.[2]

Sources

  1. 1Supplier contract template 2024Clause 14.2 · Legal / Contracts
  2. 2Supplier contract template 2021Clause 11 · Legal / Archive

Both readable by the asker in SharePoint

Sample question, fictional documents

How an answer is built

Retrieval first; the model writes only from passages the asker may read.

FIG. 1Retrieval, then a cited answer
RAG pipeline: company documents sync into an index of passages with their access lists; a question carries the asker's identity to retrieval, which keeps only readable passages; the model writes from those passages; the answer cites each one.Permissions are checked before the model sees any text. When no passage answers the question, the assistant says so instead of guessing.DocumentsWhere they liveIndexPassages, access listsRetrievalReadable passagesQuestionWho asksModelPassages onlyCited answerSource per claim

Sources and what gets indexed

SourceWhat gets indexed
Google DriveDocs, Sheets and PDFs, with the folder permissions
ConfluenceSpaces and pages, with page restrictions
SharePoint and OneDriveSites and libraries, with Microsoft 365 permissions
NotionPages and databases shared with the integration
PDFs and scansText extracted with OCR on the server, including scanned contracts

RAG vs fine-tuning for a company assistant

Fine-tuning changes how a model writes. RAG changes what it knows when the question arrives.

RAG vs fine-tuning for a company assistant
CriterionRAGFine-tuning
A document changesThe next sync picks it upThe model needs retraining
CitationsOne per claimNo reliable way
PermissionsFiltered per askerVisible to everyone
FitsPolicies, contracts, manuals, wikisA fixed style or format

Project stages: import, tune, launch

Import
Sources connected read-only and indexed, permissions included.
Tune
Real questions with known answers. It ships when it passes them.
Launch
In Slack, Teams or a web page. Unanswered questions show the gaps.

Answers only from what the asker may read

A document the asker cannot open never shapes the answer.

FIG. 2The permission filterSchematic
The permission filter: every indexed passage keeps the access list of its source, every question carries the asker's company login, and the filter passes to the model only the passages the asker may read. The rest never reach the model.PassagesAccess listsAskerCompany loginPermission filterEvery questionModelReadable onlyFiltered outNever sent

How the permission filter works

Index
Every indexed passage keeps the access list of its source document.
Filter
Each question carries the asker's identity from the company login, Google or Microsoft, and retrieval filters by it.
Answer
The model receives only the passages left after the filter. Access removed in the source leaves the index at the next sync.

Our own knowledge base, where agents file every meeting

The closest shipped work: every meeting filed as a note under its project, with the transcript as its source. LifeOS case study

Every meeting, filed under its project

Meetings filed under their project

Every action item, in writing

Action items captured in writing

One knowledge base, versioned in git

Notes in the knowledge base

Data handling: where the data goes and how ELASTO handles security.

Company knowledge: common questions

RAG or fine-tuning for a company assistant

RAG, in nearly every case. Company knowledge changes every week and has to be cited and filtered by permission, which fine-tuning cannot do. Fine-tuning helps only when the assistant must write in a fixed format, and it can sit on top of RAG.

Answers from Confluence, Google Drive or PDFs, with the source

Confluence, Google Drive, SharePoint, Notion and PDFs are read where they live, and scanned PDFs go through OCR on the server. Every answer cites the document and the section, with a link that opens it.

Documents some people are not allowed to see

The assistant filters every search by the asker's permissions in the source system before the model sees any text. A document someone cannot open in SharePoint or Drive cannot appear in their answer, and removed access leaves the index at the next sync.

Keeping answers current when documents change

Sources sync on a schedule or through each system's change notifications, so an edited policy replaces the old passages. Every answer shows when its document was last edited, and deleted documents drop out of the index.

Company data in ChatGPT vs a private assistant

Consumer chat apps can keep and use conversations under their own terms. A private assistant calls the model through a business API, where Anthropic does not train on the data without explicit permission, and the index sits on infrastructure agreed with the company. The full data path is on the AI agents page.

When the documents hold no answer

The assistant says the documents do not cover the question and lists the closest passages it found. It does not fill the gap from the model's general knowledge. Unanswered questions are collected, and they show which documents are missing.

One team's documents, mapped in the AI workshop

Which sources, which weekly questions, and the cost to build and run it.

Studio
Cluj-Napoca, Romania, EU