The AI project we nearly talked a client out of
They asked for a chatbot over eleven years of policy documents. The honest first answer was that search would be cheaper — here is why we built it anyway.
A professional services firm came to us with a familiar request: an AI assistant that could answer questions about their internal policies. Eleven years of documents, six thousand files, and staff who spent a genuinely absurd amount of time looking for the right paragraph.
Our first recommendation was not to build it.
Why we pushed back first
Before reaching for a model, it is worth asking what a good search box would do. Their existing one was terrible — it matched filenames, not contents. A decent full-text index with some filtering would have solved perhaps sixty per cent of the problem for a fraction of the cost.
We said so. Out loud, in the first meeting, before a proposal existed.
What changed our mind was watching people actually work. The questions were not lookups. They were things like “can a contractor on a fixed-term engagement claim the travel allowance, and did that change after 2021?” — a question whose answer lives in three documents, two of which supersede the third.
Search cannot do that. Search finds documents. This needed something that could read across them.
What we built
A retrieval-augmented system, but the interesting work was not the model.
Chunking was the whole ballgame. Policy documents have structure — clauses, amendments, effective dates. Splitting them every 500 tokens destroys exactly the context you need. We chunked on document structure and attached the effective date and superseding relationship as metadata on every chunk.
Retrieval had to understand time. A 2019 clause that was amended in 2021 is still in the corpus and will still match a query. The retriever had to prefer current versions while keeping the option of answering a historical question.
Every answer cites. Not as a nicety — as the safety mechanism. A staff member who cannot check the source will eventually be misled by a confident wrong answer, and in this domain that has consequences. Every response links to the clause it came from.
The part nobody puts in the case study
We built the evaluation set before the system. Two hundred real questions collected from the firm’s own support queue, each with an agreed correct answer and the clause it should cite.
That set is why the project worked. Every prompt change, every retrieval tweak, every model upgrade ran against it. Twice we made a change that felt obviously better and measurably was not, and we only know that because the numbers said so.
Final accuracy: 94% of answers cited the correct clause. The remaining 6% is not hidden — it is documented in the model card, and the interface tells a user when retrieval confidence is low rather than guessing fluently.
What it changed
Median lookup went from about fourteen minutes to forty seconds. That is the number in the case study.
The number I find more interesting is that usage kept climbing after month three. Internal tools usually spike and decay. This one did not, which suggests people trusted it — and they trusted it because it showed its working.
The lesson we keep relearning
The model is rarely the hard part now. The hard parts are: what does the data actually look like, how do we know this is working, and what happens when it is wrong.
If a vendor is enthusiastic about the first and quiet about the other two, that is worth noticing.