Notes
Numbers don't belong in a retrieval system
Why figures stay in SQL, and what goes wrong when a language model is asked to remember a total.
The most common request I get is some version of “can it just answer questions about our data?” The honest answer is yes for words and no for numbers, and the difference matters more than any model choice.
What retrieval is good at
A retrieval system takes a question, finds the passages most likely to contain the answer, and hands them to a language model to write a reply. It is very good at prose: a clause in an operating agreement, the paragraph in a policy that governs a situation, a decision recorded in minutes three years ago. Each passage arrives with its source, so the answer can say where it came from.
That is a real capability, and for most of what a company knows it is the right one. Most institutional knowledge is prose.
What goes wrong with figures
Ask the same system for last quarter’s total, and something different happens. The model receives a handful of passages that mention amounts, and it produces a confident number. It cannot sum. It cannot join two tables. It does not know that the figure in one passage was superseded by a correction in another, or that a payroll register has three rows for the same person. It will not say “I am not sure.” It will give you a total, formatted nicely, and it will sometimes be right, which is the worst possible property for a number to have.
The failure is not a bug to be fixed by a better model. It is what happens when arithmetic is done by predicting the next word.
The right store for each kind of data
The approach I use is unglamorous. Payroll registers, plan data, and financial figures are extracted from whatever they arrive as, PDFs, spreadsheets, scans, into tables with a schema, and those tables live in a relational database. Scanned pages go through optical character recognition first, on hardware I own, and the extracted rows are checked against the totals printed on the page before they are accepted.
When a question needs a number, the model does not remember one. It writes a query, the database runs it, and the result comes back with the rows that produced it. Every figure is auditable back to the cells it was built from.
Prose still goes into retrieval, where it belongs. The two are joined by the agent, which decides which store a question needs and cites what it used. The person asking gets one answer, but underneath it the words came from one place and the numbers from another.
Why this is worth the trouble
Building the extraction pipeline is more work than pointing a model at a folder. It is also the difference between a system that can be trusted with a question about money and one that cannot. In workers’ compensation and payroll administration, the numbers are the business. A system that gets them approximately right is not a convenience; it is a liability with a friendly interface.
The rule I keep to is short: a language model is never asked to remember a figure it could look up.