writing4 min read
Don’t let the model do the arithmetic
Every business that puts a language model in front of its numbers will eventually be shown a figure the model made up. The fix is not a better prompt. It is to stop giving the model anything to calculate.
// the problem
The pitch for AI over business data is that anyone can ask a question in plain English and get an answer. The risk is the same sentence: the answer arrives fluent, confident and nicely formatted, and nothing about it tells you whether the number in it is right.
That is an accounting problem before it is a technology one. The first thing I ask about any figure is where it came from: which ledger, which definition, which calculation. A number nobody can trace is not a number; it is an opinion with a decimal point. Ask a language model to derive something and that is exactly what you get back.
// what I measured
Told not to quote numbers, the model quoted 47 of them across six samples. With the numbers removed from its input, it quoted none.
That was armchair.run, a deliberately silly app that writes commentary on Strava activities. Silly or not, it publishes with nobody checking first, so it cannot get the facts wrong. Instructing the model was the obvious approach, and it failed comprehensively. Taking the figures out of its input worked completely, with no instruction at all.
Roos Research, a paid investment-research product, reached the same rule from the other direction. Live testing caught the model getting ordinary derivations wrong: how far a price sat from its moving average, whether one indicator was above another. Not exotic maths, just the kind of subtraction a spreadsheet never gets wrong.
One serious product and one joke, built for unrelated reasons, arriving at the same conclusion. That is why I treat it as a rule rather than a preference.
// the rule
Anything that can be computed gets computed in code, where it can be tested and re-run. The model only interprets finished figures.
In practice that means doing the arithmetic first (every ratio, difference, ranking and comparison the answer could need) and handing the model the results as named, finished facts. In Roos, the model is told how far the price sits above its moving average; it is never handed the price history and asked to work it out. armchair.run goes a step further: the facts are turned into plain sentences before the model sees them, so it never sees a number at all and has nothing to misquote.
It also moves judgement somewhere it can be checked. What counts as notable (a long stop, a hard climb, a heart rate drifting upward) is decided by rules in code, not by the model’s sense of drama. A rule in code can have a test. A tendency in a prompt cannot.
None of this is an argument against automation. Everything here still runs on its own, with nobody typing a figure in. The point is to know a figure is right before it ships, and I know it is right because of how it was produced: calculated the same way every time, by something that can be tested, rather than inferred by a model. It is still automated. It is just automated by something you can check.
// what it costs
- More code, and more tests. Roos Research carries 142 passing tests, because the maths is the product. That is the price of a figure you can reproduce, and it is cheap next to explaining a wrong one.
- The model only sees what you thought to compute. Roos never shows it raw price history, so it cannot spot a pattern nobody wrote code to measure. I would rather miss an insight than publish an invented one, because one figure found to be wrong breaks trust in all of them. Where did it come from? Is the calculation wrong, or the warehouse, or the chart? And how does anyone know they can trust any other number on the page? That trust is very hard to win back, if it comes back at all.
- Someone has to decide what matters, in advance. That is a real design job, and it is the same job as deciding which KPIs a business reports. It is not overhead. It is the point.
// if you are putting AI over your own numbers
- List every figure the output could contain. Each one should trace to a calculation you can point to and test, not to the model.
- Remove, don’t instruct. If the model should not be producing a number, do not give it the inputs to produce one. An instruction is a request; the input is a constraint.
- Test the calculations, not the prose. The wording can change from one run to the next. The figures must not.
- Show the working. Wherever the model has to do more than interpret, put what it did in front of the reader so they can check it. Trust that cannot be checked is just hope.
- When it must choose the calculation, let something else run it. In a finance portal I’m building, the model writes the query that answers a user’s question, but the warehouse does the arithmetic, under a read-only identity and hard limits, and the query sits under the answer for anyone to read.
// the short version
A language model is a very good writer and an unreliable bookkeeper. Give it the job it is good at.
// the work behind this
Roos Research
A subscription research desk for self-directed investors. It turns market data and company news into a plain-language brief on a company.
armchair.run
Reads a finished Strava activity, works out what actually happened in it, and writes the commentary into the description, in a voice you pick.
Finance dashboards with an analyst that shows its working
A portal where a business sees its management KPIs on one page, and can ask “who do I owe this week?” and get an answer from its own books, with the query shown underneath.