2026in development
Finance dashboards with an analyst that shows its working
Businesses were seeing their numbers through a third-party BI embed that made them log in twice and still could not answer a simple question. I am replacing it with a portal of reviewed management KPIs, and an AI analyst that answers questions from the business’s own data, and shows exactly how it got there.
- role
- solo build
- built
- 2026, in development
- running on
- two companies’ real books
- runs on
- bigquery · dbt · claude api
// the problem
Each client’s numbers lived in an embedded third-party BI tool. It made users sign in twice, broke outright in Safari, was priced per seat, and needed a separately published dashboard for every client. And it only answered the questions someone had thought to build a chart for. Anything else still meant emailing an analyst.
The vendor sold a chat-with-your-data feature that would have covered the second problem, at considerably more cost, and without fixing the first.
// the accounting calls in it
- Every KPI is a reviewed definition in code. Sales, channel mix, receivables and payables ageing, inventory days on hand, the reported P&L and balance sheet, and unit economics. A new bespoke number costs a review and a deploy, and that friction is deliberate.
- Acquisition spend is not selling cost. Marketing to win a customer is spent whether or not a sale happens; the cost of fulfilling a sale is not. They are kept apart, because blending them makes both meaningless.
- Overhead is never spread by revenue share. Contribution is shown only over overhead that genuinely belongs to a channel. A board deck that came before it had apportioned overhead by revenue, and one cell in it read −156,189%.
- Lifetime value against acquisition cost is shown as a floor, not an estimate, because the known biases in the data all push in the same direction. Saying so is more useful than a confident ratio.
- Distortions are disclosed, not hidden. Where a number is known to be skewed (an order average inflated because some “orders” are daily roll-ups), the tile says which way it is wrong, rather than going blank.
// the decision I’d defend
I withdrew a finished thirteen-week cash-flow forecast, because a straight line reads as confidence.
The tab was built and working. But the business behind it does no forecasting and no demand planning, so nothing in the data could say anything true about future cash. Every line on it was a trailing average drawn forward, and a chart like that looks exactly like a forecast while carrying none of a forecast’s information. A founder would plan against it.
So it came down. No forecast is an honest gap. A forecast that implies knowledge nobody has is a liability with a nice chart on it.
// the analyst
For questions no dashboard anticipated, there is a chat panel. The model writes a read-only query, the warehouse runs it, and the model explains the rows that come back. Under every answer, a “how I got this” panel shows the query and a sample of the rows, so a finance user can check the logic rather than take it on trust.
This bends the rule I hold everywhere else (that the model never calculates), so it is worth being precise about how far. The model still never does the arithmetic; the warehouse does. But it does choose the arithmetic, and that choice can be wrong. So the choice is made in the open, and inside hard limits.
// the limits it works inside
- One read-only question at a time. Anything that is not a single SELECT statement is rejected before it runs, checked by parsing the query, not by pattern-matching it.
- Each company can only ever see its own data. Every query runs under that company’s own read-only identity, so reading another company’s books is impossible at the permissions layer. The prompt is not trusted to enforce it.
- Capped, so a bad question cannot become a big bill. Limits on data scanned per query, queries and model usage per day, rows returned, and steps per answer.
- Every query is logged, and a question asked often enough is meant to graduate into a reviewed dashboard tile: out of the model’s hands and into tested code.
// known gaps, and how I’d close them
- There is no automated test of answer accuracy yet. That is the honest reason the query is always on screen. While building the dashboards, summing the as-reported column instead of the right one flipped the sign of a company’s profit in a hand-written tile, and nothing structural stops the chat making the same mistake except a reader checking the query. The fix is a fixed set of real finance questions with known answers, run against every change, before anyone outside the team relies on it.
- The chat does not yet share the dashboard’s definitions. The tiles carry reviewed definitions of revenue, margin and the rest; the chat has to infer them each time. The fix is to hand it those same definitions, so a number means one thing whichever way you ask for it.
- It is steered to the modelled tables, not yet confined to them. The raw tables underneath are hidden from it rather than locked. The fix belongs in permissions, not in the prompt: the same principle that keeps one company out of another’s data.
- Not hardened for release. Rate limiting and a penetration test of the query checks come before anyone outside the team uses it.
// what I deliberately didn’t build
- No “anyone with the link” dashboards. It would have been the cheap fix for the double login. A hard-to-guess URL is not an access control for a company’s financials.
- No stored conversations. Chat is ephemeral; only the query is logged, not the question a user typed.
- No admin screens for onboarding. With one operator who can run a script, a form is strictly slower than the script.