2026live
Inventory data that reconciles itself every night
Businesses running their stock and orders on Cin7 can’t report across years of sales, purchases and stock movements without exporting spreadsheets or leaning on the live system. I built the platform that keeps a complete copy of that data in each business’s own warehouse, and that knows, every morning, whether the copy is still complete.
- role
- solo, built twice
- built
- 2026
- per client
- about seventy tables
- runs on
- python · bigquery · gcp
// the problem
Reporting for a business like this is built on its inventory system’s data, so that data needs a complete, current copy somewhere it can be queried. Copying it is easy. Copying it without losing anything is not.
The source returns data a page at a time. A missed page, a timeout or a reply that is silently shorter than it should be loses records, and nobody notices until a report fails to tie, by which point the wrong number has already been in front of a client.
// the decision I’d defend
The sync is allowed to miss a record. It is never allowed to miss one silently.
Every few hours a job copies what has changed. It is built to be fast and cheap, and it is deliberately not trusted for completeness. Anything that fails goes into a retry queue that repairs itself. Then, every night, a separate check lists what exists in the source, compares it with the warehouse record by record, and queues any gap for repair.
Each client ends up with a green, amber or red status, and a person only needs to look at red. If too much is missing at once, the repair stops and raises an alert instead of re-fetching everything: a large gap means something systemic, and that needs a person, not a loop.
That is a reconciliation. It is the oldest control in accounting (don’t trust the feed, tie it back to the source) applied to a data pipeline.
// what I built
- The sync, the retry queue and the nightly reconciliation, covering roughly seventy tables of sales, purchases, stock and products per client, with every load safe to re-run without creating duplicates.
- A control panel for operators, so a client can be added, scheduled, paused, backfilled and monitored without touching code or a deploy.
- The data lands in the client’s own cloud account. They own it. Ending the relationship never means a data migration.
- A demo dataset that adds up. It relabels one real client’s data, with every name replaced and every number, date and relationship left intact, so any total on the demo equals the real one, and a demo survives a prospect checking the arithmetic. Scaling the money or shifting the dates was ruled out permanently, because either one breaks that.
// what changed
- The checks found real defects. An audit they made possible showed that in one dataset about a third of sale-transaction lines were duplicated, and that most datasets were affected to some degree. The fix went in at the source (how line items are written), and the nightly check now looks for it.
- Completeness is a status, not an assumption. Every client’s copy is checked against the source every night, and the result is visible without anyone going to look.
// what I deliberately didn’t check, and why
- Not every field of every record. The reconciliation proves records exist and line up; it does not compare every value. That is a stated limit, which is why this page says “never silently” rather than “guaranteed complete”.
- Not the largest client in full, every night. A full pass takes most of a day, so it gets a check of recent changes nightly and a full check on demand, giving up nightly detection of losses in old, untouched records.
- Not every line item, every night. Line-level checks sample a portion each night so the job never overruns its time limit.
- No modelling in here. This copies the data faithfully. Deciding what it means is a separate project, kept separate on purpose.
// built twice
The first version was a proof of concept. The source system’s API is awkward enough that whether it could be copied reliably at all was a genuine question, and the cheapest way to answer it was to try.
The second version is the one a business can rely on. The difference is almost entirely the reconciliation, the repair and the status a person can read at a glance. That is the difference between a copy of the data and a copy you can trust.