Performance
How Numeralens behaves on large ledgers, how those numbers were obtained, and where it stops being comfortable.
Every figure here is a measurement. Nothing in this document is estimated, extrapolated or quoted from a marketing deck, and where a metric could not be measured it says so rather than filling the gap.
How to reproduce every number
Two harnesses, measuring two different things. They do not replace each other.
1. Engine benchmark (Node)
npm run bench # every dataset size
npm run bench -- --bench-sizes=0,100000 # just two
Runs tests/engine-benchmark.test.ts with BENCH=1. It measures the
calculation engine alone — no React, no network, no artificial latency — which
is the honest basis for statements about query cost. Each figure is the median
of five runs.
2. Performance Lab (browser)
Open /performance in the running app. It measures what a user actually
experiences: real Web Vitals from PerformanceObserver, the real number of
<tr> elements in the document, adapter round-trips including React scheduling,
and live-update render latency. A metric the browser has not reported yet
renders as "—"; it is never back-filled with a plausible-looking number.
Measure in a production build. Development React is several times slower and its numbers are not representative:
npm run build && npm run start
Keep the tab visible. Browsers clamp setInterval in a hidden tab to roughly
once per second, which caps the live-update stress test at about four events per
second whatever rate you select. The Lab detects this and says so on screen
rather than reporting the browser's policy as an application limit.
Dataset sizes
The Performance Lab switches the in-memory demo ledger between sizes. Every size is generated from the same deterministic seed, so a benchmark is repeatable.
| Size | Ledger lines | What it represents |
|---|---|---|
| Base | ~8,900 | A realistic mid-market ledger: 25 months, 24 customers, 17 vendors |
| 10K–250K | as labelled | The base ledger plus a synthetic high-frequency self-serve channel |
What the synthetic volume is. Small-ticket online orders and card expenses, each a balanced double-entry document. Deliberate properties:
- Balanced. The trial balance balances and the balance sheet reconciles at
every size — asserted by
tests/scale.test.tsandtests/balance-integrity.test.ts. - Small amounts. A 100K-line ledger looks like a business with a busy online channel, not one whose revenue grew fifteenfold.
- Budgets follow. Synthetic monthly totals are folded into the budget lines, so budget-vs-actual variance stays meaningful instead of showing a fictitious windfall.
- Cash-settled. Synthetic sales settle immediately and their cost is paid from the bank, so receivables, payables and the stock ledger keep the readable size and balances of the base dataset. Consequence: liquidity ratios rise with dataset size, because the synthetic channel adds cash without adding liabilities. Read ratios from the base dataset; read performance from the large ones.
- Seeded anomalies. Exact duplicates, a near-duplicate pair, a material weekend posting and a round-dollar cluster, so the Anomaly Sentry has genuine findings to locate at any scale.
Measured results — calculation engine
Node v24.16.0, win32 x64, median of five runs, milliseconds.
| Ledger lines | Build | Index | Page (500) | Search | Filter | Sort | Dashboard | Ledger | Trial bal. | Anomaly scan | Risk scores | AI context | Forecast |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| 8,922 | 35 | 7.5 | 1.1 | 0.4 | 0.5 | 1.4 | 8.9 | 2.1 | 1.8 | 19.6 | 13.3 | 13.6 | 19.9 |
| 10,010 | 27.5 | 10.0 | 1.5 | 0.5 | 0.5 | 1.3 | 10.0 | 3.4 | 3.4 | 27.3 | 16.8 | 15.7 | 17.3 |
| 25,012 | 56.9 | 20.9 | 2.2 | 1.3 | 1.0 | 4.3 | 11.7 | 3.2 | 5.1 | 31.2 | 10.6 | 32.2 | 22.0 |
| 50,010 | 92.7 | 45.4 | 4.5 | 3.4 | 2.8 | 8.9 | 31.6 | 6.7 | 15.6 | 55.5 | 10.4 | 81.0 | 39.8 |
| 100,018 | 200 | 82.4 | 7.2 | 5.8 | 5.4 | 28.5 | 76.0 | 17.3 | 39.2 | 105.9 | 11.4 | 195.5 | 79.7 |
| 250,022 | 635 | 276.9 | 25.1 | 17.0 | 12.9 | 91.6 | 237.0 | 43.3 | 113.0 | 293.8 | 12.1 | 644.7 | 198.2 |
Build is one-time dataset generation; Index is the one-time build of the lookup structures including the search haystack. Both are paid once per dataset size and then cached. Anomaly scan is a single cold run — it memoises per dataset and range.
Reading the table
- Browsing a 100K ledger is fast. Paging, searching and filtering are all under 8 ms because the ledger is date-sorted and the range narrows by binary search before any predicate runs. Free-text search matches a lowercased haystack built once per dataset, so a keystroke costs a substring scan, not five string concatenations per row.
- Sorting scales with the result set, not the ledger. Sorting 40K matching rows costs ~28 ms at 100K lines. It is the most expensive table operation because it is the only one that touches every matching row.
- Aggregation is the real cost. The dashboard is 76 ms at 100K and 237 ms at 250K; the AI context — six full reports assembled into one payload — is 196 ms and 645 ms. These are the numbers that decide the ceiling.
Where it degrades
100K lines is comfortable. 250K is the practical ceiling for this architecture. At 250K every interaction a user drives directly (page, search, filter, sort) is still well inside the 100 ms that feels instant, but assembling the executive narrative takes two thirds of a second and the dashboard a quarter of a second before React has done anything. Dataset build reaches 635 ms and the ledger occupies a few hundred megabytes of browser memory.
This is a property of doing analytical aggregation in JavaScript over an in-memory array — which is what demo mode is. A deployment with a real backend pushes aggregation into the database, where a hundred million rows is unremarkable, and the numbers above become irrelevant to it. They describe the demo, and the demo is what a buyer evaluates.
Live updates
The live ledger and the user interface are deliberately two separate loops.
scheduler tick (4 Hz) publish (1 Hz)
emit events setState(rows, feed, metrics)
append to ledger + indexes invalidate aggregates (throttled)
write to refs — no render → one React commit
Emission never calls setState, so a slow render cannot starve the ledger. This
was not the original design, and the original design was measurably wrong: with
React work on the emit path, a 1,000 events/minute run in development delivered
roughly a third of the requested events and rendered at 273 ms per batch.
Measured
| What | Measurement |
|---|---|
| Ledger capacity (Node) | 246,615 events/second — 20,000 events applied to a 100K-row ledger in 81 ms, 0.004 ms each |
| Render latency, production build, 1,000/min | 11.8 ms (p95 57.2 ms) |
| Render latency, development build, 1,000/min | 258 ms — development React, shown for contrast only |
| Publish cadence | 1 per second, by design |
What "render latency" means here. Publish → React commit. It is deliberately not "emit → paint": the ledger is updated the instant an event is emitted, and measuring from emission would report the one-second publish interval as if it were processing cost. That would be flattering in one direction and alarming in the other.
Update rate matrix
| Rate | Ledger | UI |
|---|---|---|
| 100/min | 0.007% of capacity | 1 commit/second |
| 250/min | 0.017% | 1 commit/second |
| 500/min | 0.034% | 1 commit/second |
| 1,000/min | 0.068% | 1 commit/second, 11.8 ms each |
The UI column is identical at every rate because publication is time-based, not event-based. Raising the rate raises the number of events per published batch, not the number of renders. This is the whole point of the split, and it is why the rate ladder shows no degradation: the tested rates are nowhere near any limit the application has.
Honest caveat on the dropped counter. The Lab reports events the requested rate called for but the scheduler did not deliver. In a foregrounded production tab this stays at zero for the rates offered. In a hidden tab it will be large, because the browser clamps background timers — the Lab detects that case and labels it explicitly instead of presenting browser policy as an application limit. The stress-test figures above were taken with the tab hidden, so their delivered counts are not meaningful; the render latency and the Node capacity figure are, and those are the two numbers that describe the system.
What does not happen on an update
- No refetch of the transaction list.
- No invalidation of the 100K-row query.
- No recomputation of aggregates per event. Dashboard, P&L and cash-flow queries
are invalidated at most once per
aggregateRefreshMs(default 5 s, adjustable in the Lab, can be switched off), and only for queries that are currently mounted —refetchType: "active".
Table virtualization
The transactions explorer and the Performance Lab render only the rows in and near the viewport.
Measured in the browser: 26 DOM rows for a 100,006-line ledger with 40,982
rows matching the active filter and 500 loaded. The Lab counts
tbody tr[data-index] elements in the live document, so the figure is read from
the page rather than derived from the overscan setting.
The implementation is a fixed table layout with spacer rows above and below the
window (@tanstack/react-virtual). Spacer rows rather than absolute positioning
keeps real <table> semantics — screen readers still announce a table, column
widths stay stable, and the sticky header behaves normally.
Loading is windowed, not bulk. Rows arrive in server-paginated batches of 500
as you scroll (useInfiniteTransactions). A 100K ledger is never transferred in
one response, so the same page works against a real API where that would matter.
The transactions page keeps both modes:
- Paged — classic server pagination with page-size control. Unchanged from 1.0.
- Continuous — windowed loading with virtualization, and where live rows appear.
Web Vitals
The Lab reads TTFB, FCP, LCP, CLS, long tasks and JS heap from
PerformanceObserver, and the longest Event Timing duration as a proxy for INP
(labelled as a proxy, because true INP needs the full interaction lifecycle).
These are not published here as product claims, for a reason worth stating: they
depend on the machine, the build, the dataset size selected and whether a
benchmark is running. A number measured on the author's laptop is not a number
about your deployment. Open /performance on the hardware you care about; the
page will show you real values, and blanks where the browser has not reported
one.
What made the difference
The optimisations that moved the numbers, in the order they mattered. Each is a single commit-sized change, and the comments in the code say why.
- Binary search on a date-sorted ledger (
src/mock/indexes.ts). Every report narrows by date first. Turning that from a full scan into a slice is the difference between O(ledger) and O(range) on every widget. - A pre-built search haystack. One lowercased string per row, built once, lazily. Search went from re-concatenating five fields per row per keystroke to a substring test.
- Killing an N+1. The expense-anomaly rule scanned the whole ledger once per bill to build a vendor's trailing average: 1,352 ms → 167 ms at 100K.
- Single-pass balance series (
balanceSeries). The dashboard's cash line calledbalanceAtonce per chart bucket, rescanning the largest account's full history for every point: dashboard 481 ms → ~262 ms. - Per-account and per-reference indexes. Trial balance, balance sheet and the journal-entry drawer walk only the rows they need.
- Not sorting 40,000 rows to show ten.
recentTransactionstakes the tail of an already-sorted slice. - Separating the live emit loop from React (see above).
Known limits
Stated plainly, because a performance document that lists only wins is marketing.
- Aggregation above 250K lines in demo mode is slow. The AI context takes 645 ms at 250K. Connect a real backend for larger ledgers.
- Dataset build blocks the main thread. Generating 250K lines takes ~635 ms in Node and around a second in the browser; the Lab paints a pending state first, but the tab is unresponsive during the build. It happens once per size and is then cached.
- Memory. A 250K-line ledger plus its indexes and haystack occupies a few
hundred megabytes.
performance.memoryis Chromium-only, so the heap figure is blank in Firefox and Safari. - The AI endpoints run server-side against their own copy of the dataset. In
demo mode the Performance Lab's size switch is passed through as
datasetSizeso both sides agree; against a real backend the parameter is ignored entirely. - After-hours posting detection is not implemented. Ledger rows carry a posting date but no timestamp, so the rule cannot be evaluated. The Anomaly Sentry lists it as skipped, with the reason, rather than quietly omitting it.
- Mobile layout is verified by test, not by device. The card layout and its
disclosure behaviour are covered by
tests/data-table-responsive.test.tsxagainst a mocked breakpoint. Physical-device verification is worth doing before shipping to a mobile-heavy audience.