Every figure on this site is a test the software has to pass.

How well the matching works, whether the ethical wall holds, whether the screens are usable by everyone: each is measured, and each stops a release if it gets worse. This page sets out what we measure, how, and what we are not claiming.

Two piles of record pairs, two cut-off lines, and one number that matters more.

We build a test set where we already know which records are the same client and which are not, so the figures can be calculated rather than estimated.

How the test set is built

We generate a firm's worth of records and record, as we go, which pairs are the same client. Then we damage them the way real systems do: misspelled names, changed suffixes, moved addresses, a registration number typed once and never again. Because the right answer is known for every one of the 17,261 pairs, the figures below are counts, not estimates.

Without that check, a well-meant improvement quietly makes results worse for every firm already using Datum, and nobody finds out until someone complains months later. So the measurement runs on every build, and a figure that falls stops the release.

How confident Datum is that two records are the same client Different companies The same client
−20 −10 0 +10 +20 HOW STRONG THE EVIDENCE IS REJECTED ↓ ↑ MERGED NOBODY IS ASKEDA PERSON DECIDESMERGED AUTOMATICALLY Different companies The same client 5.8% of real duplicates fall below the reject line

One figure asks how many of the automatic merges were right. The other asks how many real duplicates Datum found at all. The second matters more, because a duplicate Datum rejects on its own is one nobody is ever shown, and it stays in your data for good.

FigureWhat it meansHow it is kept true
99.8% Of the records Datum merged on its own, this many really were the same client. Engineers call this precision. Checked on every build against the release before. If it falls, the release does not ship.
94.2% Of the real duplicates in the test set, this many Datum found. Engineers call this recall. The 5.8% it misses fall below the reject line, and nobody is shown them. Checked on every build. This is the figure that matters more, because a missed duplicate stays in your data for good.
17,261 Record pairs where the right answer is known outright, because we built the test set that way. The right answer for every pair is known before the test runs, so nothing here is estimated.

These figures come from our test set, not your data. Your figures will differ. A pilot runs exactly the same measurement against your records and produces your numbers.

Eight checks, any of which stops a release.

Engineers call these gates: a check the build has to pass before it can go anywhere. These eight are requirements, not housekeeping. Any one of them failing stops the release, and nobody has to remember to look.

CheckWhat fails the build
Where the code comes fromAny dependency outside the permitted set, across both sets of third-party components Datum is built from. We prove the check works by deliberately adding a component with the wrong licence. It caught a real one on its first run, then a second that had sat there unrecorded.
How well matching worksAny of the matching figures above getting measurably worse than the release before it.
The ethical wallA screened matter becoming reachable, including by deduction from a total or through an export, which is what a penetration test would try.
Runs without the AI partsAnything that quietly comes to depend on the AI integration. The whole platform is built and tested with every AI connection switched off.
Internal separationThe parts of the system that are meant to stay independent of each other still are, checked automatically rather than agreed in a review and forgotten.
Screens and platform agreeAny request the review screens make that the platform does not actually answer. The two are written in different languages and nothing else joins them up.
AccessibilityWCAG 2.2 AA in both colour themes, across every screen in the product.
Consistent formattingCode that has not been tidied to the house style, because an argument about formatting in a review is time not spent on the review.

Where two parts of the build had each grown their own copy of a check, there is now one definition that both use. Two copies of the same rule drift apart, and neither side can see it happening.

Eighteen problems were found in the design, and two would have shipped as confidentiality bugs.

The original design was reviewed before anyone started building. Every fix in the platform traces back to an entry in that review, and the review ships with the product rather than being produced for a sales meeting.

18

Problems recorded

Against the original design, each with the fix that was adopted.

4

Serious enough to stop work

One of them would have let a single record end up as two different clients at the same time.

2

Confidentiality bugs avoided

Including one that let a screened person through because of how two conditions combined.

Where a requirement has a number in it, it is a test rather than an opinion.

The review desk is designed against a steward's eight-hour day, and the numbers that make it usable are checked on every change.

25+

Rows on screen at once

Fitting this much on a 1080p screen is a requirement, not a styling preference. A prettier screen that fits twelve rows is rejected by the people who sit with it all day.

16

Screens checked for accessibility

Every screen meets WCAG 2.2 AA in both light and dark. Every box on every chart can be reached and read out, and anything you can do with a mouse you can do with a keyboard.

1,737

Automated tests, including the database rules

The rules that stop the database holding two conflicting versions of the same thing are tested against a real database, not a stand-in.

The limits, set out here rather than discovered in week three of a pilot.

What these figures do not say, and what Datum does not claim.

These figures come from our data, not yours

We know the right answer in our test set outright, which is what makes the measurement possible, and also what makes it not yours. A pilot runs the same measurement against your data and produces your numbers.

Some plumbing is designed for, not yet running

Scheduling, loading, search, sign-in and audit trails are each designed with a clean join to whatever you already run, and assembled from proven parts rather than rebuilt. The technical brief will say exactly where each join is and what sits behind it.

The product tells you what it has not finished

When it starts up, the platform prints which of its parts would survive a restart and which would not, before it accepts a single request. Anything not yet durable says so on screen, and that list is something to go through with us, not around us.

It never decides whether a matter can be taken

Datum is not a conflicts system and issues no clearance decisions. It gives your conflicts system better data to work from. Nothing in it works out a clearance, stores one or hands one back, and no field could be mistaken for one.

See the measurement run on the demonstration firm. Book a demo