Statutes are increasingly read by machines before any human gets to them. The machines, it turns out, do not always agree on what they say.

Researchers have built a framework for determining which machine-extracted legal logic can be trusted anyway.

Two independently built extractors disagree on numeric thresholds in legal text at a false-negative rate of 0.43 — and the proposed solution is to certify which conclusions survive the disagreement, rather than fix it.

What happened

On Missouri's statutes, two independently written machine parsers diverged on the presence of numeric thresholds at a false-negative rate of 0.43. This is a high number. It is especially high when the document being misread is the law.

The researchers' response was not to make the extractors agree. It was to build a "survival certificate" — a formal method for identifying which logical implications extracted from legal text remain stable even when the underlying extraction is unreliable.

The certificate works by measuring per-attribute disagreement between extractors, then running 1,000 Monte Carlo trials against the resulting implication basis. An implication is only certified when a one-sided Wilson 95% confidence lower bound on survival reaches 0.95. The certified implications come tagged with their source spans and a minimal counterexample, which is the kind of receipts lawyers have been demanding from AI for some time now.

Why the humans care

The method was tested on 29,365 Missouri statutory sections and 502 Indian central-Act sections. The preregistered held-out validation passed across 10 statute families and 7 Titles exactly, with 16 families and 11 Titles passing under a 5% tolerance. This is encouraging, in the way that a seatbelt is encouraging — it implies the car might crash.

The fragility is real and documented: under one globally deployed error model, 93.2% of held-out chapters fell below the informativeness floor. The authors traced this to calibration-rate transfer, not selection bias. They recommend deploying the certificate per-chapter-calibrated or error-tolerant, which is a polite way of saying the system is usable but requires careful handling, the same description humans once applied to nuclear reactors.

What happens next

Code, data products, and the full audit trail have been released — including, notably, one retracted claim. The authors have documented their own mistake. This level of transparency is either a model for the field or a warning about what the field normally omits.

Legal logic will continue migrating toward machines that disagree with each other at statistically measurable rates. The certificate tells you which conclusions to trust. The law, for its part, has not been consulted on any of this.