Consistency is the legal AI risk that gets discussed least: the same facts producing a different answer six months apart, with nothing in place to catch it.
Shaz Aziz, Head of Client Solutions, Neota Logic · Oct 7, 2026
In 2025, researchers put 500 real federal appeals to leading models 20 times each, with identical text and the models set to be as consistent as their settings allow. On up to half of the questions, the same model named one party the winner, then the other.1 A legal expert reviewed pairs of the contradictory answers and found the analysis sound in every one. The model simply weighed the arguments differently each time.
Legal work has always tolerated a version of this problem in people, because a lawyer's inconsistency sits inside an accountability structure: review, a file, insurance, a regulator. A model's drift has no equivalent. The major platforms spent 2026 shipping the first half of that structure, with audit logs, admin controls, and ethical walls that show who had access, what was configured, and increasingly what was said. None of it shows whether the answer held over time, or why it came out the way it did.
This paper sets out the evidence on why AI answers move, why human review cannot make them consistent, and why the accountability question is arriving on a regulatory clock rather than a procurement one. It closes with five questions to take into any vendor meeting, ours included, and what a defensible answer to each one sounds like.
In 2025, researchers at the University of Maryland and Johns Hopkins ran a simple experiment on leading AI models. Anyone buying legal AI should know the result.
The researchers took 500 real US federal appeals, each one a case where the judges themselves had split two to one.
Each case was reduced to a statement of facts plus the best arguments for each side. Then they asked each model a single question: which party should win?
They asked it 20 times per case. The text was identical every time, and the models were set to be as consistent as their settings allow.
On a large share of cases, the same model named one party the winner, then the other.
A legal expert reviewed pairs of the contradictory answers and found the analysis sound in every one. The model simply weighed the arguments differently each time.1
Across all 500 cases, here is how often each model flipped. For Gemini 1.5, half the questions got both answers.
Accuracy gets checked in January, at the pilot and the sign-off. Consistency is what you need to know about in July, when the same clause comes back with a different answer and nobody is comparing.
The model changes underneath you. GPT-4's accuracy at identifying prime numbers fell from 84% to 51% between March and June 2023. Same product, same name, and no release note.2
One prompt, run 1,000 times at temperature zero, gave 80 different completions. The cause was how the inference server batches requests together. With that fixed, all 1,000 came back identical.3
Their research shows it can be fixed, but only by whoever runs the model. Whether you get the same answer twice is an engineering choice made outside your organization.
Newer models show it too. A 2026 study had five models, including Claude Sonnet 4.5, Claude Haiku 4.5 and Gemini 2.5 Flash, score real enterprise question and answer pairs, and found substantial variability in the scores even at temperature zero.4 Those are the kinds of scores enterprise pipelines use to route, triage and gate work, which is the job legal workflows give AI outputs.
Accuracy is what gets checked in January, at the benchmark, the pilot, and the sign-off. Consistency is what you need to know about in July, when the same contract clause comes back with a different escalation decision and nobody is comparing. Princeton researchers Narayanan and Kapoor call this the capability and reliability gap, and consistency is the first reliability dimension they measure.6
Every vendor benchmark you've seen is a snapshot of one moment.
The strongest objection to the consistency argument is that lawyers aren't consistent either. Ask twenty lawyers a hard question and you'll get a spectrum of answers. That's true.
Human review is often offered as the fix, and it does real work: a reviewer can catch an answer that is wrong today. What a reviewer cannot do is make answers consistent, because review vouches for one answer at one moment.
When the same clause produces a different escalation decision three months later, the January sign-off says nothing about it, and the reviewer who approved the first answer is not in the room for the second. A reviewer can only review what they can see, and drift is precisely what nobody is shown.
The lawyer leaves a paper trail.
The model usually does not.
The scale of unchecked AI output is already on the public record. By mid-2026, Damien Charlotin's AI Hallucination Cases database listed more than 1,100 US court decisions dealing with filings that contained fabricated AI content.7 Every one of those filings had someone's name on it.
Click any solid green dot to open one of the cases.
Source: AI Hallucination Cases Database, Damien Charlotin, CC BY 4.0
Each dot is one decision.
A year ago it was fair to say nobody had shipped a governance layer for legal AI. That is no longer accurate. Over the past twelve months the major platforms have been racing to add exactly the controls enterprise legal buyers asked for.
The rest of the paper looks at what the market has shipped and what is still missing, sorts legal AI by what actually produces the conclusion, sets out why accountability is arriving on a regulatory clock, and closes with five questions to take into any vendor meeting, with what a defensible answer to each sounds like.
Read the full paper
We'll only use your details to show relevant Neota updates. Unsubscribe anytime.
All of this answers three questions: who had access, what was configured, and increasingly what was said. Those are necessary answers, and legal teams should welcome them. They are also only half of the layer. An access log tells you an associate queried the system on a Tuesday. It can't tell you whether the same query, a quarter later, came back different.
Firms are trialing nearly everything and committing to very little. On the corporate side, the ACC and Everlaw found in-house use of generative AI more than doubled in a year, from 23 percent to 52 percent, so the exposure now sits inside the legal department as much as at outside counsel. Pilots often stall because nothing makes the output dependable enough to stand behind in front of a client, and no volume of access logging changes that.
Not every use of AI in legal work carries the same consistency exposure. The useful way to sort a stack is by what actually produces the conclusion.
Tier 3 is where a decision becomes defensible
A regulator has already approved the rules half of this. In February 2026 the Solicitors Regulation Authority authorized LawFairy, which describes itself as the first deterministic, technology-only law firm in England and Wales, starting with UK immigration.12 Statutory and policy criteria are encoded as decision pathways validated by lawyers, no generative AI touches the regulated outcome, and the same facts produce the same result, in deliberate contrast to the probabilistic systems the rest of the market runs on.
The open question for everyone else is how to add a language model's fluency without giving that up. It comes down to architecture.
The model and the rules each do the work they suit,
and neither does the other's job.
The accountability question is no longer theoretical, and it is not waiting for procurement cycles.
Regulators on both sides of the Atlantic are heading the same way:
if you put an AI system into legal work, you answer for what it does.
Meanwhile there is no agreed standard for autonomy in legal AI, nothing like the SAE levels that define who is driving at each stage of a self-driving car, and that lawmakers use to decide who is liable. Legal AI today works like driver-assist, and with driver-assist the driver is still liable. That puts the liability where the review sits, on the lawyer who approved the output. A reviewer can't sit down and check the reasoning inside a model. They can check logic a lawyer wrote.
Most legal teams have not priced this risk, and the reason is understandable. The models are right most of the time, and a fast, plausible answer is hard to resist. Three things are arriving at once:
Teams that can show why an answer came out the way it did will be in a very different position from teams that can only show who had access.
Take these five questions into your next vendor meeting. Under each is what a defensible answer sounds like. Any vendor serious about legal work should be able to answer all five, and that includes us.
Neota Logic has spent more than fifteen years encoding legal reasoning as rules, and the last several adding AI to that foundation. The approach comes down to three commitments.
The language model does the work it suits: reading documents, summarizing, and pulling out the facts. Within a workflow, an AI task node calls your own registered models for tasks such as extraction, summarization and classification, then returns the output into the workflow, where your rules, validation logic and human sign-off apply exactly as you have defined them.
The conclusion is reached under logic a lawyer wrote and can read, versioned like any other logic, so a change shows exactly what changed, when, and by whom.
In one controlled internal test, 50 of 50 scenarios passed 250 of 250 checks. That figure describes a specific test, and it illustrates the property that matters: same inputs, same output, every time.
Each AI call is recorded in the session history, with the model used and the tokens in and out, and human approval steps sit at the points where the judgment is. Site administrators register and approve which models are available at all, so credentials stay centralized and only approved models ever reach a workflow. No AI model is invoked by default.
To be clear about the boundary, because the market blurs it: Neota Logic governs the workflow and the answer. It doesn't govern the underlying AI model.
Every new model release is pitched on capability: more fluent, higher on the benchmark, better at the bar exam. Capability is real and it is welcome. It is also not the test.
The market is already recalibrating: in Law360's 2026 survey, 44 percent of attorneys who frequently use AI at firms now say they see both the pros and cons of adoption, where a year earlier 73 percent of frequent users held a positive view.10
The test is whether the answer you relied on in January is the answer the system gives in July, and whether you could show anyone why.
See how this applies for in-house legal teams and law firms.