Whitepaper · A buyer's framework for defensible legal AI

Same facts,
different answer

Consistency is the legal AI risk that gets discussed least: the same facts producing a different answer six months apart, with nothing in place to catch it.

Shaz Aziz, Head of Client Solutions, Neota Logic · Oct 7, 2026

Scroll to begin
Executive summary

In 2025, researchers put 500 real federal appeals to leading models 20 times each, with identical text and the models set to be as consistent as their settings allow. On up to half of the questions, the same model named one party the winner, then the other.1 A legal expert reviewed pairs of the contradictory answers and found the analysis sound in every one. The model simply weighed the arguments differently each time.

Legal work has always tolerated a version of this problem in people, because a lawyer's inconsistency sits inside an accountability structure: review, a file, insurance, a regulator. A model's drift has no equivalent. The major platforms spent 2026 shipping the first half of that structure, with audit logs, admin controls, and ethical walls that show who had access, what was configured, and increasingly what was said. None of it shows whether the answer held over time, or why it came out the way it did.

This paper sets out the evidence on why AI answers move, why human review cannot make them consistent, and why the accountability question is arriving on a regulatory clock rather than a procurement one. It closes with five questions to take into any vendor meeting, ours included, and what a defensible answer to each one sounds like.

01

Same question, twenty times

In 2025, researchers at the University of Maryland and Johns Hopkins ran a simple experiment on leading AI models. Anyone buying legal AI should know the result.

The experiment
Input · 1 of 500

A real US federal appeal

A statement of facts, plus the best arguments for each side. The question: who should win?
Party A
vs
Party B
The judges themselves split two to one
0 / 20 runs
Party A winsParty B wins
0said Party A
0said Party B
✓ A legal expert reviewed pairs of the contradictory answers. The analysis was sound in every one.
Illustrative sequence for one case. Identical text, settings tuned for consistency.

Share of the 500 questions where the same model named both parties the winner across 20 identical runs

Claude 3.50%
GPT-4o0%
Gemini 1.50%
Source: Blair-Stanek and Van Durme, 2025
The reasoning was sound.
The answer wasn't repeatable.
  • Late-2024 models, on deliberately hard questions, so read the percentages as indicative, not universal.
  • o1 flipped on 28% of a 50-question subset, though it can't be set to temperature zero, so it isn't like-for-like.
  • What matters is that it happens at all, and that nothing in a standard deployment would surface it.

The researchers took 500 real US federal appeals, each one a case where the judges themselves had split two to one.

Each case was reduced to a statement of facts plus the best arguments for each side. Then they asked each model a single question: which party should win?

They asked it 20 times per case. The text was identical every time, and the models were set to be as consistent as their settings allow.

On a large share of cases, the same model named one party the winner, then the other.

A legal expert reviewed pairs of the contradictory answers and found the analysis sound in every one. The model simply weighed the arguments differently each time.1

Across all 500 cases, here is how often each model flipped. For Gemini 1.5, half the questions got both answers.

Accuracy gets checked in January, at the pilot and the sign-off. Consistency is what you need to know about in July, when the same clause comes back with a different answer and nobody is comparing.

Two different things move the answer

Mechanism 1

Drift

The model changes underneath you. GPT-4's accuracy at identifying prime numbers fell from 84% to 51% between March and June 2023. Same product, same name, and no release note.2

March 2023June 202384%51%
Mechanism 2

Server-side nondeterminism

One prompt, run 1,000 times at temperature zero, gave 80 different completions. The cause was how the inference server batches requests together. With that fixed, all 1,000 came back identical.3

Their research shows it can be fixed, but only by whoever runs the model. Whether you get the same answer twice is an engineering choice made outside your organization.

0runs
0distinct answers

Newer models show it too. A 2026 study had five models, including Claude Sonnet 4.5, Claude Haiku 4.5 and Gemini 2.5 Flash, score real enterprise question and answer pairs, and found substantial variability in the scores even at temperature zero.4 Those are the kinds of scores enterprise pipelines use to route, triage and gate work, which is the job legal workflows give AI outputs.

Accuracy is what gets checked in January, at the benchmark, the pilot, and the sign-off. Consistency is what you need to know about in July, when the same contract clause comes back with a different escalation decision and nobody is comparing. Princeton researchers Narayanan and Kapoor call this the capability and reliability gap, and consistency is the first reliability dimension they measure.6

Every vendor benchmark you've seen is a snapshot of one moment.

OpenAI's headline for Astra for Law, the newest legal launch in the market, is 54 percent on a legal research benchmark, on its own testing.5 A good number, and a snapshot. Legal work runs on the months in between.
02

Lawyers are inconsistent too, and the comparison fails anyway

The strongest objection to the consistency argument is that lawyers aren't consistent either. Ask twenty lawyers a hard question and you'll get a spectrum of answers. That's true.

The difference is structural

A lawyer's advice

1Produced by an associate
2Reviewed by a partner
3Recorded on the file
4Backed by malpractice insurance, bar oversight, and a license worth protecting
The system is imperfect, but errors get found, priced, and sometimes corrected.

A model's answer

1Produced by the model
?Nobody sees it change
?No record of why it came out the way it did
?No license at risk
A model's answer has no equivalent path.

Human review is often offered as the fix, and it does real work: a reviewer can catch an answer that is wrong today. What a reviewer cannot do is make answers consistent, because review vouches for one answer at one moment.

January
Same clause, same facts
Escalate to legal
✓ Reviewed and signed off
three months later
April
Same clause, same facts
Approve
No reviewer in the room

When the same clause produces a different escalation decision three months later, the January sign-off says nothing about it, and the reviewer who approved the first answer is not in the room for the second. A reviewer can only review what they can see, and drift is precisely what nobody is shown.

The lawyer leaves a paper trail.
The model usually does not.

Already on the public record
0US court decisions

The scale of unchecked AI output is already on the public record. By mid-2026, Damien Charlotin's AI Hallucination Cases database listed more than 1,100 US court decisions dealing with filings that contained fabricated AI content.7 Every one of those filings had someone's name on it.

Click any solid green dot to open one of the cases.

Each dot is one decision.

03

The half-built governance layer

A year ago it was fair to say nobody had shipped a governance layer for legal AI. That is no longer accurate. Over the past twelve months the major platforms have been racing to add exactly the controls enterprise legal buyers asked for.

Keep reading

The half-built governance layer, a three-tier way to sort your AI stack, the regulatory clock, and five questions to put to any vendor

The rest of the paper looks at what the market has shipped and what is still missing, sorts legal AI by what actually produces the conclusion, sets out why accountability is arriving on a regulatory clock, and closes with five questions to take into any vendor meeting, with what a defensible answer to each sounds like.

03The half-built governance layer2 min
04Three tiers of reliability1 min
05A risk arriving on a regulatory clock2 min
06Five questions for any vendor, including us2 min
07How Neota Logic has approached it2 min

Read the full paper

Get instant access

We'll only use your details to show relevant Neota updates. Unsubscribe anytime.

Shipped in 2026
Who had access
What was configured
Increasingly, what was said
Still missing
Did the answer hold over time?
Why did it come out the way it did?
Anthropic shipped a Compliance API that pulls activity logs, chat data, and files from its enterprise product.
Google launched Gemini Enterprise for Legal in preview, with a centralized control plane for IT and risk teams, audit logging, and ethical walls inherited from document and matter management systems through its connectors.
Harvey integrated with Intapp Walls so ethical wall policies sync into its access controls.
OpenAI is working with Latham & Watkins to design permissions, ethical walls, and firm oversight into Astra for Law.

All of this answers three questions: who had access, what was configured, and increasingly what was said. Those are necessary answers, and legal teams should welcome them. They are also only half of the layer. An access log tells you an associate queried the system on a Tuesday. It can't tell you whether the same query, a quarter later, came back different.

0%
of firms using or exploring generative AI (ILTA 2026)
0%
of firms using Harvey are still piloting it (ILTA 2026)8
0% to 0%
in-house generative AI use in one year (ACC / Everlaw)9

Firms are trialing nearly everything and committing to very little. On the corporate side, the ACC and Everlaw found in-house use of generative AI more than doubled in a year, from 23 percent to 52 percent, so the exposure now sits inside the legal department as much as at outside counsel. Pilots often stall because nothing makes the output dependable enough to stand behind in front of a client, and no volume of access logging changes that.

04

Three tiers of reliability

Not every use of AI in legal work carries the same consistency exposure. The useful way to sort a stack is by what actually produces the conclusion.

February 2026

A regulator has already approved the rules half of this. In February 2026 the Solicitors Regulation Authority authorized LawFairy, which describes itself as the first deterministic, technology-only law firm in England and Wales, starting with UK immigration.12 Statutory and policy criteria are encoded as decision pathways validated by lawyers, no generative AI touches the regulated outcome, and the same facts produce the same result, in deliberate contrast to the probabilistic systems the rest of the market runs on.

The open question for everyone else is how to add a language model's fluency without giving that up. It comes down to architecture.

The model and the rules each do the work they suit,
and neither does the other's job.

05

A risk arriving on a regulatory clock

The accountability question is no longer theoretical, and it is not waiting for procurement cycles.

0days until the EU transposition deadline
Jan 2025
FTC final order against DoNotPay, with $193,000 in monetary relief, for marketing an AI legal service on untested claims.11
Enforcement
Today
Dec 9, 2026
Deadline for EU member states to transpose the revised Product Liability Directive, bringing AI systems inside the legal definition of a product.
Product liability

Regulators on both sides of the Atlantic are heading the same way:
if you put an AI system into legal work, you answer for what it does.

Meanwhile there is no agreed standard for autonomy in legal AI, nothing like the SAE levels that define who is driving at each stage of a self-driving car, and that lawmakers use to decide who is liable. Legal AI today works like driver-assist, and with driver-assist the driver is still liable. That puts the liability where the review sits, on the lawyer who approved the output. A reviewer can't sit down and check the reasoning inside a model. They can check logic a lawyer wrote.

Most legal teams have not priced this risk, and the reason is understandable. The models are right most of the time, and a fast, plausible answer is hard to resist. Three things are arriving at once:

Product liability for AI software
Enforcement against untested claims
No autonomy standard to share out responsibility

Teams that can show why an answer came out the way it did will be in a very different position from teams that can only show who had access.

06

Five questions for any vendor, including us

Take these five questions into your next vendor meeting. Under each is what a defensible answer sounds like. Any vendor serious about legal work should be able to answer all five, and that includes us.

1
Did the answer hold six months later?
Hide
What a defensible answer sounds like
Accuracy gets checked once, at the pilot and the sign-off. Ask to see the same question asked months apart, and what came back each time. A vendor who has never run that comparison is telling you consistency has never been measured.
2
What exactly is versioned: the model, the prompt, or the rules?
What good sounds like
What a defensible answer sounds like
The model is the provider's to change and retire, whatever you have deployed on top of it. The prompt is yours to version, and worth versioning, but it controls the wording of the question rather than how the system behaves. The rules control the outcome, and a change to one shows what changed, when, and who changed it. A vendor who versions only prompts has a record of the wording and none of the decision.
3
When an answer changes, where is the record of why?
What good sounds like
What a defensible answer sounds like
Audit logs tell you who, what, and when. For a legal decision you also need to know why this answer came out this way, and to be able to reproduce it. If the reasoning lives inside the model, that record does not exist.
4
Can I inspect the reasoning, or only the output?
What good sounds like
What a defensible answer sounds like
The defensible shape lets the model gather the facts and lets versioned rules make the call. Same inputs, same output, and a path through the logic you can hand to a colleague, a client, or a regulator to check. If all you can inspect is the output, you're taking it on trust.
5
When it drifts, who is accountable: you or us?
What good sounds like
What a defensible answer sounds like
Watch for the pause. Vendors are quick to claim capability and slow to accept liability for behavior that changes over time. Given where product liability law is heading, the answer to this question will matter more every quarter.
Score a vendor
0 / 5
answered well
Mark each answer as you go.
Print the checklistReset
07

How Neota Logic has approached it

Neota Logic has spent more than fifteen years encoding legal reasoning as rules, and the last several adding AI to that foundation. The approach comes down to three commitments.

1

Language to the model

The language model does the work it suits: reading documents, summarizing, and pulling out the facts. Within a workflow, an AI task node calls your own registered models for tasks such as extraction, summarization and classification, then returns the output into the workflow, where your rules, validation logic and human sign-off apply exactly as you have defined them.

→
2

Decisions to the rules

The conclusion is reached under logic a lawyer wrote and can read, versioned like any other logic, so a change shows exactly what changed, when, and by whom.

0/50 scenarios · 0/250 checks

In one controlled internal test, 50 of 50 scenarios passed 250 of 250 checks. That figure describes a specific test, and it illustrates the property that matters: same inputs, same output, every time.

→
3

Keep the record

Each AI call is recorded in the session history, with the model used and the tokens in and out, and human approval steps sit at the points where the judgment is. Site administrators register and approve which models are available at all, so credentials stay centralized and only approved models ever reach a workflow. No AI model is invoked by default.

To be clear about the boundary, because the market blurs it: Neota Logic governs the workflow and the answer. It doesn't govern the underlying AI model.

The provider's
The model
The model remains the provider's, versioned and retired on their schedule.
Under your control, on your record
What sits under your control, and on your record, is everything that turns a model's output into a legal decision:
The rules that decide
The approvals that vouch
The record that shows how

The test that matters

Every new model release is pitched on capability: more fluent, higher on the benchmark, better at the bar exam. Capability is real and it is welcome. It is also not the test.

The market is already recalibrating: in Law360's 2026 survey, 44 percent of attorneys who frequently use AI at firms now say they see both the pros and cons of adoption, where a year earlier 73 percent of frequent users held a positive view.10

A year earlier
0%
of frequent users held a positive view
Law360, 2026
0%
of attorneys who frequently use AI at firms now say they see both the pros and cons of adoption

The test is whether the answer you relied on in January is the answer the system gives in July, and whether you could show anyone why.

An LLM gives you a confident answer. A rules engine gives you a defensible one.
Legal AI needs both.
Request a demo
Sources
1. Blair-Stanek and Van Durme, "LLMs Provide Unstable Answers to Legal Questions," University of Maryland and Johns Hopkins, arXiv:2502.05196, January 2025
2. Chen, Zaharia and Zou, "How Is ChatGPT's Behavior Changing over Time?", Stanford and Berkeley, 2023
3. Thinking Machines, "Defeating Nondeterminism in LLM Inference," 2025
4. Lau, "Same Input, Different Scores: A Multi Model Study on the Inconsistency of LLM Judge," 2026
5. OpenAI, "Introducing Astra for Law," September 17, 2026
6. Rabanser, Kapoor and Narayanan, "Towards a Science of AI Agent Reliability," Princeton, 2026
7. Damien Charlotin, AI Hallucination Cases database
8. ILTA 2026 Technology Survey
9. ACC and Everlaw, "Generative AI's Growing Strategic Value for Corporate Law Departments," 2025
10. Law360 Pulse, "The 2026 AI Survey," March 2026
11. EU Product Liability Directive 2024/2853; FTC final order against DoNotPay, January 2025
12. Artificial Lawyer, "LawFairy 'Technology-Only Law Firm' Gets Regulatory Approval," February 24, 2026