← All articles

Prompting vs Governed Workflows: Why Clever Prompts Still Get Lawyers Sanctioned

Katie Pham
·
July 28, 2026

More than 1,800 court and tribunal decisions worldwide have now dealt with AI-fabricated content in legal filings, and the count climbs almost daily. Some have ended in fines and referrals to the bar. The tools behind the fake citations did not malfunction: they did what they were asked and invented the supporting authority with complete confidence.

None of this has slowed adoption, nor should it. Ninety-two percent of legal professionals now use AI daily, according to Wolters Kluwer's 2026 Future Ready Lawyer Survey, yet only 31 percent feel prepared on information security and governance. The question is no longer whether to use AI. It is whether a team can stand behind what it produced when a regulator, a client, or a court asks how the answer was reached.

AI governance for legal teams is not about restricting AI. It is the discipline of making every AI-assisted decision traceable, reviewable, and defensible by controlling the workflow around the model. The problem was never AI itself. It is AI running without that discipline, and most legal AI today runs on prompt logic instead. Closing the gap does not mean using AI less. It means wrapping the AI you already rely on in governed workflows.

What we mean by prompt logic

Prompt logic is the model most legal teams have fallen into without naming it. A lawyer opens a chat window and types a request. An application author writes a clever system prompt and points it at a document. In both cases, reliability is treated as a prompting problem: the belief that a good enough instruction produces a good enough answer.

The instinct is understandable, but a prompt is a single instruction to a probabilistic system, with three consequences that matter in a regulated setting. The same prompt can return different answers on different runs. Nothing is recorded by default: not the model that responded, the sources it used, its confidence, or who reviewed the output. And there is no enforced checkpoint, so whether a human reviews the result before it reaches a client rests on the discipline of the person at the keyboard, under deadline.

Prompt logic optimizes the answer. It does nothing for the accountability around the answer.

The evidence that better prompts are not enough

It is tempting to believe purpose-built legal tools have engineered the risk away. The evidence says otherwise. In the first preregistered empirical study of commercial legal research tools, published in the Journal of Empirical Legal Studies in 2025, Stanford researchers found the leading products still hallucinated 17 to 33 percent of the time: more than 17 percent for Lexis+ AI, around a third for Westlaw's AI-Assisted Research, and about 43 percent for general-purpose GPT-4. Vendors had marketed these retrieval-based tools as eliminating hallucinations. The researchers concluded those claims were overstated.

The courtroom consequences are well documented. The first widely reported sanction came in Mata v. Avianca in 2023, when two lawyers were sanctioned after their brief cited six cases that did not exist. The database tracking those filings, maintained by researcher Damien Charlotin, logged more than 700 such decisions by late 2025 alone, with individual sanctions since climbing above 55,000 dollars. And the pattern is not about model quality: in Dehghani v. Castro, an attorney filed a brief bought from a freelancer without reviewing it and faced a fine, mandatory training, and a duty to self-report to the bar. The failure was a missing review step and no record of one.

None of this is an argument against AI. It is an argument for the guardrails around it. Better prompts and better models lower the error rate. They do not create the checkpoint or the audit trail that make an output defensible.

What a governed AI workflow actually does

A governed AI workflow treats each AI call as one logged step inside a deterministic process, with rules, validation, confidence thresholds, and human sign-off at defined points. This is legal AI orchestration in practice: the probabilistic model does what it is good at, such as extraction or summarization, while a deterministic workflow decides what happens to the output next.

The mechanics create the defensibility. Every call records the model used, the prompt, the sources, the tokens, a confidence score, the reviewer, and the time of sign-off. Confidence becomes a routing variable rather than a hidden number: low-risk outputs proceed, uncertain ones escalate automatically to a named human. Review is not left to discipline; it is a step in the process, and by design the last one.

The distinction is simple. The output of prompt logic is an answer. The output of a governed AI workflow is an answer plus the record that makes it defensible.

AI risk in contract review: the two models side by side

Contract review is where the difference becomes concrete, and where AI risk in contract review is most acute, because outputs move quickly toward negotiation and signature.

Under prompt logic, a lawyer pastes a draft into a chat interface, asks the model to flag onerous clauses, and works from whatever comes back. If it misses a warranty gap, nothing catches it, and there is no record of what was checked.

In a governed AI workflow, the same review runs as a sequence: an extraction step pulls the clauses, dates, and jurisdictions into a structured record; a validation step tests them against the firm's playbook; and confidence-based routing approves low-risk contracts automatically while escalating anything material to the right lawyer for sign-off. Every clause, assessment, and decision is stored against the matter with a full audit trail. The same logic underpins compliance workflow automation, from data processing agreement review to classifying a system under the EU AI Act, and legal intake automation, where a request is classified, routed to the right team, and logged before a lawyer has read a word.

Same model, same capability. The governed version produces the same work with a record you can defend.

Legal AI ROI measurement: where the value actually lands

The returns from AI are real. Wolters Kluwer's 2026 survey found most respondents reported weekly time savings of 6 to 20 percent, and roughly a third attributed an 11 to 20 percent revenue increase directly to AI. But honest legal AI ROI measurement has to account for the cost prompt logic hides: verification. When there is no record of what the model did, someone re-checks everything by hand, and the more the tool produces, the longer that takes. Time saved in generation is quietly spent again in review. Worse, a single sanctioned filing, with its fine, bar referral, and reputational damage, can erase the efficiency gains of an entire practice group.

Governed AI workflows change the arithmetic, because the checking and the record are part of the process rather than a separate manual chore. The efficiency is durable, and the audit trail that protects against the tail risk is produced automatically. Returns you can keep are worth more than returns you re-earn on every matter.

How to build a legal AI governance framework

For teams working out how to govern AI in a legal department, the shift is less about policy documents and more about making a few controls the default. A workable legal AI governance framework rests on five principles:

  1. Approved models, centrally registered, so no one is pasting API keys into a chat tool.
  2. Every call logged: model, prompt, sources, tokens, and confidence, so any output can be reconstructed later.
  3. Confidence as a routing variable, so uncertain outputs escalate instead of passing silently into a deliverable.
  4. Human review at defined points, always last, attributed to a named reviewer at a recorded time.
  5. A full audit trail per matter, tied to the matter it belongs to.

The pattern is the same across all five: accountability is designed into the process, not left to the person under deadline. That is the practical meaning of AI governance for legal teams, and it is the standard Neota Logic was founded on, that every automated legal decision is tied to a rule, a check, or an approval.

From prompt to process

Prompting is a genuine skill, and prompts will keep improving. But a better prompt improves the answer, not the accountability around it. No prompt records which model responded, scores its confidence, routes uncertain cases to a human, or produces the audit trail a regulator will ask for.

The goal is not less AI. It is AI you can defend. For legal teams, the unit that matters is not the prompt but the governed AI workflow around it. That is where a probabilistic model becomes a defensible output, and where AI stops being a source of exposure and becomes infrastructure you can stand behind. Every decision, defensible.

Sources

Ready to make your AI workflows defensible?

Book a demo and we'll walk one of your real processes through Neota.

Book demo