Confidence scores serve the buyer, not the lawyer
ChatGPT scores its own confidence at 9.4 out of 10 on legal reasoning questions, yet its high-confidence error rate is 6.7%. Meta AI's high-confidence error rate is 31.7%, and Perplexity AI's is 15.0% (arxiv 2608.21089). A system that believes it is almost always right while being wrong nearly a third of the time has a calibration problem. In contract review, that calibration problem has a specific beneficiary.
The confidence score answers the buyer's question. A general counsel evaluating AI contract review tools needs a number for the shortlist, and the score survives procurement. The LegalOn 2026 Contract Review Benchmark found that the most capable foundation models produce confident-sounding answers on provisions carrying real legal and financial exposure, and those answers are frequently wrong. General-purpose models often identified the right topic but missed the legal standard.
The buyer got a metric. The metric was miscalibrated. The user, the attorney who has to stand behind the output, needs something different. Research on AI-assisted decision-making found that miscalibrated confidence impairs appropriate reliance and reduces AI-assisted decision-making efficacy, and that miscalibration is difficult for users to detect (arxiv 2402.07632).
Here is the disclosure paradox. Telling users that the confidence score is unreliable does not help. The same research found that communicating calibration problems decreases users' trust in uncalibrated AI, leading to high under-reliance, and does not improve decision efficacy (arxiv 2402.07632). Miscalibration harms the user whether they know about it or not: disclosed, it causes under-reliance; undisclosed, it causes over-reliance.
The adoption numbers draw the line. Document review has reached 77% adoption, legal research 74%, contract drafting assistance 61%, and due diligence support 58% (AI in Legal Industry Statistics 2026). Adoption falls as the stakes of a wrong answer rise.
Accuracy and hallucination concerns top adoption barriers at 74.7%. Measured hallucination rates across legal AI tools bear this out: Lexis+ AI at 17%, Westlaw AI at 33%, GPT-4 on legal queries at 43% (AI in Legal Industry Statistics 2026). Lawyers who distrust these outputs are reading the evidence correctly.
Harvey's own account of AI contract review describes what happened next: adoption stalled in most firms and in-house teams. Their diagnosis is specific. "A flag without a source is just a guess." The fix Harvey proposes is structural: the reviewer accepts, rejects, or refines redlines in tracked changes, and the agent learns from those decisions on the next run. The confidence score is replaced by a feedback loop calibrated to the person doing the work.
The regulator public has barely entered this conversation. Only 9% of firms enforce any AI policy. Fifty-four percent offer no AI training and have no plans to (8am 2026 Legal Industry Report). When users name what makes a tool trustworthy, their answers skip the confidence score entirely: the provider understanding firm workflows at 47%, the provider understanding ethical obligations at 46% (8am 2026 Legal Industry Report). Attorneys are naming the gap between what the buyer evaluated and what the user needs.
Vaquill's synthesis of lawyer sentiment on Reddit and review platforms captures the daily reality: users describe uploading the contract, reading the AI summary, then reading the whole contract anyway to verify. Lawyers keep reporting that tools clear procurement but never enter actual workflow.
This is a four publics problem. The buyer got the metric they needed. The user got a metric designed for someone else's decision. The regulator got no governance framework. The ecosystem, the client on the other side of the contract, receives the output of a process where the quality signal was built for a purchasing committee.
The confidence score will remain decorative until it is rebuilt for the person who signs the opinion letter. Every point of self-reported confidence that exceeds actual accuracy is a liability assigned to the wrong public.








.webp)
.webp)
.webp)

