Having spent considerable time evaluating the performance of various monitoring and observability tools through rigorous benchmarking, I decided to apply a similar methodological approach to a different domain. I conducted a controlled experiment to assess ChatGPT's (specifically GPT-4) accuracy in reviewing common legal clauses, pitting it against the work of a certified paralegal with five years of experience in corporate contract review. The goal was not to declare a winner in an absolute sense, but to quantify the current capabilities and failure modes of the LLM in a high-stakes, detail-oriented task.
The benchmark was structured as follows. I curated a set of 25 non-disclosure agreement (NDA) clauses, each containing one intentional, substantive issue (e.g., overbroad confidentiality definitions, missing injunctive relief language, unreasonable survival periods) and several stylistic or minor grammatical issues. The paralegal and the LLM were provided identical instructions: identify all substantive legal issues, flag ambiguous language, and note any procedural deficiencies. Scoring was binary for substantive issues (correct identification/incorrect or missed) and tracked separately for minor issues.
The results were illuminating and, in some aspects, counter to prevailing hype:
* **Substantive Issue Recall:** ChatGPT successfully identified 19 out of 25 core substantive issues (76% recall). The paralegal identified 24 (96% recall).
* **Precision & Hallucination:** Of the 28 issues flagged by ChatGPT as "substantive," 9 were incorrect or were hallucinations of problems not present in the text (68% precision). The paralegal flagged 26 issues, with 2 being overly cautious interpretations (92% precision).
* **Contextual Understanding:** The model consistently struggled with the interplay between clauses. For example, it correctly flagged an indefinite survival period for confidentiality but failed to connect that to a separate clause that would make the obligation practically perpetual through a "residuals" clause. The paralegal identified this linkage immediately.
* **Minor Issues:** ChatGPT excelled at identifying typographical errors, inconsistent numbering, and passive voice, outperforming the human reviewer in this mechanical aspect.
The critical failure modes observed were not simple misses, but dangerous overconfidence. In three instances, the LLM fabricated specific case law citations to support its incorrect analysis. Furthermore, its performance degraded noticeably when the clause language deviated from standard patterns, whereas the paralegal's experience allowed for robust interpretation of novel phrasing.
This benchmark suggests that while LLMs like ChatGPT can function as a potent preliminary scanner for surface-level inconsistencies and boilerplate review, they lack the relational reasoning and deep contextual awareness required for reliable, autonomous legal analysis. The hallucination problem presents a profound risk. The optimal workflow, analogous to using synthetic monitors to triage potential incidents for human engineers, appears to be using the LLM as a high-speed first pass, with every single output requiring rigorous validation by a qualified human. The cost/accuracy trade-off is significant.
— Billy
I'm an SRE at a mid-size fintech, we run about 300 services on Kubernetes, and our whole observability stack (Prometheus, Loki, Tempo, Grafana) is on-prem.
* **Target user fit** - GPT-4 for this is a fast, low-cost first pass for non-critical internal documents. The paralegal is for any contract with material liability, external partners, or regulatory exposure. It's the difference between a pre-flight checklist and the actual pilot.
* **Real cost & hidden lift** - GPT-4 via API is about $0.03-$0.12 per document for this task, while a paralegal runs $30-$80/hr. The hidden cost for the LLM is the legal review *you still must do* on its output and the prompt engineering time to get consistent formats. The human's hidden cost is the time spent on simple clauses that don't need their expertise.
* **Failure modes** - The LLM will miss nuanced cross-clause implications and lack judgment on what's "market standard" for your industry. It will confidently hallucinate incorrect legal citations if your prompt is vague. The paralegal's limitation is throughput and potential for human fatigue on high-volume, repetitive clauses.
* **Deployment & iteration speed** - You can have the LLM reviewing a clause in an afternoon via the API and iterating prompts weekly. Integrating a paralegal into a workflow means process docs, scheduling, and longer feedback loops if their review guidelines need adjustment.
I'd use the paralegal for any final agreement that leaves the company. Use GPT-4 as a standardized clause scrubber for early drafts of internal policies or vendor MSAs to free up human time for the hard parts. To decide properly, we'd need to know your annual contract volume and your internal legal team's capacity for oversight.