CommentaryThe legal news is crowded with stories about generative AI. Judges impose sanctions for hallucinated citations and fabricated case law. Courts add new standing orders restricting how lawyers and litigants may use these tools. A profession built on the accuracy and grace of its written word is surrendering that core capability to machines. The volume of pro se filings, now often drafted with these generative tools, is growing rapidly. Little wonder that many in the legal community have formed a grim view of AI.
But that impression is drawn from one side of the technology. AI writes, but it also reads. Reading at scale is necessary to assess the rule of law accurately. Claims about fairness, equality, transparency, accountability, predictability, impartiality, and legitimacy concern how the system behaves across the breadth of cases, rather than in a single case. Historically, measuring that behavior was prohibitively expensive.
Consider a straightforward question about equality: do two defendants charged with the same offense, with similar records, receive the same bail decision? The detail needed for that comparison can lie buried in individual case files. Until recently, turning those files into an analyzable dataset would have required reading thousands of them and pulling out dozens of data points from each — the charge, the prior record, the conditions imposed, and everything else that might legitimately explain a difference. That meant an enormous investment of human effort, often with painstaking manual data entry. The cost limited how broadly and how frequently such studies were conducted. So instead we often formed our beliefs about the rule of law from personal experience, selected readings, and general impressions.
That has changed. Machines can now read legal documents at a cost that makes the work ordinary rather than heroic, a new capability built on decades of hard-won innovation. Those advances began with the page itself: computer scientists developed methods for separating squiggles from smudges, finding the individual letters, the start and end of a word, the punctuation, the line breaks. With text in digital form, programmers could specify the patterns to find. Consider the task of locating every reference to Rule of Professional Conduct 8.3 in a disciplinary document. Ask a program to find the characters 8.3 and the results come back cluttered with a time entry for 18.3 hours and a restitution payment of $1,048.32. Sharpening the instruction with computerized language — “\b8\.3\b” — excludes those false matches.
The next hard problem was meaning. An early technique, aptly called “Bag of Words,” simply enumerated the words in a sentence. But “the dog chases the car” and “the car chases the dog” use identical words and describe very different scenes. In contrast, “the dog chases the car” and “the canine pursues the automobile” share no key words and mean nearly the same thing. Counting words while ignoring their order and meaning was inadequate.
One important advance was to let programs learn a numerical representation of each word from its use in large bodies of text. Words used in similar ways landed near each other. Canine came to sit close to dog, automobile close to car. Different words could now reveal their resemblance through numbers.
Knowing that dog resembles canine, however, does not tell the machine who is chasing and who is being chased. That requires attention to the words’ positions and relationships. Another advance, appropriately called “attention,” allowed models to learn which other words matter when interpreting each word. The numerical representation of a word could change with its context. These representations form a mathematical portrait of the text, helping a trained model identify actions, participants, and relationships.
Many of us unwittingly rely on this kind of machine reading every day. When you type “I hope you have a” and your phone suggests “great day,” your phone has already read what you wrote.
This capability is not entirely new to the law. For decades, lawyers have worked with tools that read text — Boolean searches on Westlaw, with their miserable strings of ANDs and ORs and proximity operators, later softened by natural-language queries. More recently, e-discovery platforms have used predictive coding to sort millions of documents for responsiveness, and analytics products have tracked how particular judges rule on particular motions. These are real capabilities, and they work.
But these commercial tools were chiefly designed to help lawyers handle particular matters. The focus was this case, this judge, this motion. The commercial incentive was to help clients with their cases; the public benefit of measuring the system was harder to finance. Researchers built datasets to study the system as a whole, but assembling the records and extracting the information needed to compare them demanded substantial resources. The cost limited the work’s reach. Recent advances in large language models are changing those economics, allowing researchers to adapt existing models to questions that once demanded extensive manual work or specialized software.
Reading across the breadth of cases is necessary to assess the rule of law in practice. Return to the straightforward question about equality. Imagine collecting the electronic records of 10,000 recent cases that reached a pretrial bail decision. AI can turn that pile into a dataset — extracting the charges, recorded criminal histories, bail decisions, and conditions imposed into a consistent set of fields. After checking the extracted data against the source records, we can begin comparing outcomes. Do similar defendants receive similar conditions? Does the answer change by courthouse, by judge, or according to whether the defendant had counsel at the hearing?
So much of the legal system is unmeasured that the list of unanswered questions seems endless. We could review the outcomes of small claims cases and how they vary depending on whether the parties are individuals or businesses. We could study how long divorce cases take, and how that time varies by court location and judge. We could compare eviction outcomes for tenants with and without representation. We could examine how sentencing decisions vary among otherwise similar cases involving the same mitigating factor. For each dataset, we could ask a range of questions, testing our beliefs about the rule of law against what the records actually show.
The legal community can finally catch up to other domains. Today we can look up the average delay on a flight route between two American cities and check whether a particular flight this afternoon is running late. We can check whether a purchase made seconds ago appears in our online account. We can research a car model’s reliability or use a VIN to retrieve one vehicle’s history from hundreds of millions of records.
But I cannot readily obtain a comparable picture of the outcomes of small claims cases, the pace of divorce cases, how eviction outcomes vary with representation, or the role of a mitigating factor in sentencing. The cost of reading those records has fallen sharply. Access to them remains uneven. I encountered this recently when I requested electronic case files to study divorce cases across one state. The court system refused to release any electronic files. Although some courts release records in bulk, others still provide them only one case at a time. How readily a court permits systematic scrutiny of its work is itself a measure of transparency under the rule of law.
To put this approach to a practical test, we needed a substantial body of publicly available records. We found one, surprisingly, in attorney discipline. Many states publish disciplinary decisions and summaries reaching back decades. For Massachusetts, we assembled roughly 4,000 documents covering 24 years — a substantial public record of how the profession holds its members accountable. Using AI, we organized that record into a dataset, extracting the same categories of information from each document: findings, references to professional conduct rules, and sanctions imposed.
Before starting the analysis, we articulated our expectations. Lawyers occupy positions of trust. They owe honesty to courts and clients, and they have obligations to report serious misconduct by other lawyers. These responsibilities seem fundamental to the rule of law. We expected candor to figure prominently in the disciplinary record. That expectation gave us an initial question to investigate; the dataset gave us a way to pursue it.
Perhaps our beliefs were wrong. Perhaps other obligations would dominate the published record. To investigate, we asked a simple descriptive question: which professional conduct rules appeared most often in the documents? Counting references gave us an initial picture, with the findings and circumstances of individual cases available for more detailed analysis.
In Massachusetts, the Rules of Professional Conduct are organized into eight sections. The eighth, “Maintaining the Integrity of the Profession,” covers a broad range of conduct: criminal acts bearing on fitness to practice, dishonesty, failures to report serious professional misconduct, and violations of other professional conduct rules. Every violation linked to one of the first seven sections could trigger an additional violation of section eight. Given that breadth and overlap, we set it aside for this initial comparison.
The first seven sections address these aspects of legal practice:
- Client-lawyer relationship — competence, communication, confidentiality, fees, conflicts, and safeguarding client property.
- Counselor — advising clients, evaluating matters, and serving as a neutral.
- Advocate — candor, fairness, and lawyers’ conduct in legal proceedings.
- Transactions with persons other than clients — truthfulness, communication, and respect for the rights of others.
- Law firms and associations — supervision, professional independence, and the organization of legal practice.
- Public service — pro bono service, court appointments, and participation in legal services and law reform.
- Information about legal services — advertising, solicitation, and communications about lawyers’ services.
We counted every reference to a Rule of Professional Conduct in the published documents. After setting aside references to section eight, an even distribution across the remaining seven sections would give each about 14 percent of the total.
That leaves us with two questions. Write down your answers before seeing the results. Which of these seven sections do you predict accounts for the most references? And what share of the total does it represent?
Perhaps for the first time, you are putting a belief about the legal system to a numerical test — committing to a prediction before comparing it with the published record. This is where an empirical assessment of the rule of law begins.
The answer comes in the next article.*
David Stasior, MD, MPP, is a principal at Broken Scales, where he applies data analysis to the justice system, and the author of Slot Machine Justice.
*This is the first part of a planned four-part series, with new installments each Friday.