ChatGPT Invented the Witnesses: New Mexico's Top Court Fines a Veteran Lawyer $5,000 Over an AI-Fabricated Murder Appeal
A 40-year Santa Fe defense attorney is held in contempt and removed from a murder appeal after ChatGPT invented police testimony and witnesses that never existed — the sharpest court rebuke yet of AI hallucinations in criminal litigation.
For years, judges have warned that AI-generated legal filings can fabricate case citations. The New Mexico Supreme Court has just shown what happens when those hallucinations invade the highest-stakes corner of the law: a murder appeal. On Friday, the state’s highest court made public the full disciplining of Stephen P. Aarons, a veteran Santa Fe criminal defense attorney with more than four decades in practice, after discovering that a brief he filed was laced with police testimony and witnesses that ChatGPT invented from scratch.
The court held Aarons in contempt, fined him $5,000, removed him from the appeal, and ordered his conduct referred onward — a package of sanctions that legal observers are calling one of the harshest state supreme court responses to AI-hallucinated filings anywhere in the country.
What happened
The case, State v. Skeets, involved an appeal of a murder conviction. According to the court’s account and reporting by Reuters, the Guardian, and the ABA Journal, Aarons fed portions of the real trial transcript into ChatGPT and asked it to draft a summary of the evidence — he told the court he was trying to build what he called a “bulletproof summary” of his client’s arguments.
What came back looked polished and persuasive. It was also partly fiction. The model:
- Invented witness statements attributed to a real witness who never said them
- Fabricated witnesses entirely — people who do not exist, quoted in a brief filed in a murder appeal
- Manufactured police testimony that appears nowhere in the actual record
None of it was flagged before filing. The brief went to the court carrying Aarons’ name and reputation — and carrying quotations no human had ever spoken, in a case where a person’s freedom for life was on the line.
How it unraveled
The fabrication was discovered during review of the appeal. Once the court compared the brief’s factual assertions against the actual trial transcript, the invented material collapsed immediately: the fake witnesses could not be located in any court record because they had never existed. Aarons admitted his “stupidity” in relying on the model’s output without verification, but the court found his response fell short of genuine remorse — a factor that weighed in favor of contempt rather than a quiet admonishment.
The sanction had multiple components:
- A $5,000 fine, paid for submitting false information to the court
- A contempt finding, the court’s formal power to punish conduct that obstructs the administration of justice
- Removal from the case, stripping Aarons of the representation
- Referral, with his conduct now moving into the disciplinary system that can suspend or disbar attorneys
For a lawyer who has practiced criminal defense in Santa Fe since the early 1980s, that is a career-defining blow delivered in the final chapter of a long career.
Why this case matters more than the citation cases
Since the infamous Mata v. Avianca filing in 2023 — where lawyers cited six nonexistent cases — courts have sanctioned dozens of attorneys for AI-hallucinated citations. The Damien Charlotin database of AI hallucination cases in law now tracks incidents across multiple countries, and the pattern is well established: a lawyer asks a chatbot for help, the model confidently invents authority, and the lawyer signs the filing.
But State v. Skeets breaks new ground in three ways.
First, the fabrications went to facts, not just citations. A fake case citation misleads the court about the law. Fabricated testimony misleads the court about what happened — in a murder case. The epistemic stakes are categorically different. An appellate court weighing a murder conviction depends entirely on the integrity of the record before it. Poison the record and you attack the foundation of the proceeding itself.
Second, the source material was real. Aarons didn’t ask ChatGPT about something it couldn’t know; he fed it the actual transcript. The model still drifted into invention — interpolating plausible-sounding witness statements and stitching nonexistent people into its summary. That is precisely the failure mode that makes LLM summarization dangerous in evidentiary contexts: the output is fluent, structured, and confidently wrong in ways that are invisible without line-by-line verification against the source.
Third, it is a state supreme court acting, not a trial judge. Earlier sanctions largely came from district courts and individual judges catching errors in individual filings. Here the state’s highest court issued a formal opinion — creating precedent that every lower court in New Mexico, and courts watching nationwide, will treat as the ceiling of tolerable AI misuse, not the floor.
The verification gap
The court’s core finding is simple: the sin was not using AI, it was filing unverified AI output. Aarons’ own framing — that he wanted a “bulletproof summary” — reveals the inversion of trust that legal ethicists have warned about since chatbots entered practice. He assigned the model the job of safeguarding accuracy, and took that job himself only after the damage was done.
That gap is widening, not closing. Frontier models in 2026 are more capable than the systems that generated the first hallucinated citations, but capability does not eliminate hallucination — it makes the fabrications harder to spot, because the surrounding text is better. A fabricated quotation embedded in an otherwise accurate summary of a real transcript is far more dangerous than a fake citation in a sloppy brief, because nothing about its style signals invention.
The professional-responsibility infrastructure is scrambling. Over half of U.S. states have now issued guidance requiring lawyers to verify AI-generated citations, and several — including California, Florida, and Michigan — demand disclosure of AI use in some circumstances. But guidance aimed at citations does not automatically reach facts. A rule that says “check the cases exist” says nothing about “check whether the witness exists.” The New Mexico opinion pushes the verification duty to where it actually matters: everything.
A defense lawyer’s failure, a defendant’s cost
There is a cruel irony embedded in the sanction. The person most injured by the fabricated brief is not Aarons — it is his client, a murder appellant who has now lost his longtime lawyer mid-appeal and inherits a record contaminated by his own side’s filing. The prosecution no longer needs to argue the appeal’s merits aggressively; the defense has handed it a credibility catastrophe.
Appellate defenders have noted the systemic risk: indigent defense and court-appointed lawyers, often drowning in caseloads, are precisely the population most tempted to use AI to close the gap — and least equipped with time to verify output line by line. If the result of adopting AI is competent-seeming briefs with fabricated evidence, the technology degrades the very resource-poor corner of the system that most needs help.
The road ahead
The New Mexico Supreme Court’s opinion lands amid a broader regulatory season for AI in the professions. Courts and bar associations across the U.S. are moving from gentle guidance toward hard duties: verification, disclosure, and competence in AI use as an ethical requirement in itself. Judge-written opinions like this one do the quiet work of converting “you should check” into “you will be held in contempt if you don’t.”
For the makers of the models, cases like this sharpen an uncomfortable question. OpenAI’s usage policies require accuracy in high-stakes domains, but policy binds the user, not the model’s behavior. A system that can be fed a murder trial transcript and return invented witnesses is not a defective summarization tool — it is an ordinary LLM doing what LLMs do, with the stakes supplied by the user. The legal system’s answer, so far, is to make the human signature on the filing mean something again.
Aarons’ disciplinary referral will run its course in the months ahead. But the opinion itself is already doing its work: a searchable, citable statement from a state supreme court that the age of filing-first-verifying-later is over. The next lawyer tempted to ask a chatbot for a bulletproof summary now knows the bullet can land in their own client’s case.
Sources
- [1] https://www.reuters.com/legal/government/chatgpt-invented-fake-police-testimony-murder-appeal-new-mexico-high-court-says-2026-09-11/
- [2] https://www.theguardian.com/technology/2026/sep/11/new-mexico-lawyer-ai-chatgpt-testimony
- [3] https://www.theverge.com/ai-artificial-intelligence/994207/chatgpt-new-mexico-lawyer-fined-murder-appeal
- [4] https://www.abajournal.com/news/article/criminal-defense-attorney-admits-stupidity-over-ai-errors-but-still-receives-sharp-rebuke-from-his-states-high-court
- [5] https://caselaw.findlaw.com/court/nm-supreme-court/118143291.html