A $5,000 sanction for an AI-tainted appeal brief
The U.S. justice system is taking a new step in its judicial response to negligent uses of generative artificial intelligence. The New Mexico Supreme Court imposed a $5,000 fine on a lawyer whose appeal brief, filed in a murder case, contained fictitious witnesses and fabricated testimony attributed to a police officer. Those elements were generated with the help of artificial intelligence, according to information reported by The Verge.
The amount of the sanction is significant, but the importance of the decision extends beyond the sum demanded from the lawyer alone. The case concerned an appeal proceeding in a criminal matter, an area in which the quality of the information submitted to the court is particularly decisive. On appeal, lawyers do not argue solely on the basis of intuition or a narrative strategy: they present facts, a procedural record, quotations, references and arguments that are supposed to be verifiable. The presence of people who do not exist and fabricated police testimony therefore represents a failure that affects the very functioning of adversarial proceedings.
The decision serves as a reminder of a basic rule of professional law: a lawyer remains responsible for what they sign, file and present to a court. The use of an AI tool, even when it is used to prepare a document, summarize materials or suggest wording, creates neither immunity nor a delegation of responsibility. The tool can help produce text; it cannot assume the ethical, procedural and factual obligations attached to the legal profession.
In its article on the case, The Verge describes a brief that incorporated content invented by AI into proceedings in which the requirement for rigor cannot be downplayed. This point is central: the problem is not limited to an erroneous legal citation or awkward wording. The disputed elements concerned witnesses and police testimony, that is, components potentially decisive in the assessment of a criminal case.
The New Mexico Supreme Court is not ruling here on the existence of generative AI as such. It is sanctioning the conduct of a professional who allowed insufficiently verified information into judicial proceedings. This distinction is likely to structure a large part of the forthcoming litigation surrounding generative tools: courts do not necessarily ask professionals to give up these systems, but they require them to retain real, documented human oversight suited to the seriousness of the context.
The point is all the more important because text-generation tools are now easily accessible. They can quickly produce a coherent narrative, a structured memo, a summary or a list of references. But their ability to formulate a convincing response is no guarantee of accuracy. When a system does not have reliable information, or when it extrapolates from linguistic patterns, it can generate content that is plausible but false. In common AI terminology, this phenomenon is often referred to as a “hallucination.” In a courtroom, technological vocabulary matters less than the consequence: a document containing an unverified assertion can mislead the court.
In a murder case, a generative error takes on a different nature
The criminal context sharply distinguishes this case from many previous AI-related incidents. An error in marketing material, an incorrect response from a conversational assistant or an inaccurate summary in an internal document can result in costs, lost time or reputational risk. In an appeal proceeding involving murder, the potential consequences are of a different nature. Courts review decisions that may affect a person’s liberty, the validity of a conviction and, more broadly, the trust placed in judicial institutions.
An appeal brief is not merely a preparatory note. It is a document submitted to the court in support of a legal and factual position. It forms part of a procedural exchange in which each party must be able to respond to the other’s arguments, judges must be able to verify the references invoked, and the record must remain faithful to the material actually submitted. Introducing fictitious witnesses disrupts this framework: the other party must then identify an invention, the court must spend time verifying what should have been checked before filing, and the judicial debate may be shifted by facts that never existed.
The fabricated police testimony identified by the New Mexico Supreme Court is particularly revealing of the risk. The testimony of a law enforcement officer may occupy an important place in a criminal case, particularly when it describes an investigation, an arrest, findings or exchanges. Attributing statements to such an officer that do not appear in the record amounts to artificially creating an argumentative piece of evidence. Even if the invention is subsequently detected, it forces the justice system to deal with extraneous information and may undermine the perception of the defense presented.
The sanction also shows that generative AI cannot be treated as merely an enhanced word-processing tool. A spellchecker can flag an error, formatting software can organize paragraphs, and a search engine can point to identifiable sources. A generative model responds differently: it produces language and can give the impression of reporting a fact without reliably providing the chain of evidence that would make it possible to confirm it. This difference requires an appropriate control method.
For a lawyer, this method entails at a minimum returning to source materials: hearing transcripts, police reports, court decisions, documents exchanged by the parties, official documents and exact references. Wording suggested by AI should never replace that reading. At most, it can serve as a starting point for human work; it cannot be the end point when the text is intended for a court.
The case also highlights a tension specific to contemporary legal work. Law firms, legal departments and courts are subject to large volumes of documents and often tight deadlines. Generative AI promises to summarize, classify, rephrase and speed up drafting. But speed has value only if it does not lower the level of reliability. In a sector where the traceability of sources is essential, an initial time saving can turn into a major cost if the user subsequently has to correct an error, respond to a court’s request for an explanation or face a sanction.
The New Mexico decision therefore sends a concrete signal: the seriousness of the matter being handled will weigh in the assessment of negligence. The more a document can affect a person’s rights or the outcome of proceedings, the less acceptable it is to rely on unverified automated output. In criminal matters, this requirement is particularly high, but the principle can be applied to other sensitive sectors: healthcare, insurance, banking, human resources, public administration or financial reporting.
From the Avianca case to case law on professional responsibility
The New Mexico case is part of a series of cases that have made visible a problem long confined to technical demonstrations: a language model can confidently produce false references. The best-known episode remains the Mata v. Avianca case in the United States. In 2023, lawyers submitted a document to a federal court in New York citing nonexistent court decisions. Federal Judge P. Kevin Castel sanctioned the lawyers involved and their firm in the amount of $5,000.
In that now emblematic case, the main issue was already not AI as an entity bearing responsibility. The question was the oversight exercised by the people who had relied on a conversational tool to research or draft legal arguments. The false decisions cited had not been identified before filing. The judge stressed the seriousness of presenting nonexistent case law to a court, while noting that the use of new technologies was not, in itself, prohibited.
The New Mexico sanction extends this logic, but its subject appears even more troubling in light of the facts reported by The Verge. In Mata v. Avianca, the fabricated content concerned case-law references. In the criminal appeal case examined in New Mexico, it concerned people and police testimony. In both cases, the failure is the same: generated text was treated as though it were a reliable research result. But criminal matters reinforce the institutional dimension of the error.
These decisions are gradually drawing a useful boundary for regulated professions. On one hand, AI may be used as assistance, subject to rules on confidentiality, security, professional secrecy and verification. On the other, it cannot become an autonomous source whose outputs are incorporated as-is into a professional document. The professional must be able to explain the origin of every important fact and find its evidence in the relevant file.
This requirement is not limited to lawyers. Journalists must verify information produced by a generative tool before publication. Doctors must check the elements used in a clinical decision. Accountants, consultants, compliance officers and public officials cannot hide behind software if a recommendation, analysis or statement proves erroneous. AI changes work tools, not the chain of responsibility.
For the justice system, the issue is also organizational. When false elements appear in proceedings, courts must detect them, characterize them and sometimes put corrective measures in place. This takes time for judges, court registries and opposing parties. The promise of efficiency associated with generative models can then turn against the judicial system: a document produced more quickly but insufficiently checked ultimately lengthens the handling of the dispute.
The multiplication of sanctions, warnings and procedural decisions could thus create a kind of practical standard even before highly detailed rules are adopted everywhere. Courts do not need to wait for AI-specific legislation to reiterate existing obligations of fairness, diligence and truthfulness. Professional law, procedural rules and courts’ sanctioning powers already provide instruments for responding to a deficient brief.
The scope of this development nevertheless needs to be understood precisely. It is not a matter of asserting that every error in a document prepared with AI automatically constitutes a disciplinary offense. Errors exist in human work, and proceedings provide correction mechanisms. What the publicly reported cases highlight is the particular risk of a use in which the user does not verify results before presenting them as accurate. Negligence does not lie in having sought help from a system; it lies in abandoning the oversight that must accompany that help.
For French and European legal professions, an operational warning
The U.S. decision does not apply directly to French lawyers. Nonetheless, it constitutes a very concrete warning for all legal professionals in France and Europe. Generative assistants are now capable of drafting email templates, summarizing documents, suggesting outlines for submissions or extracting recurring themes from a corpus. Their spread in everyday practices therefore raises less the abstract question of their arrival than that of their governance.
In France, lawyers are already subject to professional obligations that make it difficult to defend filing a document containing unverified facts or sources. The confidentiality of exchanges, professional secrecy, fairness in proceedings and the quality of legal advice remain structuring requirements. AI available online can also pose an additional problem if confidential information, personal data or non-public documents are copied into its interface without sufficient guarantees as to their processing.
The New Mexico case thus illustrates two distinct risks that may overlap. The first is the reliability risk: the tool invents a reference, a fact, a person or a quotation. The second is the confidentiality risk: the user transmits information to an external service that they should not have disclosed. Even when the final text is factually correct, the production process can therefore raise compliance and data-protection issues.
The European framework is evolving in parallel. The European regulation on artificial intelligence, often called the AI Act, introduces phased obligations for actors concerned by AI systems. Its approach is based in particular on the risk level of uses. It does not turn every professional user into a technical specialist, but it increases attention to governance, documentation and control of risks associated with automated systems.
For French law firms, the lesson from this case is above all operational. Responsible use of generative AI requires defining the tasks for which it may be used and those requiring enhanced review. Rephrasing a non-sensitive passage, preparing a work plan or creating a list of questions may fall under supervised assistance. By contrast, case-law citations, doctrinal references, testimony quotations, procedural elements and decisive facts must be checked against their primary sources before any use.
This distinction should be formalized rather than left to the isolated judgment of each associate. An internal policy may provide for a ban on submitting certain categories of data to public tools, an obligation to disclose the use of AI in the drafting process where relevant, and the retention of sources used to verify important assertions. The aim is not to bureaucratize every task, but to prevent plausible text from being mistaken for reliable text simply because it is well written.
French legal technology publishers also have an interest in drawing the consequences of these cases. The value of a legal assistant does not depend solely on the fluency of its responses. It depends on its ability to link generated information to identifiable databases, cite sources that can be consulted and enable the user to clearly distinguish a generated summary from an original document. In law, an answer without a source is less useful than it may appear: it forces the professional to redo the verification work in full.
The French-speaking market has a practical specificity in this respect. French law, European Union law, national case law and sector-specific regulations form a documentary environment in which terminological precision is decisive. A poorly reproduced reference, a decision confused with another, or a provision cited outside its context can alter the scope of an argument. The linguistic quality of a French-language model therefore never removes the need for legal verification, just as the availability of a French-language interface does not guarantee that the tool knows the latest version of an applicable text.
In this context, training appears to be an issue at least as important as the choice of tool. Young professionals may be tempted to view conversational assistants as search engines. Yet the two tools do not operate in the same way. A search engine points to pages or documents that the user can consult; a generative model directly produces an answer, sometimes without correctly indicating the limits of its knowledge. Confusing the two uses is one of the quickest paths to professional error.
Human responsibility becomes the real market for legal AI
In the short term, cases such as the New Mexico case should reinforce caution among law firms and courts. Some organizations will choose to sharply limit the use of generative assistants, especially in criminal, litigation or highly confidential matters. Others will invest in private environments, tools connected to controlled documentary sources and systematic review processes. Both responses reflect the same observation: text production is not the core of the problem; the reliability of what is asserted is.
This development could favor solutions that provide access to source documents rather than those that merely provide a drafted answer. In the legal sector, a useful system should ideally make it possible to trace information back to the decision, the official text, the cited document or the relevant transcript. Generating a paragraph will remain attractive, but its professional value will increasingly depend on the ability to audit it.
The concept of auditability is essential. When faced with a response produced by a model, the professional must be able to ask: which documents does this sentence rely on? Is it a faithful rephrasing, an inference or an invention? Which sources confirm this fact? Which colleague validated the passage before filing? In sensitive proceedings, these questions should precede the signature, rather than arise after an error is discovered.
The New Mexico Supreme Court’s decision shows that courts can already turn these requirements into financial and professional consequences. The $5,000 fine is not merely an individual penalty. It becomes a benchmark for a market still seeking its legitimate uses. It reminds publishers, employers and users that the argument of innovation is insufficient in the face of an obligation of competence.
This reality could change how companies sell legal AI. Promises of instant drafting or massive automation may seem insufficient if they are not accompanied by guarantees of traceability, access control, data management and human validation. Professional users are not merely looking for fast text. They are looking for text they can defend before a client, regulator, judge or professional body.
For courts, the next step could be the adaptation of procedural practices. Without presuming specific rules, courts may be led to seek clarification when a reference appears impossible to find, reiterate verification obligations in their decisions or sanction the most serious failures. The objective would not be to police every use of AI, but to prevent automation from serving as a pretext for degrading the quality of written submissions.
The long-term outlook is therefore less one of a simple opposition between justice and artificial intelligence than of stricter conditions of use. Generative tools will probably continue to become integrated into research, drafting and document analysis. But their adoption in high-responsibility professions will depend on their ability to fit into a chain of evidence and human oversight. The New Mexico case, as reported by The Verge, sets a clear limit: when a professional presents nonexistent witnesses and fabricated testimony to a court, AI is not an excuse. On the contrary, it becomes the indicator of a methodological failure for which the professional remains fully accountable.
Comments· 2 comments
The article gets the headline-grabbing AI failure across, but it feels a little thin on the practical safeguards that could prevent this kind of mistake. I would have liked more discussion of where professional responsibility ends and technological overconfidence begins.
I see that point, although the central lesson may be simple enough: a lawyer should verify every source and factual claim before filing. More detail on safeguards would be useful, but it should not obscure that basic duty.