ArXiv cracks down on AI-generated papers
ArXiv may ban for one year authors of papers massively generated by AI, a turning point for the use of LLMs in scientific research.
ArXiv tightens its rules in response to the rise of scientific papers produced by AI
The world’s leading repository of scientific preprints, arXiv, is preparing to take a symbolic step in regulating generative artificial intelligence. According to TechCrunch, the platform plans to ban for one year authors who submit papers largely written by AI systems, when those tools have done most of the intellectual or writing work. The signal is strong: after a phase of enthusiasm around large language models, the academic world is beginning to formalize red lines and attach explicit sanctions to them.
Created in 1991 and administered by Cornell University, arXiv occupies a central place in the rapid circulation of work in physics, mathematics, computer science, statistics, quantitative biology, and economics. The platform now hosts more than 2 million articles and receives tens of thousands of new submissions each month. In fields such as AI, where competition is often measured in just a few weeks, publishing on arXiv has become a reflex even before peer review.
It is precisely this strategic position that gives arXiv’s new direction a much broader significance than a simple moderation adjustment. By targeting what many researchers now call AI slop — texts produced quickly, barely reviewed, often padded with generic phrasing and sometimes scientifically weak — the platform is intervening at the heart of a problem that is already affecting the credibility of the scientific literature.
What the new policy provides for and what it seeks to prevent
According to information reported by TechCrunch AI, arXiv wants to sanction authors who let AI “do all the work,” with a stated penalty of one year of exclusion. The goal is not to ban all use of generative models, but to target submissions in which the machine effectively replaces the author instead of assisting them. The distinction is important: correcting wording, rephrasing an abstract, or improving the English of a text is not equivalent to producing a proof, a literature review, or a scientific discussion without substantial human oversight.
This development comes amid a climate of heightened vigilance. Since the explosion of ChatGPT at the end of 2022, followed by the arrival of tools such as Claude, Gemini, or open-source models from the Llama family, academic uses have multiplied. Many labs already use LLMs to translate, summarize papers, generate code, prepare figures, or rephrase passages. But this rapid adoption has also opened the door to abuses: verbose texts, invented references, methodological errors masked by fluent prose, and even quasi-automatic submissions intended to artificially pad a CV.
ArXiv is not a traditional peer-reviewed journal, but a dissemination server. Its role is nevertheless crucial, because it often serves as the first public showcase. If that showcase fills up with weak or semi-automated content, the entire chain of trust deteriorates: readers, journalists, investors, recruiters, and other researchers rely on these texts to follow the state of the art. In this sense, the announced decision is as much a matter of editorial regulation as it is of defending a scientific infrastructure.
The implicit message is clear: AI can assist research, but it must not replace the scientific responsibility of authors.
After the euphoria of LLMs, the return of boundaries in academic production
The shift is as much cultural as it is technical. For the past two years, the dominant narrative around LLMs in research has rested on productivity gains: writing faster, synthesizing more literature, speeding up the drafting of projects, reports, or preprints. In teams most exposed to publication pressure, the promise was appealing. But textual productivity guarantees neither scientific quality nor intellectual traceability.
The problem is particularly acute in disciplines close to computer science and AI, where the pace of publication is high and the barrier to entry for writing can seem lower. A generative model can produce in a few seconds an introduction that appears credible, with the expected academic tone, standardized phrasing, and a familiar structure. This veneer may be enough to circulate mediocre work, especially in prepublication spaces where evaluation comes after dissemination.
ArXiv’s decision therefore reflects a shift: the use of LLMs is no longer merely a matter of personal tooling, but a subject of scientific governance. Who is the real author of a text? What share of the reasoning has been outsourced? How can it be verified that the author understands what they are submitting? At what threshold does assistance become abusive delegation? These questions, long theoretical, are now taking on a disciplinary and operational dimension.
The term “AI slop,” popularized on the web to describe low-quality content generated at scale, is now entering the academic debate. Its appearance in the research sphere is in itself revealing: scientific institutions now consider that they can be contaminated by the same dynamics as mainstream content platforms, namely abundance, speed, and the dilution of responsibility.
A sensitive issue for French and European research
For the French-speaking ecosystem, this development is far from abstract. French, Belgian, Swiss, and more broadly European labs also use generative assistants in their daily workflows. In France, where public teams must contend with strong publication pressure, constrained budgets, and growing internationalization of exchanges, help with writing in English is often seen as a practical lever. LLMs can reduce a real linguistic disadvantage for non-native researchers.
But this usefulness comes into tension with several structuring principles of European research: scientific integrity, individual responsibility, methodological transparency, and reproducibility. French institutions have already begun publishing internal recommendations on the use of generative AI, often inspired by the work of ethics committees, the CNRS, universities, or engineering schools. ArXiv’s tightening could accelerate this formalization by pushing institutions to specify what is allowed, what must be disclosed, and what constitutes fraud.
The issue also affects young researchers. PhD students, postdoctoral researchers, and candidates for academic positions are among those most exposed to the temptation to automate writing. Yet they are also the ones for whom a one-year exclusion from arXiv can have very concrete effects on visibility, the search for collaborations, or career timelines. In some subfields of AI and machine learning, not being able to publish a preprint for twelve months amounts to temporarily disappearing from the international radar.
- For laboratories: the need to document acceptable uses of LLMs.
- For authors: direct disciplinary risk in the event of a submission deemed excessively generated.
- For institutions: the obligation to clarify the boundary between linguistic assistance and intellectual production.
- For readers: increased expectations of transparency about the conditions under which a text was written.
Regulation that is difficult to enforce, but already structuring
One central question remains: how will arXiv determine that a paper was “largely produced” by an AI? Automatic detectors of generated text are notoriously fragile, with frequent false positives and false negatives. Human-written texts can be flagged incorrectly, while generated content that has then been reworked goes unnoticed. Implementation will therefore probably rely on a mix of signals: an unusually generic style, errors typical of LLMs, bibliographic inconsistencies, anomalies in proofs, and human reports.
This difficulty does not cancel out the effect of the rule. In moderation, the norm often matters as much as the detection tool. By announcing a clear sanction, arXiv is creating a precedent and shifting the center of gravity of the debate. The question is no longer “can an LLM be used to write a paper?” but “how far can one go without breaking the contract of trust with the scientific community?”
It should also be noted that the measure comes in a broader context of regaining control. Scientific publishers, conferences, and universities are gradually adjusting their policies. Some journals already require disclosure of AI use. Others prohibit generative models from being listed as co-authors, on the grounds that they can neither assume legal responsibility nor respond to criticism. ArXiv, for its part, adds a more visible punitive dimension, suited to its role as a gateway to scientific dissemination.
For AI companies, including those selling research assistants, the message is ambiguous. On the one hand, their tools remain useful for peripheral tasks. On the other, the idea of largely automated scientific writing is becoming institutionally suspect. This could encourage the emergence of more specialized solutions focused on verification, sourced citation, auditing of changes, and traceability of human contributions.
Toward augmented science, but under a burden of proof
ArXiv’s tightening probably marks the beginning of a more mature phase in the relationship between research and generative AI. The challenge will not be to eliminate LLMs from academic practices, an unrealistic scenario, but to reintegrate them into a chain of demonstrable responsibility. In other words, assistance will remain tolerated, even encouraged, provided that the author can prove that they control the content, verify the references, understand the results, and stand behind the argumentation.
This development could transform publication norms well beyond arXiv. One can imagine, in the medium term, standardized disclosure forms, editing logs integrated into writing tools, or institutional policies requiring the precise documentation of model use. In Europe, where the regulatory culture is more pronounced than in the United States, this logic could find fertile ground, particularly in public institutions and projects funded with European money.
The paradox is that generative AI, designed to streamline text production, could ultimately lead science to demand more proof about the origin of phrasing, reasoning, and editorial choices. The more machines can write like researchers, the more researchers will have to show how their work remains irreducibly human. That is probably where the real turning point lies: not the end of LLMs in research, but the end of the illusion that writing faster is still enough to produce credible science.
Comments· No comments yet
Be the first to react.