Radar Perene / Archive / science
AI authorship and detection in science: who signs, and the real limits of detectors
◦ Index methodology v2.2 (working papers with DOI). See the methodology.
Science
In the first months of generative AI, a few scientific articles were actually published with an AI tool listed among the authors — name on the cover, next to the humans. The system's reaction was swift and nearly unanimous: the listings were corrected, and the prohibition became the norm. The episode, curious in itself, opened the two questions this article describes — prescribing an answer to neither. What exactly makes someone an author? And, if the rule exists, how would anyone verify that it was broken?
In the consensus of publishers and publication-ethics committees, scientific authorship requires four capacities: contributing substantially to the work, approving the final version, answering publicly for errors and declaring conflicts of interest. The exclusion of AI follows from there — not from the tool's performance, but from its inability to answer for anything.
The authorship debate: deeper than the rule
The canonical formulation — AI does not sign because it does not answer — settles the easy case and leaves the hard ones open, and the honesty of the current debate lies in admitting it. The public positions of publishers and of the international committee on publication ethics (COPE) converge on the rule and diverge at the boundary: how much tool use is compatible with the statement "this work is mine"?
One side of the spectrum argues from historical continuity: researchers have always used instruments that take part in intellectual work — libraries, calculators, statistical software, spell-checkers — and no one ever listed the regression package as an author. Through the lens of responsibility, the new tool is one more layer: whoever signs, verifies and answers. The other side points to a real discontinuity: earlier tools executed operations the researcher specified; generative AI produces formulations — candidate text, argument structures — and the border between assistance and intellectual contribution becomes genuinely blurred. Between the two poles, the policies chose a pragmatic cut, described in the previous article on this trail: full human responsibility, declaration of relevant uses. It is an administrable cut, not a philosophical answer — and the published positions have the modesty to say so.
Detectors: what the public evidence shows
If the rule prohibits, the institutional temptation is to buy automatic enforcement — and the supply exists: tools that promise to estimate whether a text was machine-written. The public literature on these detectors, however, documents three limits any reader of research should know, because decisions about careers have already been made on top of them.
The first limit is false positives — human text classified as synthetic. The public record includes emblematic cases, such as historical texts, written decades before any AI, flagged as generated. Studies have further pointed to a graver pattern: texts by non-native speakers of the language are flagged disproportionately often, because more regular prose with a more contained vocabulary statistically resembles the machines' average output. In a global science written mostly in English by non-natives, that is not a side defect.
The second limit is evasion: paraphrase, light human editing or style instructions degrade detector accuracy. The third is structural and summarizes the first two — detection is a race between models, with no finish line: each generation of writing tools moves the target of the detection tools. One of the largest AI developers launched and then discontinued its own text detector, citing the low accuracy rate; the episode is frequently cited in the debate as the honest summary of the state of the art.
The final asymmetry deserves the record: an accusation based on a detector is nearly impossible to refute. There is no conclusive proof that a text was not generated — and systems that punish on irrefutable evidence collide with elementary principles of due process. It is for that reason, documented in the public discussions of editorial ethics, that the journal policies described on this trail bet on declaration and responsibility, not on statistical policing of text.
What remains when detection fails
The debate's conclusion, in its current state, has an involuntary elegance: without a reliable detector, the system falls back on what it always had — verification of content, not of origin. An article with checkable data, described method and reproducible results can be evaluated whichever hand typed it; an article without those properties is weak even if written in longhand. The classic integrity mechanisms — replication, open data, review — police what matters without needing to guess the tool. The question "was it a machine?" ages; the question "does this hold up when checked?" does not.
Frequently asked questions
Why can't an AI simply be a co-author?
Because scientific authorship embeds legal and moral responsibility: answering for error, fraud and conflicts of interest. A system that cannot be held responsible does not meet the definition — this is the declared foundation of the policies, uniform across publishers.
Do AI text detectors work?
With relevant error rates in both directions, degradation under editing and paraphrase, and documented bias against non-native speakers. The public evidence places their usefulness at indicative triage, very far from proof.
Has anyone been harmed by a false positive?
The documented cases concentrate in education, where detectors were used in disciplinary decisions — and prompted policy revisions at several institutions. The scientific editorial debate cites those episodes as a warning.
How does a journal enforce the rule, then?
In practice: author declarations, contractual responsibility, editorial screening for inconsistencies (nonexistent references are the most common trace) and the ordinary post-publication integrity mechanisms. Statistical verification of the text is not the pillar — for the fragility described above.
---
The integrity trail closes here. The next one is made of more tangible matter: Brazil's open databases — beginning with the table behind the country's most cited price index: IBGE SIDRA: how a table becomes research input.
House reading: today's reading is in the Diário.
This is the Radar’s memory. Today’s reading — regime, 5 lenses and the day’s analogs — is live, free.