Bot or not?
Readers rate AI stories higher than human literature – and cannot tell them apart – what publishing professionals should take away from a new study
Published: 15.9.2026 | Foto / Video: AI generated, Magnific
A new study in the journal "Judgment and Decision Making" presented short stories to 2,587 people, some written by authors, some generated by ChatGPT. The findings are uncomfortable for the book trade on two counts. The AI texts scored higher on perceived quality. And nobody could reliably say which text came from whom.
Sydney Sears and Deena Skolnick Weisberg of Villanova University ran three experiments on how readers respond to AI-generated fiction. The paper is titled "Bot or not", was accepted for publication in June 2026 and is freely available. It follows a series of similar studies that had recently shown, mainly for poetry, that machine-generated texts are barely distinguishable from human ones.
The setup
Three published short stories served as source material. "FISH" by Emilie Fox, which appeared in Longleaf Review in 2021, "High Heels" by Susie Maguire from the 2000 collection "The Short Hello", and "Inisfree" by Patrick Smyth from a 1993 collection. For each of these the researchers had ChatGPT 4.0 produce a counterpart of roughly 1,000 words, each time with an instruction that reproduced the theme, narrative perspective and imagery of the human original.
The paper consists of three self-contained experiments, each with its own group of participants. None of them is a pilot, each answers a different question. In the first experiment, 1,682 US adults read only a single text. Half were told truthfully who had written it, the other half were told the opposite. Ratings were given on two scales running from minus three to plus three. One captured how far readers were drawn into the story, the other the perceived quality of plot, characters, theme and style.
The AI wins on perceived quality
On the quality scale the ChatGPT texts averaged 1.54, the human-written ones 0.97. On being drawn into the story the AI led with 1.42 against 1.00. Both gaps are large enough not to be down to chance.
The label worked in parallel. Anyone who believed they were reading a human-written text gave better marks throughout, regardless of what was actually in front of them. On quality that meant 1.40 against 1.12, on absorption 1.32 against 1.10. So the two effects pull against each other. Participants in fact preferred the machine while penalising anything that looked like a machine.
This matches a pattern the research community now knows well. Disclosing that AI was involved is paid for in lower ratings. A large-scale study by Raj, Berg and Seamans concludes that this discount cannot be argued away. Even when participants are told in advance how capable generative systems are, it persists.
Worse than chance
The second and third experiments, with 905 new participants between them, asked the question directly. Each person received a pair of texts, one human and one generated, and had to assign them.
In experiment 2, 39.4 percent got it right. That is worse than if participants had tossed a coin. Experiment 3, run several months later, came out at 51.97 percent, exactly the level of pure guessing. The wording of the question made no difference. Whether people were asked to pick the human or the machine text, the hit rate was the same. Confidence was no help either. Those who felt sure were not right any more often.
The intermediate step is striking. In experiment 2, participants systematically took the human stories for AI products and the other way round. They were not merely off, they were reliably off in the same direction.
The wrong cues
The study also asked what participants based their decision on. Those who cited the language were markedly more often wrong. In experiment 3 the same held for those who went by their own enjoyment of the text. Only those who attended to symbolism and imagery were right somewhat more often, and even that edge was weak.
A further finding is of interest to the trade. Those who credited themselves with a lot of experience handling AI systems made the right call more often. That held in experiment 2 with a purpose-built questionnaire and equally in experiment 3, which used a questionnaire already tested in other research. Experience with literary fiction, by contrast, was no help at all. People who read, write or work with texts professionally did not spot AI prose any better than people with no literary background whatsoever.
So if you want to recognise what comes out of the machine, you need experience with the machine, not with literature. That is a concrete statement about which kind of experience will count in editorial and copy-editing work.
What publishers can take from this
Three points matter in practice.
First, trained reading experience alone does not work as a detection instrument. Anyone in copy-editing or manuscript screening who assumes that machine texts give themselves away by their sound is working against the evidence. The features that intuitively sound like AI are apparently precisely the wrong ones.
Second, the labelling question is not a purely legal one. The markdown applied to disclosed AI involvement is real and measurable. Publishers meeting transparency obligations or labelling voluntarily should factor in that the label shapes perception, independently of actual text quality. This applies to cover copy, metadata and press work alike.
Third, the apparent quality advantage should be treated with care. Wenger and Kenett have shown that generative systems convince on the single piece but produce very similar output across many texts. Human authorship varies far more widely. For a publishing programme that defines itself through distinctiveness, that is the real point. A single AI story survives the comparison. Twenty of them in one seasonal catalogue probably do not.
Limits of the research
The authors name the constraints themselves, and they weigh heavily enough that the findings should not be overstretched.
At around 1,000 words the texts are very short. Whether an AI can carry characters and plot arcs equally convincingly across the length of a novel is not answered here. All the stories are realistic fiction; genres such as romance or science fiction were not tested. The matching task also told participants that exactly one text per pair was machine-generated. No such information is available in everyday life, so real conditions are likely to be less favourable still.
There is also a point of method that the paper does not discuss. The instructions given to ChatGPT were reverse-engineered from the human originals, including motif, perspective and tone. What is being compared, then, is not free machine creativity against free human creativity, but a very tightly steered imitation against its original. The quality questionnaire was moreover newly devised and had never been tested before, though the authors credit it with a degree of face validity. And none of the three experiments was registered in advance. The researchers therefore did not publicly commit to their procedure and their expectations before they began collecting data, something now regarded as a mark of quality in psychology.
All participants were based in the United States. The texts were generated between November 2023 and December 2024, which makes the model used long since superseded.
That last point hardly counts against the findings. Today's models would be expected to do better rather than worse.
---
Source: Sears S, Weisberg DS. Bot or not: Can people tell the difference between stories written by a human or by an AI system? Judgment and Decision Making. 2026;21:e21. doi:10.1017/jdm.2026.10042

