What the experiment actually found

The study, published by Cambridge University Press in the journal Judgment and Decision Making and led by Sydney Sears and Dr. Deena Skolnick Weisberg from Villanova University's Department of Psychological and Brain Sciences, set out to answer a simple question: can today's readers still distinguish fiction written by humans from fiction generated by large language models?

To find out, the researchers selected six short stories—three written by published authors and three generated by ChatGPT—and presented them to different groups of readers across three separate experiments.

In the first experiment, 1,682 participants read a single story accompanied by an author label that was either accurate or deliberately misleading, before rating how engaging they found the text. In the following experiments, involving 424 and 481 participants respectively, each reader received one human-written story and one AI-generated story without being told which was which and was asked to identify the AI-generated one.

The results were remarkably close to random guessing. Identification accuracy reached 39.93% in the second experiment—significantly worse than chance—and 51.97% in the third, statistically indistinguishable from random choice.

The most revealing finding emerged once the author label entered the equation.

"Participants gave the highest ratings to stories they were told had been written by humans, even though they had actually been generated by AI," Weisberg explained.

According to the researchers, this pattern reveals a bias in favor of stories believed to be written by real people. The preference is not for authenticity itself, but for the label. When an AI-generated story already benefits from the clarity and fluency typical of modern language models while also carrying a false "human-written" label, human authors end up competing not only with AI, but with the reputation that human creativity has built over centuries—a reputation temporarily borrowed by a machine.

Why AI-generated writing often appeals to readers

The researchers argue that the explanation lies more in the psychology of reading than in any artistic superiority of AI.

According to Weisberg, AI-generated texts tend to be clearer, more direct, and easier to read, whereas stories written by humans are often more subtle and more complex. In other words, the average reader may naturally prefer the predictability and accessibility of an easy-to-read text over one that demands greater cognitive effort.

Interestingly, the people who proved most capable of identifying AI-generated stories were not literature experts, but participants with a high degree of familiarity with artificial intelligence systems. They had learned to recognize stylistic patterns characteristic of language models, suggesting that detecting AI-generated writing depends more on AI literacy than on refined literary taste.

The study has important limitations

The authors acknowledge several limitations in the experimental design.

The stories used in the study were very short—requiring only a few minutes to read—and the results could differ for longer works, where AI systems may struggle more with sustained character development and narrative structure. In addition, all of the stories belonged to the realistic fiction genre, and the researchers note that different genres could produce different outcomes.

Finally, participants in Experiments 2 and 3 were explicitly told that one of the two stories had been generated by AI, information that readers would not normally have in real-world situations.

An increasingly difficult creative landscape to map

The implications extend well beyond fiction.

The same difficulty in reliably distinguishing human-written from AI-generated text has also emerged in higher education, where automated detection tools such as Turnitin have proven almost as unreliable as the readers in this study. According to Inside Higher Ed, universities including Yale, Vanderbilt, Johns Hopkins, and Northwestern have gradually moved away from AI detection tools after experiencing false-positive rates significantly higher than initially promised by their developers.

The study adds to a growing body of research challenging the long-held assumption that creative writing depends exclusively on uniquely human qualities such as emotional understanding or lived experience. Weisberg argues that public perceptions of AI's creative abilities are becoming "increasingly outdated."

The research does not show that AI-written stories are inherently better. Instead, it demonstrates that the "written by a human" label influences readers' judgments more than the content itself. Our perception of a text's value is strongly shaped by who we believe wrote it. As AI-generated text detection systems also continue to struggle with reliability, distinguishing between human and artificial authors is likely to become increasingly difficult in the years ahead.

Sources