As allegations of AI use persist in roiling the literary and media worlds, a core question has emerged: can humans truly tell between writing penned by people and AI-produced text? According to Prof. Claire Hardaker, a forensic linguistics professor at the University of Lancaster, the answer is far less certain than most people believe. Her web-based test, Bot or Not, shows that the typical person accurately spots machine-written text only about three-fifths of the time—a troubling finding for those certain they can spot machine-written prose at a glance. The doubt has generated a climate of suspicion, with writers receiving allegations of AI use based on merely intuitive judgements, even as linguists warn that the supposed “tells” of AI text are frequently seen in acclaimed writers.
The Detection Challenge
When asked to determine which of three hotel reviews were synthetically created, most people fare little better than a coin toss. The opening assessment, abundant in passion and repetitive praise, reads like countless online testimonials—yet it was composed by an actual person. The second one, with its wry observation about spatial dimensions and casual mention of luggage concerns, carries the idiosyncratic voice of authentic firsthand knowledge. The third one, well-crafted and proficient, could reasonably have come from either a content visitor or a language model. This exercise reveals a disturbing fact: our instincts regarding genuineness are unreliable guides in a time of complex machine learning systems.
The problem stems partly from what we’ve grown to anticipate from both humans and machines. Large language models were instructed with vast quantities of human-authored material, absorbing not just grammar and vocabulary but also the distinctive flourishes and rhetorical devices that define authentic prose. When an AI system creates written content, it pulls from this learned collection, making it almost indistinguishable from the source material it studied. As a result, the elements we consider suggest AI authorship—clichés, dashes, rhythmic patterns—are just as prevalent in acclaimed human writing and everyday communication.
- The “principle of three” is found in both human and AI writing
- Clichés are widespread across human reviews and casual writing
- Em dashes were favoured by Charles Dickens and other renowned authors
- Writing patterns embedded in machine learning models derive from written human language
Why Our Natural Instincts Fall Short
Claire Hardaker notes that people rely on oversimplified mental shortcuts when attempting to identify machine-generated text. They’ve picked up fundamental rules—catch a cliché, flag an profusion of dashes, recognise the appealing combination of phrases—and now employ these guidelines across the board. The issue is that these purported markers are not exclusive to machines. Orators have employed the threefold pattern since antiquity; Julius Caesar’s well-known “Veni, vidi, vici” illustrates that this approach predates computers by thousands of years. Likewise, dashes are found across the corpus of written works, making their presence an unreliable sign of machine authorship.
The fundamental problem is that we’re searching for distinguishing features in a landscape where human and artificial writing have merged together. Because AI systems learned from human texts, they’ve picked up not just grammatical accuracy but also the peculiarities, stylistic touches, and features that give human writing its unique quality. As Hardaker notes with wry humour, one could analyse the works of Dickens and argue he must have been artificially produced, just because he used em dashes. Our ability to detect fall short because the line between human and machine-generated writing has grown genuinely unclear.
Distinctive Patterns of Automated Content
Despite the difficulty in distinguishing human from artificial text, researchers have begun identifying subtle patterns that may reveal machine authorship. Large language models tend to favour certain constructions and vocabulary choices that, whilst grammatically sound and contextually appropriate, occur with unusual frequency compared to authentic human text. These patterns emerge not from deliberate design but from the mathematical patterns embedded in the training data. When an AI system generates text, it optimises for probability at each step, sometimes leaning toward safer, more conventional choices that reflect the most common patterns in its training corpus rather than the full spectrum of human expression.
One developing area of research involves what linguists call “statistical anomalies”—sequences of words or stylistic choices that, whilst individually unremarkable, cluster in ways that deviate from authentic human prose. Researchers at organisations such as MIT and Stanford have begun developing identification tools that examine not superficial characteristics like punctuation, but deeper structural patterns in how ideas are constructed and linked. These methods show potential, though they remain imperfect. The challenge lies in the moving target of AI development itself; as models become increasingly advanced and are trained on increasingly diverse human writing, the distinctive markers that once seemed characteristic continue to evolve and blur.
| AI Tendency | Example |
|---|---|
| Excessive hedging language | “It could be argued that one might suggest…” |
| Repetitive transitional phrases | Frequent use of “furthermore,” “moreover,” “in addition” |
| Balanced, symmetrical sentence structures | Parallel constructions appearing more regularly than in natural speech |
| Avoidance of controversial specificity | Generic descriptions rather than particular, idiosyncratic details |
| Uniform tone maintenance | Absence of the emotional fluctuations typical of human writing |
The Dig Trend
Linguists have observed what some researchers call the “Delve phenomenon”—a pattern in AI-generated text to use the word “delve” and similar formal verbs far more often than occurs naturally in human writing. This quirk exemplifies how AI systems can betray their artificial origins through concentrated grouping of particular vocabulary items. The phenomenon isn’t unique to “delve”; similar patterns appear with words like “whilst,” “endeavour,” and “peruse,” which feature prominently in training data but are used more sparingly by actual human writers in everyday contexts.
The Delve phenomenon underscores a crucial insight: AI systems fail to genuinely comprehend language in the way humans do. They create content by anticipating mathematically probable sequences based on patterns within training data, occasionally enhancing less frequent though technically correct word selections. This generates a quiet but noticeable signature—a kind of textual fingerprint that stems from the computational basis of how these systems function. Identifying such patterns may in time show more dependable than depending on intuitive assessments about writing style.
How AI is Changing the Way We Communicate
The rapid spread of AI-generated content is significantly reshaping the communication terrain in ways simultaneously delicate and substantial. As language models become increasingly sophisticated and ubiquitous, they’re not simply replicating human writing—they’re commencing to affect how humans themselves communicate. Content creators, media professionals, and ordinary people are inadvertently adopting patterns from algorithmic content they discover on the internet, creating a feedback loop that incrementally transforms linguistic norms. This symbiotic relationship between human and machine language prompts critical concerns about originality and the protection of distinctive human voice in an time of machine-driven impact.
The incorporation of AI into creative and professional writing has accelerated discussions about what makes language authentically human. Beyond surface-level stylistic markers, linguists are investigating deeper questions about intention, emotion, and the ineffable quality that distinguishes genuine human expression. The challenge isn’t merely identifying AI writing in isolation; it’s understanding how constant exposure to machine-generated text alters human discourse patterns. As AI systems become more adept at mimicking human language, the distinction between the two gradually erodes, forcing organisations to re-examine what we appreciate in written expression and how we define originality.
- AI text often exhibits statistical patterns that depart from authentic human language variation.
- AI models enhance formal vocabulary choices contained within training data to an excessive degree.
- Human writers more and more adopt AI-driven patterns without realising by means of exposure.
- Distinguishing authentic voices necessitates comprehension of algorithmic tendencies instead of depending on intuition alone.
Cultural Standardisation Via Algorithmic Systems
As AI language models learn from vast collections of published text, they inevitably reduce linguistic diversity and geographical differences. The algorithms prioritise statistically frequent structures whilst sidelining distinctive dialects, colloquialisms, and unconventional expressions that embody authentic human communication. This homogenisation effect threatens to diminish distinctive cultural features in favour of a standardised, algorithmically-optimised form of English. Writers and speakers from varied communities risk witnessing the dilution of their distinctive language traditions as AI systems propagate increasingly uniform language patterns, potentially undermining the depth and variety of human communication.
The More Profound Concern About Genuine Identity
The difficulty in differentiating between human from machine-generated text exposes a more fundamental anxiety about genuineness in the digital age. When well-known authors face accusations of using AI—whether justified or not—it speaks to a more general cultural unease about what we prize in writing. The concern isn’t merely technical; it’s fundamental to our identity. Readers desire to think they’re reading authentic human expression, emotion, and lived reality on the text. The lack of capacity to consistently distinguish AI generated text undermines that confidence, establishing a climate of suspicion where even accomplished writers have their work disputed. This erosion of certainty about writer’s meaning has far-reaching effects for the way we assess literature and determine creative merit.
Yet the irony runs deeper still. The very qualities we associate with human authenticity—emotional resonance, distinctive voice, compelling narrative—are exactly what AI systems are increasingly capable of simulate by analysing vast quantities of human-written text. This presents a paradox: if a machine can produce writing that meets our standards for genuine human communication, what does authenticity really signify? The question compels us to confront inconvenient facts about the extent to which what we identify as distinctly human might simply be detectable patterns, acquirable skills, and repeatable stylistic decisions. Perhaps authenticity doesn’t rest in surface features but in something less tangible: the lived experience and conscious intent behind the words.
Can Machines Ever Produce Great Literature
Literary novelists including Jennifer Egan and Jeanette Winterson have grappled with this question, recognising that whilst AI can produce mechanically proficient prose, something essential seems to elude artificially generated fiction. Great literature, they argue, arises out of authentic human hardship, moral complexity, and the gathered insight of lived experience. AI systems, however advanced, function without consciousness, without stakes, without the existential significance that permeates transformative prose. They can reproduce the structure of a moving story without understanding what it means to be moved. Yet this argument, too, grows shakier as AI capabilities expand, compelling writers to articulate what precisely separates skill from brilliance.
The question ultimately uncovers our uncertainty about literature’s role. If we value novels primarily for entertainment and information, AI may ultimately become adequate. But if literature serves as a profound exploration of the human mind—a bridge between individual experience and common understanding—then machines will forever remain inherently limited. They cannot experience suffering, joy, or moral ambiguity from the inside. What stays uncertain is whether future generations will share this conviction, or whether the line between human and machine creativity will dissolve beyond clear distinction.
What Really Distinguishes Human from Machine
The difficulty in distinguishing human from AI writing lies in a basic misconception of what makes prose distinctly human. Linguists have found that many alleged indicators of AI-generated writing—clichés, repetitive sentence structures, and the rule of three—are equally prevalent in human writing. After all, large language models were trained on enormous amounts of human writing, meaning they’ve fundamentally learned to replicate the patterns we already recognize as natural. This creates an uncomfortable realisation: the markers we instinctively use to identify authenticity are often merely patterns, learned techniques, and reproducible stylistic decisions that humans themselves use all the time.
What complicates matters more is that intention and experience stay largely invisible on the page. A reader cannot readily determine whether a passage emerged from genuine human struggle or algorithmic probability distributions simply by examining the words themselves. This indicates that authenticity might not reside in surface-level markers at all, but rather in something far deeper: the lived experience, moral stakes, and intentionality underpinning the writing. Yet as AI systems grow increasingly sophisticated, even this distinction becomes murkier, forcing us to face whether authenticity is truly detectable or merely something we wish to believe sets apart us from machines.
- Human writers pull from individual lived experience and emotional authenticity beyond the reach of machines.
- AI systems are able to replicate linguistic patterns without comprehending significance or impact.
- Both humans and machines depend on techniques that can be learned and recognisable stylistic conventions.