1 Oct 2026, Thu

Unmasking the AI Pen: New Research Reveals Evolving "Tells" in Large Language Model Prose

The digital landscape is now awash in text generated by artificial intelligence, a phenomenon that has sparked a keen interest among human readers and creators alike in developing the ability to distinguish AI-authored content from human writing. While earlier attempts to identify AI writing relied on easily recognizable markers like excessive em-dash usage or a predilection for the word "delve," these superficial tells have largely been eradicated by increasingly sophisticated language models. However, a groundbreaking new study from the marketing firm Graphite suggests that despite these advancements, AI models still exhibit distinct, albeit subtler, writing habits that betray their artificial origin.

This comprehensive research, detailed in a report accessible via a link to Graphite’s website, delved into the writing characteristics of leading frontier AI models. The study meticulously analyzed the linguistic patterns of these models, aiming to pinpoint their unique vocabulary preferences and stylistic tendencies. The findings indicate that while the most obvious giveaways, such as overuse of em-dashes, have been largely addressed by model developers, AI prose continues to fall back on specific structural and lexical patterns. Crucially, each iteration of an AI model demonstrates its own set of idiosyncratic "tells," offering a unique fingerprint for identification. The sheer breadth of these linguistic markers is perhaps the most surprising revelation; Graphite identified an astonishing 13,000 phrases that appeared at least twice as frequently in AI-generated content compared to human-authored text – their operational definition of a "tell."

Greg Druck, Chief AI Officer at Graphite, shared his insights with TechCrunch, highlighting a notable divergence in the development trajectories of major AI models. "It turns out that Claude models are actually getting closer to the human word distribution over time," Druck observed, suggesting a deliberate effort by Anthropic to imbue their AI with more natural human linguistic patterns. Conversely, he noted a concerning trend for OpenAI’s models: "And for the GPT models, it’s getting further away." This observation points to a potential widening gap in the ability of different AI developers to achieve authentic human-like prose.

The scientific rigor of Graphite’s study was paramount to its validity. To achieve a robust comparison, the researchers employed a carefully designed methodology. They began by compiling a substantial corpus of 10,000 articles published prior to the widespread release of ChatGPT. This collection served as a critical control group, representing a benchmark of genuine human-generated writing. Subsequently, the researchers tasked various AI models with rewriting these same articles based on provided summaries. This approach was designed to minimize potential bias stemming from the original source material, allowing for a more direct comparison of the AI’s inherent writing style. By generating parallel samples from both human and AI sources, Graphite’s team was able to quantitatively assess the frequency of specific words and phrases in AI writing, as well as identify broader patterns in sentence construction and rhetorical devices.

Delving into the specifics, the study revealed particularly striking "tells" for Anthropic’s Claude Opus 5.5. The word "dependable" emerged as its most prominent identifier, appearing a staggering 23 times more frequently than in human-written samples. While Opus 5.5 has apparently moved beyond the simplistic "it’s not X, it’s Y" sentence construction, a common pitfall in earlier AI models, it still exhibits a tendency towards a closely related pattern: "is more than an X, it’s a Y." This subtle linguistic tic suggests a lingering inclination to define concepts by contrasting them with a less sophisticated alternative, albeit in a more nuanced manner.

Furthermore, Opus 5.5 demonstrates a pronounced affinity for emphasizing the significance of its subject matter. The phrase "this matters" was found to be used an overwhelming 116 times more often in Opus 5.5’s output compared to human writing. Similarly, the construction "why X matters" appeared 92 times more frequently. This persistent emphasis on the importance of the topic at hand could be interpreted as an AI’s attempt to convey the perceived value or relevance of the information, a characteristic that, while not inherently negative, distinguishes it from more organic human discourse where such emphasis might be more varied.

OpenAI’s Astra, on the other hand, exhibits a distinct set of telltale signs. This model shows a penchant for introducing "another dimension" to the topics it discusses, a phrase that frequently surfaces in its prose. Additionally, Astra tends to qualify its assertions by stating that an action "may provide" or "can provide" a particular benefit. This hedging language, while often used by humans to express uncertainty or caution, appears to be a more consistent feature of Astra’s writing. Its most significant identifier, according to Graphite’s analysis, is what they term "corrective framing." This involves defining a topic not as "simply X" but as something else, or offering an alternative such as "rather than relying on X." The research indicates that these corrective framing constructions were over 100 times more prevalent in Astra-generated prose than in human writing, suggesting a systematic approach to defining concepts through negation or contrast.

A notable observation across all the frontier AI labs studied is their apparent response to the widespread criticism regarding the overuse of em-dashes. In Graphite’s samples, Opus 5.5 dramatically reduced its em-dash usage, employing it 99% less often than its predecessor. Astra has also significantly curbed its em-dash use, employing it 88% less than human samples. Gemini 3.1 Pro, another leading model, has reportedly almost entirely purged the em-dash from its generated text. This indicates a concerted effort by AI developers to address and rectify stylistic quirks that were readily identified by the public and researchers.

However, despite these targeted improvements in specific areas, Graphite’s research suggests that the overall number of AI "tells" remains remarkably stable. Druck elaborated on this point, stating, "It’s not like the tells are decreasing. They are managing to remove the most well-known tells, but other ones pop up. And every model version has its own." This suggests a dynamic rather than static evolution of AI writing characteristics. As developers eliminate one set of predictable markers, others emerge, often tied to the underlying architecture and training data of each specific model.

The persistence of these telltale linguistic patterns is particularly surprising given the stated goals of AI developers to produce writing that closely mimics human styles. In the release notes for Opus 5.5, Anthropic explicitly boasted that the model "communicates more naturally than prior models," with early users reportedly finding its writing "clearer and easier to follow." Similarly, OpenAI made comparable claims when introducing GPT-6 versions, Sol and Luna, stating that users could "expect to see more clarity, less jargon, [and] fewer odd turns of phrase." These marketing assertions highlight the industry’s commitment to achieving human-level fluency and naturalness.

Despite these ambitious claims, Druck remains skeptical about the ability of AI labs to completely eliminate these distinctive writing constructions and phrases. He posits a general hypothesis: "The labs are less able to control some of these things than you might expect." Druck attributes this limitation to the sheer scale and complexity of these models. "These are giant models with billions of parameters," he explained. "They have some finite number of tests they can run, and things slip through." This suggests that the intricate, emergent properties of massive neural networks may inherently lead to subtle, persistent deviations from human linguistic norms, even with extensive fine-tuning and development efforts. The vastness of the parameter space and the emergent behaviors within these models create a complex system where complete control over every linguistic nuance might be an unattainable ideal.

The implications of this research are far-reaching, impacting fields ranging from content creation and journalism to education and academic integrity. As AI-generated text becomes increasingly indistinguishable from human writing at a superficial level, the identification of these deeper, systemic "tells" becomes crucial for maintaining authenticity and trust in digital communication. The ongoing arms race between AI developers striving for human-like prose and researchers seeking to unmask AI authorship highlights the dynamic nature of artificial intelligence and its continuous evolution. The challenge for the future lies not just in identifying what AI writes, but in understanding the subtle, inherent characteristics that continue to differentiate it from the uniquely human capacity for expression. The pursuit of truly indistinguishable AI writing may be a technological frontier, but the subtle signatures it leaves behind offer a valuable window into the nature of machine learning and its ever-evolving relationship with human language.

Leave a Reply

Your email address will not be published. Required fields are marked *