Writing · AI
A little comparison between human writing — and AI writing…
The evolution of AI-assisted writing — and the gap between AI and humans — in the writing process.
Human writing and tools: what the research really shows
When comparing "the art of human writing" with the existence of writing tools, it is tempting to fall into a Manichean view — either celebrating the tool as a form of emancipation or lamenting it as a sign of decline. Recent academic literature, however, paints a more interesting and less binary picture: the difference lies not so much in the final product as in the cognitive process and the writer's stylistic signature.
Regarding the product itself, stylometric studies show that human writing and text generated by language models remain distinguishable. By applying Burrows' Delta to a balanced corpus of short stories, researchers observed that human texts form "broader and more heterogeneous" clusters — reflecting the diversity of individual expression — whereas model outputs, while fluent, exhibit "greater stylistic uniformity" (Nature — Humanities and Social Sciences Communications, 2025). Similarly, the study Art or Artifice?, presented at CHI 2024, adapted the Torrance Test to evaluate creativity as a product; it concluded that stories generated by LLMs pass creativity tests 3 to 10 times less frequently than stories written by professional authors (Chakrabarty et al., 2024). Ismayilzada, Stevenson, and van der Plas (arXiv, 2024) reinforce this finding: models produce linguistically complex texts but lag behind humans in terms of novelty, surprise, and diversity — precisely the qualities we tend to associate with "art."
Yet, the most provocative point concerns not the text itself, but the writer's brain. The study Your Brain on ChatGPT by the MIT Media Lab (Kosmyna et al., 2025) used EEG to measure the neural activity of three groups: those writing with an LLM, those using a search engine, and those using no tools at all. The "brain-only" group displayed the "strongest and most distributed" connectivity networks, whereas LLM users exhibited the weakest connectivity. The authors coin the term "cognitive debt": mental activity "scales down" in proportion to the use of the external tool. This resonates with the philosophical analysis suggesting that tools like ChatGPT and Word Copilot can "disburden" the user of the effort inherent in reflective writing — drawing on Borgmann's "device paradigm" and Ihde's post-phenomenology (AI & Society, 2025).
However — and here lies the nuance that prevents a simplistic, pessimistic interpretation — the same literature shows that the effect depends on the design of the interaction, not the tool itself. Dhillon et al. (CHI 2024, N=131) found a "U-shaped" impact: sentence suggestions (low scaffolding) offer little help, whereas paragraph suggestions (high scaffolding) significantly improve quality and productivity, especially for infrequent writers. Li, Liang, Peng, and Yin (CHI 2024) observe that people are willing to forgo financial payment to access AI assistance, gaining productivity and confidence — albeit at the cost of a reduced sense of responsibility and less textual diversity. In education, there is evidence of tangible gains: Khojasteh et al. (Frontiers in Education, 2024) recorded significant improvements — with large effect sizes — in the content, organization, and vocabulary of medical students who used ChatGPT.
The common thread linking all this is the distinction between a tool that replaces the user and one that works alongside them. Research by Hwang et al. (CSCW 2025) on co-writing reveals that authors locate authenticity not in the final output, but in the "process and experience of constructing one's authorial self" — the classic "80% me, 20% AI" sentiment. In this view, human writing is not threatened by the tool as long as the tool mediates thought rather than executing it in the author's place. It is the difference between Copilot automatically producing a document and ChatGPT being used conversationally to interrogate one's own beliefs (AI & Society, 2025).
Citeable references: Nature, Humanities & Social Sciences Communications (2025) — stylometric comparison · Chakrabarty et al. (CHI 2024), Art or Artifice? — arXiv:2309.14556 · Ismayilzada, Stevenson & van der Plas (2024) — arXiv:2411.02316 · Kosmyna et al. (2025), Your Brain on ChatGPT — arXiv:2506.08872 (MIT Media Lab) · Dhillon et al. (CHI 2024) — arXiv:2402.11723 / DOI 10.1145/3613904.3642134 · Li, Liang, Peng & Yin (CHI 2024) · Hwang et al. (CSCW 2025) — arXiv:2411.13032 · Khojasteh et al. (2024), Frontiers in Education, 9 — DOI 10.3389/feduc.2024.1457744.
The Psychology of Writing: How AI (Doesn't) Understand the Human Mind
An exploration of AI's cognitive limitations in literary creation.
Paragraph 1 (Theory): Literary writing is not just about what is said, but how and why it is said. A human author writes based on:
- Memories (personal experiences that color the narrative)
- Traumas (wounds that give depth to characters)
- Desires (hidden motivations that drive the plot)
AI, on the other hand, simulates these layers based on data but does not live them. When a model generates dialogue like "— I love you, but I can't stay," it lacks the subtextual tension a human author would inject: the tremor in the voice, the averted gaze, the object in hand being squeezed unnecessarily.
This intuition is not merely a reader's impression; it has been measured in the laboratory. Stylometric studies applying Burrows' Delta to short story corpora show that human texts form "broader and more heterogeneous" clusters — reflecting the diversity of individual experience — whereas model outputs, however fluent, exhibit "greater stylistic uniformity" (Nature, Humanities & Social Sciences Communications, 2025). And when creativity is evaluated as a product, the study Art or Artifice? (Chakrabarty et al., CHI 2024) found that LLM-generated stories pass creativity tests 3 to 10 times less frequently than those by professional authors, falling short specifically in novelty and surprise (Ismayilzada et al., 2024). In other words: what machine-generated dialogue lacks is measurable — and lies precisely in the layers of intention described above.
Paragraph 2 (FolioTest Solution): At FolioTest, we use "literary prompt engineering" techniques to compensate for this limitation. For example:
- Layer 1 (Surface): "Write a farewell dialogue between two lovers."
- Layer 2 (Deep): "Write the dialogue, but include: (a) a symbolic object one of them is holding, (b) a lie both know to be a lie, (c) a promise that will not be kept."
Paragraph 3 (Example): A FolioTest user generated this excerpt using our suggestions:
"He held the pocket watch she had given him for their tenth wedding anniversary. 'I'm not leaving,' he said, as the second hand ticked like a heart about to stop. She smiled, knowing he had already bought the train ticket to Paris."
Here, the depth comes from details the AI wouldn't have invented on its own: the watch (symbolizing time running out), the lie ("I'm not leaving"), and the irony of the purchased ticket.
Note the mechanism: the density didn't originate from the model itself, but from the scaffolding imposed by the prompt. This aligns directly with the findings of Dhillon et al. (CHI 2024, N=131), who identified a "U-shaped" effect in co-writing: shallow suggestions barely improve the text, whereas structured, high-level suggestions significantly elevate its quality. That is the philosophy behind our tool.
→ Connection to the original: "…and our chat was designed around this awareness."
Although AI may seem far removed from human writing, there are already numerous AIs and platforms dedicated to general, literary, and creative writing.
However, we do not believe AI will lag behind human writing for long. It is only a matter of time before all AIs achieve the full status of writing considered "human."