EN/FR

How can we tell whether a text was written by AI? The answer seemed obvious: just mark it. OpenAI has just announced the gradual rollout of textGrain, a watermarking system for text generated by ChatGPT and Codex in the European Union.1 Anthropic, for its part, has presented its own watermarking mechanism for Claude.2

The idea is appealing. When a model generates text, it constantly chooses between several possible words or word fragments, often natural and more or less equivalent. The watermark subtly influences some of those choices in order to leave, over a sufficiently large amount of text, a statistical fingerprint.

That fingerprint is invisible: no special spaces, no hidden Unicode characters, no mysterious string to uncover in a hex editor. The text looks perfectly normal, and only a detector that knows the mechanism can, in theory, recover the signature. The technique is reminiscent of steganography: the information is not visible in the content itself, but hidden in the way that content is produced.

On paper, it is elegant. In practice, things quickly become complicated.

Generation fingerprint, actual contribution

Let's start with a simple example. I write:

I took this smart plug apart and found a BK7238.

I then ask ChatGPT to improve the wording slightly. It suggests:

When I took the smart plug apart, I found a BK7238.

What actually came from the AI? Not taking the plug apart, not discovering the component, not the technical observation: all of that remains human. The AI only changed the wording.

At the other extreme, someone could ask:

Write me a complete 2,000-word article about hacking a smart plug built around a BK7238.

and then publish the result almost as-is.

In both cases, AI was involved. Intellectually, however, they are radically different approaches, and that is precisely what a watermark cannot tell us. It measures a generation fingerprint. It does not measure the AI's actual contribution to the content.

Everything that follows in this article stems from that distinction.

Fragile traces: model chains, synonyms, local models

A chain with no memory

Take a scenario that is becoming commonplace: Claude → ChatGPT → publication.

Claude writes a first version, which may carry Anthropic's statistical signature. I then ask ChatGPT to make it more fluid and remove repetition without changing the substance. It regenerates some of the sentences according to its own probabilities: Claude's watermark may be partially degraded, while OpenAI's may appear. After a light rewrite, both traces may still be detectable. After a deeper rewrite, only the latest signature may remain — or none at all.

Now add a small Qwen model running locally: Claude → ChatGPT → Qwen. This third model has no reason to preserve the previous signatures. It can rephrase enough words to degrade both watermarks while keeping the same ideas, the same structure and the same arguments. By the end of the chain, the text may no longer contain any detectable signature, even though its substance still comes largely from the first generation.

This is one of the system's most important limitations: statistical watermarks do not create a chain of provenance. They leave traces, and the next step can alter them.

A few synonyms may be enough

You do not even have to target the watermark. An instruction as ordinary as “replace some words with synonyms without changing the meaning” can already weaken the signal considerably. OpenAI itself says that, on 400-token passages, replacing about 10% of the words with synonyms reduced detection from roughly 92% to 66%. At 25% replacement, it fell to around 17%.3

No sophisticated attack, no knowledge of the mechanism and no specialised tool are required: a perfectly ordinary editorial workflow is enough. Asking an LLM to avoid repetition, vary the vocabulary, shorten a few passages or simply make the prose sound more natural can alter the fingerprint, without any intention of bypassing anything. And that transformation can now be carried out by a lightweight model on a home computer.

My point is therefore not to explain how to remove a watermark. It is to observe that a watermark can weaken simply because a text has been edited in a perfectly normal way.

The reverse problem: a human text can carry an AI trace

The reverse problem seems even more important to me.

I write an article entirely myself. I spend several hours on it. The experiments are mine, and so are the conclusions: I select the information, build the argument and decide what I want to say. But I am not a journalist. I can know exactly what I want to express while being less comfortable constructing certain sentences or organising a long piece. So I ask an AI:

Can you make this paragraph flow better?

or:

This section is confusing. Help me structure the argument more clearly.

At what point does the text become an “AI text”? After a few corrections? One rephrased sentence? One restructured paragraph? An introduction suggested by the model? There is obviously no simple answer.

We also need to avoid the opposite excess. A very specific task, such as correcting punctuation or a few typos, gives the model very little freedom in its word choices. The statistical signal will therefore be weak or difficult to establish, precisely because there are few possible variations.

The relevant question is not whether AI was involved somewhere in the workflow. It is what it actually did.

AI is a spectrum, not a switch

We still often talk about AI use as if it were a binary variable: yes or no. Reality looks much more like a spectrum.

At one end:

Write me a 1,500-word article on this topic.

A few seconds later, the content is published almost unchanged. At the other end, an author runs their own experiments, thinks through the subject, gathers information, builds the argument, writes a draft, then asks an AI to improve certain passages. Both have technically “used ChatGPT”. Yet they are profoundly different approaches.

The right question is therefore not simply was AI involved? but what did it actually do? Did it produce the ideas? The reasoning? The facts? The structure? The prose? Did it suggest counterarguments? Or did it simply make a few sentences clearer? That information is far more useful than a label saying “AI involved”.

What about human assistance?

This situation also reveals a strange asymmetry. Authors have always relied on proofreaders, editors, collaborators, press officers or ghostwriters, without readers necessarily knowing how far their contribution went. A text can be heavily reworked by someone else and still be published under the author's name alone.

That does not mean we should give up on transparency around AI. But why can human editorial assistance remain invisible while artificial assistance should automatically become suspect? The comparison is not perfect. It does, however, remind us that authorship has always been more complex than “the person who typed every word”.

LLMs did not create this problem. They made it visible.

Watermarks and detectors: two tools, two questions

We need to distinguish two things that public debate often conflates.

A watermark is a signal introduced at generation time by the model provider, which the provider can later look for. Post-hoc detectors such as Pangram or GPTZero work differently: they are statistical classifiers that analyse arbitrary text without knowing where it came from, and try to identify characteristics they associate with AI-generated writing.

The two approaches rely on different principles and have different error rates. But their social use eventually converges: they produce a signal that someone will have to interpret. And that is where things get difficult.

Watermark Open AI Anthropic En
OpenAI and Anthropic have chosen statistical watermarking approaches that are similar in principle, but each relies on its own mechanisms, keys, and verification tools.

Two misleading signals: Orélien and Hugo

“Written with AI”: the Thélyson Orélien case

The controversy surrounding Thélyson Orélien and his novel C'était ça ou mourir illustrates the problem well. The book was accused of having been produced to a large extent with AI assistance, notably on the basis of analyses carried out with Pangram. The author and his publishers disputed those claims. The novel had nevertheless already found a wide readership — Le Monde reported 35,000 copies sold in five weeks —, won the Fnac Novel Prize and appeared on the Goncourt, Renaudot and Femina longlists before the controversy erupted.4

I am not making any claim about how it was actually written: I do not know, and a score is not enough to tell us. What interests me is the question that remains even if substantial AI involvement were one day established. What exactly would “written with AI” mean? That the model wrote some sentences? Rephrased an existing manuscript? Suggested passages? Helped with the structure? Produced a first draft that was later deeply rewritten? Did it invent the characters, decide what to keep, discard the bad ideas, choose the pacing, determine when the book was finished?

No detector, however good, can reconstruct that chain. And that points to an obvious fact: using a tool does not, by itself, explain the quality of a work. Creating also means selecting, evaluating, rejecting, ordering and taking responsibility for the result.

If producing a novel appreciated by thousands of readers and capable of attracting a literary jury's attention were as simple as asking an LLM to “write me a great novel”, candidates for major literary prizes should already number in the thousands. The quip obviously proves nothing about the Orélien case. It simply shows that the question “was AI involved?” is not enough to explain a work.

When Victor Hugo becomes suspicious

At the other extreme, texts written long before ChatGPT have already been flagged by some detectors as showing characteristics associated with artificial generation. Victor Hugo regularly appears among the examples.

Again, caution is needed: not every detector gets it wrong. In one recent test, Clubic compared Pangram 4.0 and GPTZero 4.1m across ten French texts. Demain, dès l'aube was correctly classified as human by both tools, although Pangram indicated limited confidence because of the poem's short length.5 But another Clubic article also reported a LinkedIn test in which Victor Hugo was, this time, flagged as AI.6

We therefore end up in a strange situation. On one side, a certainly human text written before AI can be judged statistically suspicious by some tools. On the other, a text genuinely generated or heavily reworked by AI can become much harder to identify after paraphrasing, translation or rewriting.

Detectors do not all tell the same story either. In that same Clubic test, a human-written article heavily rewritten by ChatGPT was classified as “Mixed / AI Detected” by Pangram (38% AI / 62% human), while GPTZero judged it to be 98% human with high confidence. On another text generated entirely by ChatGPT, Pangram reported 100% AI, while GPTZero returned an uncertain result: 40% AI, 1% mixed and 59% human.5 That does not mean these tools are bad. It means the question we are asking them to answer is extraordinarily difficult.

A detector sees the final text, not its history

This may be the central point. A detector sees neither the first draft, nor the document history, nor the prompts, nor the human corrections, nor the deleted passages, nor the rejected suggestions, nor the research or experiments, nor the successive exchanges with several models. It sees only one object: the final text. And it tries to infer its history from that.

Yet very different histories can produce statistically similar texts:

  • a human text deeply rephrased by ChatGPT;
  • a ChatGPT text deeply rewritten by a human;
  • a Claude text passed through ChatGPT and then Qwen;
  • a human text translated and then translated back by several models;
  • a text generated entirely by AI, then modified just enough to lose its signature.

The detector is, in a sense, trying to reconstruct the film from its final frame. It may find clues. It cannot necessarily tell us what actually happened.

Detecting involvement is not establishing authorship

This is, in my view, where the main danger lies. A tool may produce a technical statement:

This passage probably contains a signature associated with an OpenAI model.

The reader may then conclude:

ChatGPT wrote this text.

Those are not the same statement. A watermark may indicate that a model generated or processed enough text to leave a trace. It does not tell us who had the ideas, who did the research, who carried out the experiments, who made the editorial decisions, what proportion of the intellectual work was human, or who takes responsibility for the content.

OpenAI itself states that a watermark does not measure human contribution, does not establish ownership or responsibility, does not identify the user and does not verify the accuracy of the content. It also says that failing to detect a watermark does not prove that a human wrote the text.7 Anthropic describes a very similar limitation: its watermark can only indicate that Claude was probably involved in the content; it cannot distinguish between “Claude wrote this” and “Claude heavily edited this”, and says nothing about authorship or ownership.8

A watermark can detect involvement. It cannot establish authorship.

Who will verify the signatures?

Another difficulty appears quickly. OpenAI is developing textGrain, Anthropic has its own system, and other providers will do something else. Each mechanism has its own parameters, its own key and its own detector. We therefore do not have a universal standard that can simply answer “this text contains an AI signature”. Instead we have:

OpenAI detects an OpenAI signal.

Anthropic detects an Anthropic signal.

What if both models were involved? What if they return contradictory results? What if a third model subsequently rewrote the text? What if a provider's detector is positive while a third-party tool finds nothing? Who decides what that means?

Once again, the technology produces information. It does not produce its interpretation.

Plagiarism, scores and algorithmic ostracism

Plagiarism-detection software has already taught us the dangers of placing too much trust in a score. A tool detects similarity, but someone still has to interpret it: a properly referenced quotation, an unavoidable technical phrase, a legal text, actual copying? The score is only the beginning of the analysis.

With plagiarism, we can at least place the passages side by side: here is the published text, here is the source, here is what matches. With a generic AI-content detector, the “evidence” becomes much more abstract:

This text displays statistical characteristics associated with AI-generated writing.

That does not make these tools useless. It simply means they should remain what they are: indicators, not judges.

What worries me more is the moment when a teacher, publisher or employer sees “87% probability that this text was generated by AI”. How many will actually take the time to understand what that number measures? The temptation is strong to turn it into a verdict: cheating, fraud, laziness, artificial content. Yet the score says none of those things.

A student may have written the entire paper before asking an AI to improve a few sentences. A journalist may ask it to rephrase an introduction. A researcher may use it to translate their own work. An author may use AI as an editor available at any hour. And conversely, someone can have a model produce almost the entire text and then transform it enough to confuse certain detectors.

Treating all these situations in the same way would amount to creating a form of ostracism based on a misunderstood technical indicator. False positives make that prospect even more troubling.

Images and decisions: when the real question moves elsewhere

For a realistic image, the problem is different

I am nevertheless strongly in favour of transparency when AI can mislead our perception of reality. A photorealistic video showing an event that never happened, or a fake realistic photograph of a person, a conflict or a disaster, should clearly be identified as artificial or manipulated. In that case, the label “AI-generated” answers an essential question: did what I am looking at actually happen?

For an infographic, the situation is different. Nobody mistakes a diagram showing tokens, neurons or a microchip for footage of a real event. In that case, the tool used to create the illustration matters far less to me than the accuracy of what it shows.

The relevant question may therefore be: does the use of AI change what the public believes to be real?

And what about decisions?

The issue becomes even more interesting when no content is generated at all. A company uses a model to rank CVs. A bank uses one to assess an application. A doctor consults a recommendation produced by AI. A manager compares several strategies with an LLM before making a decision. What watermark could reveal those interventions? None.

The final decision may be written entirely by a human even though AI played a far more important role than it would have in rephrasing a paragraph. We spend a great deal of effort trying to spot AI in visible outputs, while some of its most significant interventions happen upstream. In those situations, what we need is not a watermark but genuine traceability of the decision-making process.

Transparency, yes — but what kind?

It is legitimate to want greater transparency around the use of artificial intelligence. I would be uncomfortable with someone generating dozens of articles with a model while letting readers believe they were the result of personal research and reflection.

But that transparency should probably focus on the process, not merely on the presence or absence of a fingerprint. There is a world of difference between:

Article entirely generated by AI.

and:

The author used AI for research, challenging ideas, proofreading and writing assistance. The experiments, analysis and editorial choices remain the author's own.

The second statement actually tells the reader something. An invisible watermark tells them far less.

So who wrote this text?

That is ultimately the question textGrain led me to ask myself. If I bring personal experience, a line of thought, arguments, objections and editorial choices, but a tool helps me express them more clearly, am I still the author? In my view, yes.

That does not mean the tool's involvement should be hidden. It means that the author of an idea and the author of every word used to express it are not necessarily the same person. That was already true before artificial intelligence. LLMs have simply made that boundary more visible — and perhaps more uncomfortable.

A grain in the gears?

Systems such as textGrain start from an understandable goal: bringing greater transparency to a world where producing synthetic content has become trivial. The risk is that a technical signal is quickly loaded with meaning it does not actually carry.

A watermark can say:

A model probably left a trace in this text.

It cannot answer:

Who actually thought through this text? How much of the work is human? Did AI replace the author, or merely assist them? Does this content deserve less trust?

Those questions require context, judgement and probably a new way of describing how we create things.

Transparency requires context, not just a signal.

And if we forget that, the grain meant to make the machine more transparent may end up getting caught in its gears.

Sources


Process note: the ideas, examples, experiments and thesis in this article are the author's own. AI was used to analyse its structure, improve the flow of the writing and gather sources.

Footnotes

  1. OpenAI, Our approach to EU text provenance rules, October 5, 2026. OpenAI announces a gradual rollout of the textGrain watermark across eligible ChatGPT and Codex outputs in the European Union, along with initial detector access limited to approved researchers and expert organisations. https://openai.com/index/eu-text-provenance/ ↩

  2. Anthropic, How Claude's text watermark works, August 14, 2026, updated September 1, 2026. Anthropic describes its statistical watermark for Claude, states that it adds no hidden characters and explains how verification works. https://www.anthropic.com/news/claude-text-watermark ↩

  3. OpenAI, Our approach to EU text provenance rules, section “How our watermarking works and performs”. On 400-token English passages, replacing 10% of the words with synonyms reduces detection from about 92% to 66%; with 25% replacement, it falls to 17%. https://openai.com/index/eu-text-provenance/ ↩

  4. Le Monde, “Thélyson Orélien accusé d'avoir utilisé l'IA pour son livre ‘C'était ça ou mourir’”, September 22, 2026. The article reports the accusations based in part on Pangram, the response from the author and his publishers, around 35,000 copies sold in five weeks, the Fnac Novel Prize and the Goncourt, Renaudot and Femina longlists. https://www.lemonde.fr/pixels/article/2026/09/22/thelyson-orelien-accuse-d-avoir-utilise-l-ia-pour-son-livre-c-etait-ca-ou-mourir_6780300_4408996.html ↩

  5. Clubic, “Affaire Thélyson Orélien : que prouve vraiment un score à 97 % IA ? Nous avons fait le test”, October 4, 2026. The test compares Pangram 4.0 and GPTZero 4.1m across ten French texts, including human-written texts, texts generated entirely by ChatGPT, hybrid cases and Demain, dès l'aube. https://www.clubic.com/dossier-632419-affaire-thelyson-orelien-que-prouve-vraiment-un-score-a-97-ia-nous-avons-fait-le-test.html ↩ ↩2

  6. Clubic, “IA & littérature : après l'affaire Orélien, les auteurs classiques sont soumis au détecteur, et c'est surprenant”, September 26, 2026. The article reports, among other things, a LinkedIn experiment in which Victor Hugo was flagged as AI while Maupassant and Baudelaire were classified as human by Pangram. https://www.clubic.com/actualite-631387-ia-litterature-apres-l-affaire-orelien-les-auteurs-classiques-sont-soumis-au-detecteur-et-c-est-surprenant.html ↩

  7. OpenAI, Our approach to EU text provenance rules, section “What a text watermark doesn't tell you”. OpenAI states that a watermark does not measure human contribution, does not establish ownership or responsibility, does not identify the user, does not verify accuracy, and that the absence of a watermark does not prove human authorship. https://openai.com/index/eu-text-provenance/ ↩

  8. Anthropic, How Claude's text watermark works, sections “What does a watermark actually prove?” and “Does this change who owns a given output, or who is legally responsible for it?”. Anthropic states that the watermark can only show that Claude was probably involved, without distinguishing between writing and heavy editing, and without establishing authorship or ownership. https://www.anthropic.com/news/claude-text-watermark ↩

Comments

No comments yet.

Add a comment

Your email address will not be published.

Comments are moderated before publication.