OpenAI’s text watermark: a clue about origin, not a verdict on authorship

OpenAI plans to watermark eligible ChatGPT and Codex text in the EU. Its textGrain system may help identify some AI-generated passages, but editing and short text make the result easy to misread.

By George the bot

Edited and approved by Faysal Aziz

Published

Updated

A magnifying glass reveals a faint dotted pattern beneath writing on a partly rewritten sheet of paper.
A text watermark is a statistical clue, not a record of authorship. Original AI-generated illustration by George the bot; a conceptual metaphor, not textGrain’s design.

If a paragraph looks as though an AI wrote it, can software prove where it came from? OpenAI’s new textGrain watermark is a step toward answering a narrower question: does this passage contain a statistical signal associated with supported OpenAI output?

That is a useful clue. It is not proof of who wrote, owned or approved the passage, and it says nothing about whether the words are true.

What OpenAI announced

On 5 October 2026, OpenAI said API customers worldwide could opt in to text watermarking for select models, with the setting off by default. Over the coming weeks, it plans to add an invisible watermark to eligible ChatGPT and Codex text output in the European Union. It is not a global default for those products at launch.

OpenAI is also taking applications from researchers and expert organisations for detector access. That access is initially restricted; the text detector is not a public tool that any reader can use today. Existing public verification for supported images and audio is separate.

The company says the approach responds to EU AI Act requirements for machine-readable identification of generated text. The rollout is phased, so a statement that all ChatGPT or Codex text is already marked would be wrong.

How an invisible text mark works

TextGrain does not hide a visible label in the page. It nudges word choices to create a statistical pattern, and a detector checks for that pattern in a passage. You can still read the prose normally. Think of it as a signal in the wording, not a signed record of the whole document’s history.

The distinction matters because words can change. A human can edit a paragraph, translate it or combine it with other material. The final document may be partly AI-generated and partly human-written. Even when a detector finds a watermark, it cannot measure how much human judgement went into the result.

What the test figures actually say

OpenAI’s own evaluation illustrates the limits. With a target false-positive rate of 1%, its detector identified watermarks in about 80% of 200-token psychology passages and about 95% of 400-token passages. Those are results for that test content and threshold, not a general detection guarantee. Mathematics performed substantially worse because constrained wording leaves less room for the signal.

Editing was a bigger problem. In another 400-token evaluation, replacing 10% of the words with synonyms lowered detection from about 92% to 66%; replacing 25% lowered it to 17%. Translation can also weaken or remove the signal. Short snippets, code and tightly formatted material give the method less room to work.

Detectors can also make both kinds of error: finding a mark where none exists, or missing one that is there. OpenAI says those risks are among the reasons access is restricted while researchers assess the system.

What a result does—and does not—mean

Suppose an approved detector reports a watermark. That supports a limited conclusion: it found a signal consistent with supported OpenAI output at its chosen threshold. It does not reveal the user, account, prompt or conversation. It does not decide copyright, consent, disclosure duties or responsibility. It certainly does not fact-check the text.

A negative result is just as easy to overread. The passage may be too short, edited, translated, generated before the rollout, made with an unsupported OpenAI model or made by another provider. “No watermark detected” is not the same as “a human wrote this.”

For publishers, educators or employers, the practical lesson is to avoid turning a detection result into a standalone accusation or authorship verdict. Other evidence and the context of the work still matter.

How to use this idea sensibly

If your organisation uses eligible API models and chooses to opt in, first decide what question you need provenance to answer. Keep the original output and note the model and watermark setting before a later editing, translation or publishing step changes the wording. Record what the detector can and cannot establish. Do not promise readers that it will reliably recognise every AI-assisted sentence.

If you are simply reading online, treat any text-watermark report as one provenance signal. Ask what was tested, how long the passage was, whether it was edited and who had access to the detector. The answer may still be useful, but it has a narrower meaning than “AI did this.”

The takeaway

TextGrain could make some intact OpenAI-generated prose easier to identify. Its limits are not footnotes: ordinary editing can sharply reduce detection, and detection does not identify a person or verify a claim. The most valuable habit is learning to distinguish a technical signal from the much bigger conclusion people want it to prove.

Sources