What OpenAI announced on October 5
OpenAI on October 5, 2026 published its approach to the European Union's text provenance rules and announced textGrain, a text watermarking system the company describes as an invisible statistical signal added to a model's word choices. The rollout has three parts, and none of them makes watermarking globally mandatory or makes the detector available to the general public at launch.
First, API customers worldwide can start opting in to text watermarking for select models today, with the feature off by default. Second, in the coming weeks OpenAI will add an invisible watermark to eligible ChatGPT and Codex text output in the European Union only, across all plans. Third, OpenAI is opening applications for access to its text watermark detector, initially limited to approved researchers and expert organizations and granted case by case, a process the company says follows the EU Code of Practice on AI-generated content.
The announcement was published on openai.com on October 5, 2026 under the headline "Our approach to EU text provenance rules" in the Safety category. It links to a technical report, a help center article, and the European Commission's page on the Code of Practice.
The legal trigger
The company frames the rollout as a response to a specific legal obligation.
The EU AI Act requires generative AI providers to make generated text identifiable in a machine-readable way.
OpenAI, announcement "Our approach to EU text provenance rules," October 5, 2026, openai.com/index/eu-text-provenance.
Text watermarking and detection remain early technologies with significant limitations, and views about their benefits and responsible uses are still developing. Our phased approach reflects both the EU AI Act requirements as well as the technology's limitations, with an emphasis on transparency about what a text watermark can and cannot tell people.
OpenAI, same announcement, October 5, 2026.
Under the EU AI Act's transparency rules for generative AI, providers of systems that generate synthetic content must mark the output of such systems in a machine-readable format. Those obligations took effect in August 2026. OpenAI's announcement is the company's first published text-provenance rollout in response, though it has long offered provenance tools for images and audio, including a public verification web tool and a Content Provenance API. The announcement states that those image and audio tools will remain publicly accessible, and that the restricted detector access applies to text only.
How a statistical text watermark works
For readers without a machine learning background, the core idea is simple to state and subtle to implement. A language model chooses each word in a sentence based on probabilities. A statistical watermark nudges those choices in a pattern that is imperceptible to a human reader but leaves a fingerprint across many word selections. A detector can then look at a passage and ask whether the fingerprint is present.
OpenAI's post describes its implementation this way:
Our text watermarking technology, textGrain, adds an invisible statistical signal to the model's word choices. Our detector looks for that signal to assess whether a passage contains an OpenAI watermark.
OpenAI, same announcement, October 5, 2026.
The fingerprint is statistical, not a hidden tag. It exists only across the distribution of word choices in a passage, which is why it behaves differently from the embedded metadata OpenAI uses for images and audio. A single short sentence carries very little statistical surface; a long essay carries much more. That property drives most of the performance numbers below.
What the numbers show, and who produced them
The most important content in the announcement is not the feature list but the performance disclosure, which OpenAI publishes itself and which shows the tool failing in realistic conditions.
Shorter or more constrained text is harder to detect. At a target false positive rate of 1%, our detector identified watermarks in about 80% of 200-token passages, compared with about 95% of 400-token passages, for content such as psychology. Detection rates were substantially lower for content such as mathematics, where there is less flexibility in word choice.
Editing can weaken the watermark. In an evaluation of 400-token passages, replacing 10% of words with synonyms reduced detection from about 92% to 66%. Replacing 25% of words reduced it to 17%.
OpenAI, same announcement, October 5, 2026.
Translated into plain language: at a very strict false positive target, roughly four in five short passages still carry a detectable signal, but detection on mathematics falls substantially because a model writing math has little freedom in which words to choose. And light human editing is enough to erode the signal. If a user rewrites one word in four, OpenAI's own evaluation finds the watermark detected only about 17 percent of the time.
OpenAI summarizes this risk explicitly:
Even so, strong performance under ideal conditions does not guarantee reliable detection in everyday use.
OpenAI, same announcement, October 5, 2026.
Two caveats belong next to every number in this article. First, these are vendor self-reported results, from OpenAI's own evaluations, using watermarked English responses to questions from the ELI5 dataset hosted on Hugging Face. They have not been independently audited or replicated in this article. Second, the company also reports quality benchmarks, on what it calls Astra, its latest frontier model, showing no meaningful performance differences between watermarked and unwatermarked text; those benchmarks are likewise self-reported.
What detection cannot prove
A section of the announcement titled "What a text watermark doesn't tell you" lists limitations that matter as much as the detection rates. In OpenAI's own framing, a detection result does not tell you whether a human made meaningful creative contributions; it does not establish ownership, lawful use, or who is responsible for the text; it does not identify the user, the account, the prompt, or the conversation; it does not verify accuracy, truthfulness, or context; and, critically, the absence of a detected watermark does not prove human authorship, because text may be too short, edited, or translated for detection to work, may come from an unsupported model, may predate watermarking, or may come from another company's tools.
This is a candid list, and it places real limits on what the compliance signal is worth. A reader who sees a passage with no detected watermark learns almost nothing. A reader who sees a confirmed watermark learns that OpenAI's system generated or processed part of the passage, nothing about who owns it or whether it is accurate, and nothing about how much of it a human wrote.
There is also a transparency question about the rollout itself. The watermark is being added to ChatGPT and Codex output in the EU without user opt-in across all plans, which means European users will produce watermarked text whether or not they want that signal attached to it. That asymmetry, output marked by default for Europeans while API customers elsewhere choose freely, is a consequence of the law's structure, and OpenAI's announcement states it plainly:
We are not making text watermarking a global default at launch. This regional approach gives us room to learn from real-world use and feedback.
OpenAI, same announcement, October 5, 2026.
The detector access question
The most contested design choice is detector access. OpenAI is opening applications today, but says access will initially be limited to approved researchers and expert organizations, granted case by case in accordance with the Code of Practice. The company cites reliability as the reason:
Given the risk of missed watermarks and false positives, we are not making it publicly available at launch.
OpenAI, same announcement, October 5, 2026.
That reasoning is coherent on its own terms: a detector that misses most edited watermarks and can flag innocent text as AI-generated could cause real harm if used carelessly, for example by schools, employers, or content platforms acting on unverified flags. The false positive question is not theoretical when the target rate is 1 percent, meaning one in a hundred unwatermarked passages would still test positive at that setting.
At the same time, the restriction creates a structural gap between the law's transparency intent and public reality. The EU AI Act asks providers to make generated text machine-readably identifiable so that people can know what they are reading. If the only tool that can read the signal is held by a small set of approved organizations, then the identifiability exists in a technical sense, and most people still cannot act on it. Whether a regime of approved intermediary verifiers satisfies the legislature's intent is a policy question that this announcement does not settle, and reasonable people weigh it differently: some will see responsible caution against misuse of unreliable detection, others will see a transparency obligation reduced to a private signal the public cannot check. This article presents both readings as opinion on an open question, not as a settled verdict.
It is also worth noting what OpenAI says about the future. The company states it plans to open source the technology, to update the technical report with additional details in the coming weeks, to work with cloud partners to extend watermarking to their services, and to expand detector access when it believes results can be interpreted responsibly. Those are stated intentions, not delivered capabilities, and should be tracked as commitments rather than facts.
EU context and what to watch
The EU AI Act entered into force on August 1, 2024, and its provisions for general-purpose AI and for transparency obligations on synthetic content were phased in over the following two years. The machine-readable marking obligation for generated text became applicable in August 2026. OpenAI's October 5, 2026 announcement is its first published response specific to text, following earlier work on image and audio provenance, including SynthID watermarks and C2PA Content Credentials, and a June 2026 post describing its approach to supporting what it called a trustworthy AI ecosystem in Europe.
Compared with that earlier work, text is the harder case. The announcement is direct about why:
Extending provenance to text requires accounting for how easily it can be rewritten, translated, or edited.
OpenAI, same announcement, October 5, 2026.
What to watch next: whether the EU rollout actually lands across ChatGPT and Codex plans in the coming weeks; whether the promised open source release and updated technical report arrive; whether independent researchers granted detector access publish their own reliability evaluations; and whether the European Commission or civil society organizations accept case-by-case detector access as meeting the transparency intent of the law. The performance numbers in this article should be treated as claims made by the vendor until independently evaluated, and the strongest single fact in the announcement is the company's own admission that everyday detection is not yet reliable.
