OpenAI Will Watermark ChatGPT and Codex Text in the EU, With a Worldwide API Opt-In
TL;DR
OpenAI said on October 5 that it will add an invisible statistical watermark, which it calls textGrain, to eligible ChatGPT and Codex text generated in the European Union over the coming weeks. API customers anywhere can opt in for select models starting now, but the setting is off by default and OpenAI says it "will not become a global default at launch." The detector stays with approved researchers and expert organizations. OpenAI's own figures: about 95% detection on 400-token passages at a 1% false-positive target, 80% at 200 tokens, and a slide from 92% to 66% once 10% of the words are swapped for synonyms. A nine-author technical report explains the mechanism: optimal transport over keyed vocabulary blocks, with an explicit budget for how much sampling randomness the watermark is allowed to remove.
What ships, where, and to whom
The rollout, per OpenAI's announcement and the staff post on its developer forum, has three tiers. EU users of ChatGPT and Codex get the watermark on eligible text "over the coming weeks," on every plan, with no toggle mentioned. API customers worldwide can opt in today for select models, off by default. Cloud partners get the option in the coming weeks, scope unstated.
"Eligible" is doing work in that sentence. OpenAI says the EU Code of Practice does not require watermarks on outputs shorter than 200 tokens, roughly 150 English words, or on code snippets, and that code is harder to watermark anyway because there are fewer plausible choices for what comes next. So a one-line answer, a regex, and a SQL query likely carry nothing. A 600-word essay does.
The watermark can indicate that an OpenAI system generated or processed part of a text. OpenAI is explicit about what it cannot do: it cannot tell who used the system, how much a person contributed, who owns the text, or whether the content is accurate. And a missing watermark does not prove a human wrote anything, since the text might be short, edited, translated, or from another model entirely.
How textGrain works
A language model samples each token from its next-token distribution. Every text watermark since Kirchenbauer et al. and Scott Aaronson's 2022 OpenAI scheme works the same way at the top level: use a secret key plus the last few tokens to generate pseudorandom values, then let those values tilt which token gets picked. The detector regenerates the same values from the text and key and tests whether the chosen tokens line up with them more often than chance.
textGrain's twist, per the technical report by Xiang Li, Garrett Wen, Xiaohong Chen and Qi Long (Penn and Yale) with five OpenAI authors, is how it tilts. The key and a context window of recent tokens partition the vocabulary into blocks; each block inherits the summed probability of its tokens. A keyed cost table pairs those blocks with a handful of equally likely "columns," and one column is selected pseudorandomly. An optimal-transport solver chooses a coupling that favors low-cost block-column pairs, while a KL penalty caps how far the coupling drifts from independence. Within the chosen block, tokens keep their original relative probabilities.
The result the paper leans on: that KL penalty is mathematically identical to the average entropy removed from sampling. So the watermark's strength is literally a budget, a fraction of the model's next-token entropy that the watermark may spend. Averaged over all possible keys, the output distribution is unchanged, which the paper calls unbiased.
If that is hard to picture: imagine a casino allowed to weight its dice, but only a little, and the weighting changes every roll according to a secret schedule. Any single roll still looks fair. Across a few hundred rolls, someone holding the schedule can see the lean; nobody else can, and averaged across all possible schedules the dice really are fair.
Why not just use Aaronson's older Gumbel-max trick? The report says it plainly: at a fixed key and context, Gumbel-max always picks the same token, so the same prompt produces identical responses, which matters when an application needs several distinct answers. textGrain keeps some sampling randomness under the budget. The report also includes a speculative-sampling variant with a shared key, which is the version you would actually deploy on a serving stack that cares about throughput.
Detection is a statistical test, so length is everything
The detector needs only the text and the key. It does not need the model, the coupling, or the entropy budget used at generation time. At each scored position it rebuilds the partition, the cost table and the selected column from the key and the preceding tokens, finds which block the observed token sits in, and converts that block's cost into a score. Under unwatermarked text each score is a unit exponential variable and their sum follows a Gamma distribution, so "is this watermarked?" becomes "does the sum exceed the 99th percentile of that null?" That is the 1% false-positive target in OpenAI's numbers. Only the first occurrence of each distinct context window is scored, to avoid double-counting repeated phrases.
This is also why short text is a dead end: it is checking whether a coin is loaded by flipping it 400 times rather than 40. OpenAI reports about 80% detection on 200-token psychology answers and about 95% at 400 tokens, with "substantially lower" rates on math, where the model has little freedom in word choice.
Edits hurt faster than length helps. In OpenAI's test on 400-token passages, detection ran at 92% on untouched text, 66% after 10% of the words were replaced with synonyms, and 17% after 25% were replaced. Translation, shortening and rewriting all degrade the signal. OpenAI itself cautions that strong results under ideal conditions do not guarantee reliable detection in everyday use.
Quality cost: none OpenAI could measure
OpenAI ran its newest frontier model, GPT-6 Astra, with and without the watermark. On the Artificial Analysis Intelligence Index the two runs landed within 0.2 points of each other (49.57 and 49.76), and on GPQA Diamond the watermarked run scored 93.94% against 94.44% unwatermarked. OpenAI calls that no meaningful difference. It is a self-reported result on a single model.
The rule behind it
Article 50(2) of the EU AI Act requires providers of generative systems to ensure outputs are "marked in a machine-readable format and detectable as artificially generated or manipulated," with solutions that are "effective, interoperable, robust and reliable as far as this is technically feasible." It applies from August 2, 2026. The Commission's Code of Practice on marking and labelling AI-generated content, published June 10, 2026 and found adequate by the Commission and the AI Board in July, is the operating manual.
Three Code provisions explain OpenAI's design choices, according to Gleiss Lutz's analysis. For free-form text, a watermark alone satisfies the marking duty, because text cannot carry metadata the way a PNG can. Reliability requirements are less strict for texts under 200 tokens, and where a mark is less reliable, access to the detection results may be restricted to verified expert users, which is the hook OpenAI hangs its researcher-only detector on. And providers must implement an interoperability solution for watermarks by February 2, 2027, which is the part nobody has publicly solved.
Three labs, three answers to the same law
The same deadline produced three different products. Anthropic went furthest: as AI Bacon covered in August, Claude models launched on or after August 2 weave a text watermark into output "wherever Claude is offered, worldwide," and its help article lists no opt-out, with detection in a private preview for regulators, law enforcement and researchers. Google published SynthID Text in Nature in October 2024 and open-sourced the reference implementation, so anyone can study the method, although Google's production keys stay private.
OpenAI drew the narrowest line it could: consumer products marked only in the EU, the API off by default everywhere, no global default "at launch." This is the same company that retired its own AI-text classifier in 2023 for being too inaccurate to trust. It now ships a detector, with the caveat that it does not work on short text, edited text, math, code, or any other lab's model. The math is better this time; the candor about the limits is the real upgrade.
What this means if you build on the API
- The default-off setting moves the decision to you. If you serve EU users through the API and ship long-form generated text, read the Code before assuming your vendor has covered Article 50 on your behalf. The opt-in exists for exactly that case.
- You cannot test it yourself. Detection needs the key, and the key is OpenAI's. The content provenance docs currently cover images and audio and say text verification is available only to approved organizations by application. Do not build an "AI detector" product on a signal you cannot read.
- Short outputs and code are mostly outside the blast radius. Under 200 tokens the Code relaxes, and OpenAI says code snippets are not required to carry a mark. Chat UIs with long answers are where this lives.
- Diversity is budgeted, not free. Under a fixed key the watermark spends a fraction of the model's sampling entropy. textGrain keeps more variety than Gumbel-max, but if your product relies on sampling many distinct completions from one prompt, measure it.
- Absence proves nothing. OpenAI says so itself. A missing watermark is consistent with a human, a short answer, a paraphrase pass, or a different model.
Caveats
Every detection and benchmark number above comes from OpenAI's own evaluation; no third party has run the detector. OpenAI has not published which "select models" support the API opt-in, the parameter name, or which cloud partners are getting it, and the public API docs had no text-watermark parameter at the time of writing. "Eligible" text in ChatGPT and Codex is defined by OpenAI's reading of the Code, not by a published spec. The technical report proves the entropy-budget identity and the unbiasedness property under stated assumptions; it does not report the empirical detection figures, which live only in the announcement.
Key Takeaways
- OpenAI's textGrain watermark reaches ChatGPT and Codex text in the EU over the coming weeks; API customers worldwide can opt in now, off by default, with no global default planned at launch.
- The mechanism is optimal transport over keyed vocabulary blocks with an entropy budget, and the detector needs only the text and the secret key.
- OpenAI reports roughly 80% detection at 200 tokens and 95% at 400 at a 1% false-positive target, dropping to 66% after 10% synonym swaps and 17% after 25%.
- Benchmark cost on GPT-6 Astra was within 0.2 points on the Artificial Analysis index and 0.5 points on GPQA Diamond, by OpenAI's own measurement.
- Anthropic marks all Claude text worldwide with no opt-out; Google open-sourced SynthID Text; OpenAI marks only where the EU requires it and keeps the detector gated.
- The watermark cannot identify a user, prove ownership, or verify accuracy, and a missing watermark is not evidence of human authorship.
Sources: OpenAI, Our approach to EU text provenance rules, textGrain technical report (PDF), OpenAI developer forum announcement, OpenAI content provenance docs, TechCrunch, Unite.AI, Tech Startups, EU AI Act Article 50, European Commission Code of Practice, Gleiss Lutz, IPTC, Anthropic help center, SynthID Text in Nature, synthid-text on GitHub, Kirchenbauer et al. 2023, Scott Aaronson