A watermark shaped by European rules
OpenAI has outlined how it will add an invisible watermark to eligible text produced by ChatGPT and Codex in the European Union. The 5 October announcement responds to text-provenance requirements under the EU AI Act. Rollout is planned over the coming weeks. The signal changes model token choices in a pattern that an authorised detector can recognise, rather than adding a visible label to every paragraph.
The company presents watermarking as one component of a broader provenance approach, not proof that a passage was wholly created by AI. Text can be edited, translated, copied through another model or combined with human writing. A detector can therefore find a signal associated with eligible OpenAI output, but absence of the signal cannot establish human authorship. That limitation is important for schools, publishers and employers considering how to use the technology.
Quality and robustness pull in opposite directions
A stronger watermark is generally easier to detect but can constrain which words a model chooses. OpenAI says its selected setting has a modest effect on quality and publishes results from internal evaluations. It also says the impact varies with task and language. Highly constrained writing, code and short responses may provide less room to embed a reliable statistical pattern than longer natural-language output.
Robustness is the second trade-off. Rewriting or translating text can weaken the signal, and deliberate attackers may try to remove it. OpenAI acknowledges that no current text watermark survives every transformation. The system is better suited to supporting an investigation than automatically policing all online content. Organisations should avoid treating a negative result as clearance or a positive result as evidence of misconduct without context.
Verification will not be open to everyone
Unlike OpenAI's public verification tools for its generated images and audio, text-watermark checking will be limited to authorised organisations. The company cites the risk that unrestricted access would help people learn how to erase or evade the signal. Eligible organisations are expected to use it for specific integrity purposes and under controls intended to reduce misuse.
Restricted access creates its own accountability questions. People affected by a detection should know what was tested, which threshold was applied and how they can challenge a conclusion. Authorised verifiers need documented procedures, data retention limits and training on false positives and false negatives. A secret detector should not become an unreviewable basis for discipline, assessment or denial of an opportunity.
Privacy and authorship remain separate issues
An invisible watermark is embedded in generated text, not a record of the user's identity. It should not be interpreted as showing who requested the text, whether they edited it or whether use of AI was allowed. Provenance answers a narrow question about a technical signal. Authorship, originality and compliance with a policy require additional evidence and human judgement.
The rollout also needs clear privacy boundaries. A verifier may receive documents containing personal or confidential information, so access and logging should be proportionate. Organisations should minimise the text they submit and avoid centralising sensitive material merely to check for a watermark. OpenAI should disclose enough about governance and error rates for independent experts to assess whether the restricted system is being used fairly.
A limited signal can still be useful
Appeal procedures should be designed before the first consequential detection, not improvised after somebody disputes a result.
Implementation details will matter across different kinds of text. A long essay provides many token choices in which to carry a signal, while a brief caption, a rigid form or source code may not. OpenAI uses the word eligible, so users and verifiers need a clear account of which products, languages and output types are covered. They also need version information because detector performance may change as models and watermark settings evolve. Publishers that receive mixed human and AI drafts should preserve editorial history rather than relying on the final text alone. That record can show where assistance was used without forcing a probabilistic detector to answer questions it was not designed to resolve.
Text watermarking may help platforms study coordinated synthetic-content campaigns, support investigations into undisclosed bulk generation or meet a regulatory transparency requirement. It is less suited to deciding whether one student used an assistant for one assignment. The strongest application combines the signal with source records, editing history and an opportunity for the person involved to explain their process.
OpenAI's announcement is unusually direct about technical limits, which is welcome. The EU rollout creates a large real-world test of whether a moderate watermark provides useful provenance without noticeably degrading output. Success should be measured by validated investigations and low harm from mistaken conclusions, not by the number of texts flagged. Transparency about performance across languages will be particularly important in a multilingual European market.