New Delhi: OpenAI has announced the deployment of an invisible text watermarking system called textGrain, designed to identify and verify AI-generated content in compliance with the European Union’s evolving artificial intelligence provenance regulations. Unlike digital watermarks applied to images or video, textGrain embeds undetectable statistical signals directly into the model’s linguistic choices without introducing extraneous characters, unusual symbols, or altered whitespace. The feature is rolling out across eligible ChatGPT and Codex outputs within the EU over the coming weeks, while remaining available as an opt-in integration for global API customers.
According to OpenAI, textGrain operates during generation by slightly shifting the probability distribution of synonymous word and sub-word tokens. This leaves a mathematical signature across the passage that human readers cannot perceive, but which OpenAI’s specialized detection system can identify. Because the watermark resides in the selection of words rather than metadata or formatting tags, standard copying and pasting between documents or applications does not remove the tracking signal.
However, the technology’s effectiveness depends heavily on text length and domain flexibility. In internal benchmarks set to a 1% false-positive threshold, textGrain achieved an 80% detection rate on 200-token passages and rose to roughly 95% on 400-token samples. Conversely, in technical fields like mathematics or boilerplate coding—where lexical variety is constrained—detection rates fall significantly. Post-generation edits also erode the signal: replacing 10% of tokens drops detection accuracy to approximately 66%, while modifying 25% reduces it to around 17%.
OpenAI emphasized that textGrain is intended strictly as a technical provenance check and cannot identify individual user accounts, reconstruct original prompts, determine human contribution ratios, or verify factual accuracy. To avoid false accusations and the misinterpretation of unedited text, the company is withholding the detection tool from the general public, initially restricting access to accredited academic researchers and specialist organizations on an application basis.