What Is textGrain?
textGrain is OpenAI's watermark for ChatGPT and Codex text in the EU. It lives in word choice, not hidden characters, so stripping invisible characters does not remove it.
textGrain is OpenAI's watermark for text produced by its models. It does not add anything visible or invisible to your writing. Instead it changes which words the model picks: where several words fit a sentence about equally well, a secret key nudges the choice toward one of them. A detector holding the same key can read that pattern back out of a long enough passage. OpenAI announced it on 5 October 2026 as the way it will meet EU transparency rules for AI-generated text.
Most people find it typed as one word, "textgrain", and want two answers: does my text carry it, and can it be taken out. Both answers turn on a detail that most explainers skip. textGrain is not a hidden character.
Why textGrain exists
The EU's digital services rules require that text generated by AI systems be discoverable as such. OpenAI's answer is to put the signal inside the text rather than rely on a label beside it, the same instinct behind watermarking AI images and synthetic media generally. How the company described the rollout:
- Who gets it: eligible ChatGPT and Codex users across all plans, in the EU.
- When: announced 5 October 2026, added gradually over the following weeks.
- Where it does not apply: outside the EU at launch.
- API access: customers anywhere can opt in for select models. It is off by default.
That scope is worth remembering. Text written in ChatGPT outside the EU, or through the API with default settings, does not carry a textGrain mark.
How textGrain works
A language model predicts one token at a time, and for many of them more than one continuation is reasonable. "The report was strong" and "the report was solid" fit the same sentence. Where the choice is genuinely open, the watermark's key biases which option gets taken. Where a word is forced, such as a name, a figure, or a term of art, the model is left alone. That is why the mark costs almost nothing in readability: it only spends its influence where the model was already unsure.
Because the mark lives in selection rather than insertion, nothing is appended. There is no zero-width space, no direction mark, no stray punctuation waiting to be found. The method is described in OpenAI's paper, textGrain, entropy-calibrated watermarking for language model text.
textGrain versus the invisible characters it gets confused with
| textGrain | Copy-paste leftovers | |
|---|---|---|
| What it is | A statistical pattern in which words the model chose | Actual characters: zero-width spaces, direction marks, hidden tag characters, soft hyphens |
| How it arrives | Word-by-word choices at generation time, guided by a secret key | Rides along when text is copied out of a chat app or a rich-text editor |
| Can you see it | No. Only a detector holding the key can read it | No, but any tool that inspects characters finds them at once |
| Can you confirm removal | No, not outside OpenAI | Yes. The characters are gone, and you can check that in your browser |
This distinction is why "strip invisible characters" cleaners do not touch textGrain. They remove something else entirely, which is still worth doing: those leftovers break pastes into other tools, show up as odd spacing, and survive into AI citations of your page.
How textGrain is detected
Reading the mark requires OpenAI's detector, and that detector is not public. OpenAI is taking applications for access, initially limited to approved researchers and expert organizations. Three consequences follow:
- A site that claims to detect textGrain for the general public is not using OpenAI's detector.
- A site that claims to confirm the watermark has been removed is making the same unsupported claim.
- Detection is statistical, so it needs a passage long enough to hold a signal. Short snippets are unreliable in both directions.
What weakens the signal
OpenAI's own testing on 400-token passages, measured at a 1% false-positive target, found detection of about 92% in unedited text. Replacing roughly 10% of the words with synonyms brought that to about 66%. At roughly 25% replaced it fell to about 17%. Shorter text is harder to read either way: about 80% on 200-token passages. The company also states that substantial paraphrasing or translation can make the watermark undetectable.
So removal is a spectrum, not a switch. Light editing leaves the pattern intact. A genuine rewrite of most sentences does not. Our free TextGrain remover works along that spectrum: a separate AI model rebuilds the wording sentence by sentence while holding every fact, number, name, link and quote in place, and the result tells you how much of the wording changed and which details survived. What it cannot do is promise the mark is gone. Nobody outside OpenAI can check that.
How to handle it in a publishing workflow
- Know where the text came from. Note whether a draft was written in ChatGPT or Codex, and from where, since the EU rollout does not describe every account.
- Clean the visible problem first. Run the draft through the Clean only mode. It deletes hidden characters and normalises odd spaces in your browser, changes no words, and gives you a list of what was removed.
- Rewrite where you want your own phrasing. The rewrite pass rewords the text and reports the change rate, so you can see how far you have moved from the model's wording.
- Keep the disclosure decision separate from the text. If your school, employer, publisher or local rules require you to say AI was involved, rewording does not lift that duty.
- Move on to the part you can control. Whether an answer engine can use and credit your page depends on how it is published, not on a watermark. See robots.txt and llms.txt for the signals you can set yourself.
What textGrain does not mean
- It is not a quality signal. It records where text came from, not whether it is accurate. A watermarked passage can be wrong, and an unwatermarked one can be wrong too; models produce AI hallucinations with or without a mark.
- It is not attribution. Provenance tells you a system generated the text. Source attribution tells readers where the claims came from, and only you can add that.
- It is not a machine-readability setting. A watermark says nothing about whether AI systems can crawl, retrieve or cite your site.
- It is not a public record. Nothing published says how other platforms read, ignore or act on the mark.
Frequently asked questions
Is textGrain a hidden character?
No. It is a pattern in word choice. Invisible characters can ride along when you copy text out of a chat app, but they are a separate problem, and deleting them has no effect on textGrain.
Does ChatGPT watermark everything I write?
Only in the EU, for eligible ChatGPT and Codex users, following the announcement on 5 October 2026. Outside the EU the mark was not applied at launch, and API use is opt-in and off by default.
Can I check whether my own text has it?
Not with OpenAI's detector, which is not public. You can check the other thing yourself: the free tool lists any hidden characters in your paste before you decide whether to rewrite. It runs in your browser and sends nothing anywhere for the clean pass.
Can textGrain be removed?
Rewording weakens it, and OpenAI reports that heavy paraphrasing or translation can make it undetectable. No tool can confirm removal, because no public detector holds the key.
Does textGrain affect search rankings?
OpenAI describes it as a provenance signal for disclosure under EU rules. Nothing published says search engines read it, and no ranking factor has been documented for it.
Related Terms
llms.txt
llms.txt is a proposed plain-text file at the root of a site. It gives large language models a curated, machine-readable map of the pages that matter most.
What Are AI Citations?
AI citations are the source links AI models attach to their answers. They show where a fact came from and decide which brands get credit.
What Is robots.txt?
robots.txt is a plain-text file at the root of a site that tells crawlers which paths they may fetch.
What Is an AI Hallucination?
An AI hallucination is when a model states something false with full confidence. It happens when the model fills gaps with plausible-sounding text instead of grounded facts.
Token
A token is the smallest piece of text an AI model reads at a time. Sometimes a word, often a fragment of one.
Synthetic Media
Synthetic media is image, video, audio, or text generated by AI rather than captured or written by a person.
What Is Source Attribution?
Source attribution is the practice of an AI system naming and linking the sources it used to generate an answer.