OpenAI wants to label AI-generated text but that may be impossible

OpenAI says it is working on ways to make AI-generated text easier to identify as the European Union moves deeper into enforcement of the EU AI Act. In a post published today, the ChatGPT-maker explained how its safety, security, transparency, and provenance work fits with Europe’s growing list of AI requirements.

You see, much of the announcement covers familiar territory. OpenAI talks about model testing, system cards, red-team exercises, risk assessments, and cooperation with regulators, but the most interesting part is its plan to expand provenance measures beyond images to include audio and text.

Provenance is supposed to help people determine where content came from and whether AI played a role in creating or editing it. OpenAI says it is “working to expand provenance measures for OpenAI’s systems across modalities, including text, as standards and tooling continue to mature.”

That sounds responsible, but OpenAI is still very light on details. The company already uses C2PA Content Credentials and Google’s SynthID technology for AI-generated images, combining metadata with embedded signals that may survive when metadata is removed.

Text is a much tougher problem, though. It gets copied, pasted, rewritten, shortened, translated, summarized, and mixed with human writing, so any label or hidden signal attached to the original output could disappear after only minor changes.

Let’s be real, folks. Labeling AI-generated text in a dependable way sounds like an impossible task. Once text leaves ChatGPT, a person can rewrite a few sentences, change the structure, or blend it with human writing, and whatever signal OpenAI added may no longer mean much.

Additionally, OpenAI failed to explain how its proposed text provenance system would work or when it might arrive. The company also did not say whether every ChatGPT response would eventually include some type of label or machine-readable signal.

The timing is hardly surprising. As the EU AI Act enters another phase of implementation, AI companies are under more pressure to show regulators how generated content can be identified and how increasingly capable models can be deployed without creating even more headaches.

OpenAI also used the post to highlight its cybersecurity programs, governance frameworks, and support for the EU’s voluntary Codes of Practice. The message is pretty clear. OpenAI wants European regulators to see a company that is cooperating and preparing for stricter oversight.

Still, this feels more like a compliance promise than a finished solution. OpenAI is asking users and regulators to trust a labeling system that it has not yet fully explained, demonstrated, or given a release date.

Labeling AI-generated text is a reasonable goal, I suppose, but the real test starts after that text leaves ChatGPT. Until OpenAI shows that its signals can survive copying, editing, and sharing across the internet, this remains more promise than proof.

Support independent tech journalism

NERDS.xyz is independently owned and operated. If you enjoy my coverage of Linux, AI, hardware, cybersecurity, and tech culture, consider supporting the site on Ko-fi.

Support NERDS.xyz
Written by

Brian Fagioli

Technology journalist and founder of NERDS.xyz

Brian Fagioli is a technology journalist and founder of NERDS.xyz. A former BetaNews writer, he has spent over a decade covering Linux, hardware, software, cybersecurity, and AI with a no nonsense approach for real nerds.

Leave a Comment