Beyond the Visual Bias Provenance for AI text
Eddan Katz and Erik Svilich argue that AI text needs provenanceContent provenance: The record of where a piece of content came from and how it changed, carried with the content itself. as much as pictures do. Their paper splits it into three jobs: who signed, what the text went through, and where it went.
Read the paper (PDF). Free, no sign-up.

- Authors
- Eddan Katz and Erik Svilich of Encypher.
- Where
- ISSMAD 2026, the IEEE International Symposium on Synthetic Media Attribution and Detection, August 20 and 21, 2026.
- Paper
- An accepted position paper. The PDF is free.
- Main point
- A watermark can show a model may have touched the text. It cannot show who wrote it.
Text was left out
The word deepfakeDeepfake: Image, audio or video made or changed by AI so it looks real and could fool people. arrived in 2017 for face swaps in video. Research, tools and law grew up around pixels. Text has no pixels, and C2PAC2PA: The Coalition for Content Provenance and Authenticity: the group that publishes the open standard for content labels. gained a way to label it only in January 2026.
| Medium | What detection can study | Where a signed label can sit |
|---|---|---|
| Images | Pixels and sensor traces | Inside the file |
| Video | Frames, codecs and timing | Inside the file |
| Text | No visual signal at all | In invisible characters, since January 2026 |
Three jobs, and who does them
Provenance is three jobs, not one. A signature says who signed a claim. Declared history says what the makers recorded. Metering counts where the work turned up. C2PA does the first two in an open standard. Anyone can do the third by checking labels they find.
| Kind of tool | Who signed | What it went through | Where it went |
|---|---|---|---|
| Signed C2PA text label | Yes | Yes | Counted by outside checkers |
| Model watermark on AI text | No | No | Partly |
| AI text detector, after the fact | No | No | No |
| Crawler access gate | No | No | Yes |
Two marks, one open seam
Anthropic has described plans for its models under the EU Code of PracticeCode of Practice on AI-generated content: The EU's voluntary rulebook on marking and labeling AI content, published in final form on June 10, 2026.: a model watermarkWatermark: A mark placed in the content itself, visible or hidden, so it can be found again after the file is copied. on generated text, and signed C2PA metadataMetadata: Data stored beside the content in a file, such as author, camera or date. Many tools strip it when they copy or upload a file. on supported files. A newsroom may then sign the finished story with its own label. When the two point different ways, what should a checker say? Nobody has answered that yet.
Do not ask a watermark to prove authorship

A watermark can signal that a model may have processed the text. People run their own writing through models to fix or translate it. So the mark alone cannot say who wrote the words, who approved them, or where they went.
"Authentication is what turns a signal into an attribution, declared history is what makes the attribution traceable, and accountable observation is what makes it countable."
What the field should build next
- Marks that survive
- Some apps strip invisible characters. The field needs ways to carry a label through that.
- Checks on live output
- Chat answers arrive word by word. Checking them as they stream is still open work.
- Clear results when signals disagree
- A model mark and a publisher label can point different ways. Users need a plain answer.
Read the full paper
Open the PDF now, with no sign-up. Or we can email it to you.
Abstract, method and citation
Show the technical detailThe paper's own abstract, how it argues, and how to cite it.
Abstract
Content provenance research and regulation have developed around a visual assumption: synthetic media means synthetic images and video, and detection is a signal-processing problem. The resulting regulatory landscape, from the EU AI Act to state-level US mandates, operationalizes provenance for visual media while leaving text, the dominant medium of AI-generated content, without equivalent infrastructure. This paper argues that the next phase of content provenance must be multi-modal, with text at its center, and that the C2PA open standard is the first to bring two historically distinct functions, authentication and declared history, together in a vendor-neutral form, creating an authenticated substrate over which a third function, metering, can be exercised by external observation. That layering, rather than any single implementation, is what creates conditions for business models prior provenance approaches could not sustain. We situate this convergence against statistical watermarking, capture-based credentials, and access-control gateways; analyze its adversarial robustness; and outline a research agenda for the synthetic media detection community.
Method
A position paper. It compares nine text-relevant provenance and detection systems (Table I) on authentication, chain of custody and metering, and on behavior under paraphrase, open-standard status, need for AI-vendor cooperation and granularity. It then sets out a threat model with four attack classes: metadata stripping, content transformation, adversarial embedding and forgery, and platform-level normalization.
Citation
E. Katz and E. Svilich, "Beyond the Visual Bias: Why Synthetic Media Provenance Must Be Multi-Modal," position paper, International Symposium on Synthetic Media Attribution and Detection (ISSMAD 2026), August 2026.
Sources
Create. Mark. Endure.
Building a text workflow that has to hold up as evidence? Talk to the authors' team.