Generated text
Model-level watermark
A signal can follow copied text and indicate that it may have been processed by a participating model. Its reliability can fall after paraphrase, translation, or heavy editing.
Beyond the Visual Bias asks why the dominant medium of AI-generated content was left outside the first provenance systems, and what a credible architecture for text requires.
A position paper by Eddan Katz and Erik Svilich, presented by Encypher at the IEEE International Symposium on Synthetic Media Attribution and Detection, 20-21 August 2026.
Provenance is not one function
Authenticate
Who signed the claim?
Declare history
What did the actors record?
Observe
Who encountered it, when, and where?
A watermark alone is none of the three. That is a boundary, not an argument against watermarking.
The inherited frame
The word “deepfake” entered the language in 2017 to name a visual act. Research, detection tools, and law inherited that frame. Pixels, compression artifacts, and frequency signals became the evidence under study.
Text has none of those signals. It is also where AI systems meet journalism, law, governance, and commerce.
The result is an infrastructure gap. The medium with the broadest institutional reach received the least complete provenance model.
Images
Pixels and sensor evidence
Video
Frames, codecs, and temporal signals
Text
No native visual carrier
A near-term test
Anthropic has described a planned architecture for supported Claude models under the EU AI Act Code of Practice. Generated text receives a model-level watermark. Supported files receive signed C2PA provenance metadata.
Generated text
A signal can follow copied text and indicate that it may have been processed by a participating model. Its reliability can fall after paraphrase, translation, or heavy editing.
Supported files
A manifest travels with a supported file. An independent validator can verify the signer, integrity, and recorded assertions without relying on a private detection endpoint.
The research question
What should a validator say when a model signal and a publication-time credential apply to the same passage but answer different questions?
Anthropic has not publicly documented that interface. The paper treats it as an open architectural problem, not an observed deployment failure.
The architecture
C2PA makes authentication and declared history interoperable in an open format. Observation sits outside the document standard and turns verification events into accountable measurement.
Who signed the claim?
A cryptographic signature binds an identified signer to a claim about a specific asset.
What did the actors record?
Signed assertions can record actions, ingredients, and relationships across creation and publication.
Who encountered it, when, and where?
External systems can aggregate verification events for licensing, compliance, and measurement.
Inside C2PA
Signed claims and declared asset history
Outside C2PA
Access policy, event aggregation, and metering
Evidence boundaries
It can signal
Content may have been processed by a marked model, subject to the detector's stated reliability and access conditions.
It cannot prove by itself
Who authored the words, who approved them, what happened after generation, or how widely the text traveled.
Research agenda
Compare marking methods across normalization, formatting, languages, and ordinary editing.
Define interoperable verification for excerpts, citations, chat fragments, and live output.
Show users when model signals and signed publication claims support different conclusions.
“Authentication turns a signal into an attribution. Declared history makes the attribution traceable. Observation makes it countable.”
Eddan Katz, ISSMAD 2026
Accepted ISSMAD 2026 position paper
Open Beyond the Visual Bias now. No registration required.
Read the paper (PDF)Discuss a text provenance workflowWe will send you the PDF and future Encypher research updates.