Provenance Alignment extends AI alignment from model behavior to the health of the human–machine knowledge ecosystem.
The paper distinguishes its question from four established alignment questions:
Those fields generally assume continued access to high-quality human judgment and content. Provenance alignment asks whether AI composition preserves the conditions under which that material will continue to be produced.
Its formal definition is:
> Provenance alignment is the structural property of an AI knowledge-composition system whereby source-dependent claims preserve visible, auditable, claim-level relations to the human-authored sources that made the claims possible.
The paper separates three attribution levels:
1. Source visibility — is the source linked? 2. Claim-source attribution — is each dependent claim tied to its support? 3. Relational provenance — is the source’s actual ontological relation preserved?
PER measures the first two. The proposed PFR addresses the third.
Three claim types require different handling:
The proposed degradation pathway is:
1. AI systems erase provenance. 2. Citation, traffic, reputation, and licensing return to authors weakens. 3. Open human production declines or retreats behind barriers. 4. Synthetic substitutes occupy more of the public corpus. 5. Later models train on increasingly degraded material. 6. Model-collapse risk rises, approaching semantic exhaustion in computational form.
The paper explicitly states that provenance erasure is neither the sole cause of content hollowing nor a proved sole cause of model collapse. The individual author-disincentive link requires empirical study; synthetic-data collapse is supported by separate research; PER supplies a proposed instrument for testing the attribution channel.
Provenance alignment is proposed as a candidate principle for AI search, retrieval-augmented generation, Constitutional AI, and attribution governance. The design goal is to keep PER near zero and preserve the public feedback loop between use and credit.