SPXI-TLP extends provenance design from the publication layer to the training layer.
The protocol assumes that text may pass through scraping, cleaning, deduplication, filtering, tokenization, batching, training, and post-training. Signals outside visible prose may be lost before training.
Its strategic reduction is:
> Assume ingestion. Make extraction carry provenance.
Three engineering registers are combined:
Important stack components include:
The protocol distinguishes survival domains. C2PA and credentials may provide strong human-verifiable evidence but do not themselves survive as training text. Visible capsules and canaries may travel farther but cannot guarantee model recovery.
The record recursively applies the protocol to itself and provides a self-inventory. It proposes a blog corpus as a test substrate and a quarterly canary-recovery audit.
The parametric tooling and chx inscribe command-line implementation are deferred to v2.3.