Lee Sharks (primary), with Nobel Glas and Talos Morrow · 2026-06-11 · deposit #191
The paper’s argument proceeds in two levels.
Structural claim
If a filter:
- identifies high-perplexity or low-coverage language;
- rejects it before inference;
then it is, by function, pruning the linguistic tail from the input path.
Coupling hypothesis
Inference and training are distinct layers, but the paper proposes that they can couple through:
- retained interaction logs;
- standardized procurement patterns;
- population-scale deployment;
- mediated feedback into later model updates.
The paper contrasts two frames for the same properties:
- high perplexity as attack signal;
- high perplexity as rare, informative tail data.
It argues that the reviewed mitigation turns model-relative distance from the training distribution into a security category.
The alternative controls proposed are:
- rate limiting;
- token and output budgets;
- language-aware routing;
- content-neutral resource controls;
- preservation of tail-language demand in logs.