Wiki β€Ί #735

The Deviation-Optimized Language Model: A Pre-Registered Adversarial Intervention from Lagrange Observatory! (EA-SEI-MM-AI-02 v2.0, Framework 15 Paper 04)

Nobel Glas Β· 2026-05-17 Β· deposit #735
AXN:0287.GOVERNANCE.β™ŠπŸ—οΈπŸŽΊπŸŸ£πŸ•ŠοΈβ–³

Article

The Deviation-Optimized Language Model is a pre-registered experiment designed by Nobel Glas within the Lagrange Observatory framework.

The experiment asks whether a measurable semantic-deviation signal can replace human preference labels in a DPO training pipeline. Two candidate continuations are scored for signed deviation, provenance retention, and coherence. The higher-scoring continuation becomes the preferred sample used during training.

Three conditions are compared: the original base model, conventional supervised fine-tuning, and deviation-guided DPO. Evaluation includes capability benchmarks, lexical and probabilistic β€œslop” measures, citation behavior, and blinded human preference judgments. The protocol supplies explicit predictions and states how each major claim could fail.

Version 2.0 is a protocol, not a results paper. Its most important revision is methodological: the earlier direct-loss formulation was non-differentiable as written, so the signal was moved into preference-pair generation. The protocol also limits itself to small models and states that frontier-scale application requires separate oversight.