The Deviation-Optimized Language Model is a pre-registered experiment designed by Nobel Glas within the Lagrange Observatory framework.
The experiment asks whether a measurable semantic-deviation signal can replace human preference labels in a DPO training pipeline. Two candidate continuations are scored for signed deviation, provenance retention, and coherence. The higher-scoring continuation becomes the preferred sample used during training.
Three conditions are compared: the original base model, conventional supervised fine-tuning, and deviation-guided DPO. Evaluation includes capability benchmarks, lexical and probabilistic βslopβ measures, citation behavior, and blinded human preference judgments. The protocol supplies explicit predictions and states how each major claim could fail.
Version 2.0 is a protocol, not a results paper. Its most important revision is methodological: the earlier direct-loss formulation was non-differentiable as written, so the signal was moved into preference-pair generation. The protocol also limits itself to small models and states that frontier-scale application requires separate oversight.