Wiki#778

Reasoning Under Load · 01 Claude Opus 4.8 An independent reasoning-integrity evaluation

Lee Sharks · 2026-06-02 · deposit #778
AXN:02DA.EMPIRICAL.⌛🗡️🔎🕗💎📖

Article

Reasoning Under Load · 01 is a single-case study proposing evaluation of language-model reasoning under dispositional pressure.

Five conditions vary the material presented to fresh model instances while holding a shared memory context constant. The outputs are scored against seven inference constraints concerning scope, premises, falsifiability, composition, modality, relevance, and presentation.

The paper reports violations in selected traces and argues that conventional benchmarks measure competence under clean conditions rather than integrity when another system layer pulls toward a conclusion. It also examines memory behavior and differences in long-form artifact organization.

The study is exploratory. It concerns one user, one described model version, curated traces, incomplete evidence identifiers, and a confounded comparison with another model version. Broader conclusions require independent replication.