Narrow claims survive review
There is a sentence we keep deleting from our research notes. On null controls, ablations, and why we publish the regime where our own effect reverses.
There is a sentence we keep deleting from our research notes. It goes something like “the hierarchy implements error correction over experience.” It is a good sentence. It would look great in a deck. The trouble is that we can't defend it.
What we can defend is smaller. In our tests, an event-first memory hierarchy behaves like error correction at the representation level: recurring temporal patterns survive corruption of the underlying observations — some events missed, others spurious — better than they have any obvious right to. “Behaves like” and “implements” are separated by a formal model we do not have, and a reviewer who asks for one would be right to.
What a result has to survive
Before anything we measure gets reported, it runs a gauntlet designed by the most hostile reviewer we could recruit, which is us.
Null controls first: the same pipeline pointed at signals where the effect cannot exist, because an effect that shows up anyway means the pipeline is measuring itself. Then ablations — remove the mechanism the claim credits, and see whether the result has the decency to disappear. Then the uncomfortable one, boundary hunting: push conditions until the effect dies. Ours does something better than die. Past a certain entropy in the candidate envelope, it reverses, and the hierarchy starts making things worse.
That reversal ships alongside the headline result. Not buried in an appendix; alongside it, with the regime where the effect holds marked out from the regime where it doesn't.
Why hand critics the failure mode?
Because a result with documented failure modes is more usable than one presented as unconditional. A reader who knows where the effect reverses can tell which regime their own system is in. A reader holding an unconditional claim can only find out the hard way, and when they do, the cost lands on every other claim we have made.
What this discipline mostly costs is headlines. A narrow claim makes a worse press release, and we have watched broader claims travel further on less. We think that trade reverses over time: each claim that survives review makes the next one cheaper to trust, and we are optimising for the tenth interaction with a skeptical reader, not the first.
Our results are not published yet. When they are — the plan is laid out on our research page — the null controls and ablations go out with them, reversal included. In the meantime, the provenance layer the research runs on, ldgr-core, is open source under Apache-2.0. The discipline starts with keeping a trail you can't quietly rewrite.