Radar Perene / Archive / science
Reproducibility in action: what to do when a central statistic does not repeat
◦ Index methodology v2.2 (working papers with DOI). See the methodology.
Science
There is a moment in the life of anyone who works with data that no manual romanticizes: the number holding up the analysis is recalculated under a legitimate variation of the same test — and comes back different. Not scandalously different; different enough that the conclusion no longer stands on its own. What happens in the following minutes defines the seriousness of a research operation better than any statement of principles.
The house's discipline for that moment is a stopping rule: when a central statistic does not reproduce under legitimate variants of its own test, the line of analysis that depended on it is recorded as unconfirmed and does not move forward — not toward publication, not into the house's readings. The record stays; the line stops. Reproducibility is not a decorative ideal: it is a gate with a closing rule.
Why the moment is more dangerous than it looks
The danger is not in the divergent number — it is in the repertoire of rationalizations available to whoever finds it. The variant that contradicted the result can be reclassified as "less appropriate"; the original cut can be retroactively promoted to "the correct one"; the test can be run again, with adjustments, until it returns the expected value. Each of these moves has a technically defensible justification in isolation. Their sum has a name, and the reader of this trail already knows it: the pursuit of the result, disguised as refinement.
That is why the stopping rule must exist before the episode. Decided in the heat of a threatened finding, the decision leans toward saving the finding — always. Decided beforehand, as cold protocol, it does not ask what the house would like to be true; it asks only whether the number came back or did not.
How the discipline works, without the manual
The mechanism is the natural extension of the robustness test into routine: central results are re-executed under reasonable variants of the same calculation, and survival is a condition of passage, not a formality. When the statistic comes back, the line proceeds. When it does not, the outcome is binary by design — the non-confirmation is recorded, dated, and the line is closed or demoted to a hypothesis under observation. What does not exist is the third way: moving forward with a number that only works when calculated one particular way.
The house has exercised this rule — internal lines of analysis have been closed exactly this way, for non-confirmation on retest, and the record of those stops belongs to the internal quality-control history. Which lines, under how many variants, on what dates: that is bench work, and detailing it publicly would serve curiosity more than verification. What matters to the reader is the asymmetry of the commitment: the cost of a stop is private and immediate; the cost of not stopping would be public and deferred — it would land on whoever reads.
The protocol working is not bad news
There is a mistaken reading of this kind of account, and it is worth disarming: that admitting to statistics that do not repeat weakens the credibility of whoever admits it. The mature reading is the reverse. Every quantitative operation in the world routinely produces numbers that do not survive the retest — the difference between operations is not in the existence of those numbers, it is in what happens to them. Where there is no stopping rule, they become publications; where there is one, they become internal records and closed lines. The house that never closes any line is not the one that never errs — it is the one that does not recalculate, or recalculates and does not tell. The subject has its general treatment in reproducibility in finance; here, the point is the habit.
The rule has a corollary that goes unnoticed: it also protects the results that survive. In an archive where lines die of non-confirmation, the numbers that remain carry an implicit attestation — they stood under the same threat and came back. Without the gate, surviving would mean nothing, because nothing could die.
Frequently asked questions
Does a statistic that fails to repeat prove the analysis was wrong?
Not necessarily — it proves it was not established. Non-reproduction demotes the finding from result to hypothesis; the line can be resumed if new evidence supports it. What the rule prevents is treating it as a result until that happens.
How does this rule differ from an ordinary robustness test?
The robustness test is the instrument; the stopping rule is the commitment about the outcome. Many operations run the test and negotiate with the result. The discipline described here removes the negotiation: did not come back, does not proceed.
Why not publish every closed line, in detail?
Because the public value is in the standard, not the inventory. Internal lines closed by quality control are the protocol's normal functioning; what the house publishes are the materials that crossed the gate — and when an already-published finding falls, the revision is made in the open, in the repository itself, as the house's series has already recorded.
Doesn't this slow production down?
It does — and that is the correct price. An operation that publishes at the speed of its raw findings is publishing noise punctually.
---
Continue the trail: Who pays the bill if a market index errs: the institutional case for public self-assessment →
House reading: today's reading is in the Diário; the patterns that survived the gate, in the Atlas.
Running a reader's specific hypothesis through this same reproduction sieve is bench work — the house makes a routine of it, in conversation.
This is the Radar’s memory. Today’s reading — regime, 5 lenses and the day’s analogs — is live, free.