Radar Perene / Archive / science
Robustness testing: why a result that only works one way is not a result
◦ Index methodology v2.2 (working papers with DOI). See the methodology.
Science
There is a question that dismantles most exciting market findings, and it fits on one line: does the result survive if something changes? Shift the window by a month, swap the index for a neighboring one, start the series two years later — and watch. A surprising share of "discovered patterns" evaporates at the first alteration. They were never patterns; they were properties of the exact configuration in which they were born.
A robustness test is the procedure of redoing a result under deliberately altered conditions — another window, another subsample, another definition of the variables, another calculation method — to check whether the finding survives outside the configuration in which it was found. What exists only one way is not a fact about the world; it is a fact about that way.
Why fragile results are born all the time
No one needs bad faith to produce a fragile finding. The normal research process suffices: reasonable choices made one at a time — this period because the data is better, this definition because it is the most common, this window because it is the one the literature uses — which, added up, shape the result without the author noticing. It is the silent relative of p-hacking: not the active hunt for the desired result, but the passive settling into it.
The robustness test is the structural antidote. It asks, in an organized way, what happens when the reasonable choices are traded for other equally reasonable ones. If the finding depends on one specific choice, that does not necessarily kill it — but it demotes the claim: the text stops saying "A relates to B" and starts saying "A relates to B in this window, under this definition". The difference between the two sentences is the difference between science and a dated anecdote.
The repertoire, in broad strokes
The classic repertoire has four families, and describing them gives away no method — the craft lies in the dosage and the order, not in the list. The first is temporal variation: cutting the series into subperiods and checking whether the finding lives in each one, or only in the average that blends them. The second is definitional variation: measuring the same phenomenon with an alternative ruler. The third is method variation: if the conclusion changes when the calculation changes, the conclusion belonged to the calculation. The fourth is the placebo test: applying the same procedure where nothing should appear — if something does, the procedure manufactures results, and the original finding loses its credit.
A result that crosses all four families is not proven — it is harder to refute, which is the most an empirical study can aspire to.
What the ruler did to our own folklore
The proof that this house applies the ruler is not a promise — it is a public, uncomfortable outcome. The working paper Brazilian Intramarket Relationships started from a set of relationships between sectors and asset classes of the Brazilian market, pairs the folklore treats as reliable gears and which a first version of the study had identified as candidates. Retested under altered conditions, nearly all of them fell. In numbers: of the relationships examined, a single one survived the retest — the one linking the utilities sector to the market's defensive moves — and the second version of the study openly revises a finding the first one upheld, with both versions preserved in the repository.
The detail that matters: the study published the mortality, not just the survivor. A text presenting only the winning relationship would tell half the story — and the omitted half is precisely the one that says how much the winner is worth. Which battery of variations was applied, in what order and with what cut-off criteria, is bench work; the outcome is public and verifiable.
What robustness is not
It is not repeating the test until the result pleases — that is the vice under another name. It is not a guarantee of truth: a finding can survive every variation and still fall when new data arrives, as discussed in Reproducibility in finance. And it is not a table ritual at the end of the paper: decorative robustness, done for the record, is the theater of method without the method.
Frequently asked questions
Are robustness and reproducibility the same thing?
No. Robustness asks whether the result survives variations of the study itself; reproducibility asks whether someone else, with the same data and methods, reaches the same number. A study can be reproducible and fragile — or the reverse.
How many robustness tests are enough?
There is no number. The better question is qualitative: do the tested variations cover the choices that most influence the result? Ten irrelevant variations are worth less than two at the right joints.
Is a result that fails one test dead?
Not necessarily — it is resized. The failure marks out where the finding holds, and that declared boundary is more useful than an untested generality.
Why do so many market studies publish no robustness tests?
Because the incentive points toward the finding, not toward its demolition. It is one of the reasons showy market results age badly — and one of the most useful filters a reader can apply.
---
Continue the trail: Sample size in financial series: why "20 years of data" can be too little →
House reading: today's reading is in the Diário; the patterns that survived retesting, in the Atlas.
Going deep on a robustness test for a specific relationship is the kind of exercise that fits a consultation — the house bench makes a routine of it.
Characters: Method
This is the Radar’s memory. Today’s reading — regime, 5 lenses and the day’s analogs — is live, free.