Radar PereneRadar Perene
← home

Radar Perene / Archive / science

What p-hacking is and why it undermines market models

◦ Index methodology v2.2 (working papers with DOI). See the methodology.

Science

A researcher tests twenty versions of the same idea against the same price series. Nineteen fail. The twentieth passes the statistical test — and is the only one that appears in the final report, presented as if it had been the first and only question asked. No number was falsified. No spreadsheet was tampered with. And the result is still worthless.

P-hacking is the practice of working the data until a statistically significant result appears — testing many variations of a study and reporting only the one that passed, as if it had been the only one tried. The name comes from the p-value, the significance measure the maneuver corrupts.

The term barely exists in Portuguese, and that absence has a practical consequence: a good share of the "models that work" sold or reported in the Brazilian market carries exactly this defect, with no public vocabulary to name it.

The mechanics: why nineteen failures disappear

The most common convention in research calls a finding "significant" when the p-value falls below 0.05 — which means, in plain language, that a result as strong as the one observed would appear by pure chance in roughly one attempt out of twenty, if there were no effect at all.

That sentence contains the key to the problem. If chance approves one attempt in twenty, whoever makes twenty attempts and shows only the approved one has not discovered an effect: they have manufactured one. The statistical test remains mathematically correct — the fraud lives in what was left out of the report. That is why p-hacking is so hard to detect from the outside: the final paper is impeccable; the graveyard of discarded tests is invisible.

The variations are many and not always conscious: shifting the date window until the result shows up, swapping the indicator for a close cousin, excluding inconvenient "outliers", ending data collection on the day the numbers come out right. Each decision looks innocent in isolation. Their sum is a study that only says what the author wanted to hear.

Why the market is the perfect habitat

Financial market research combines the three ingredients that make p-hacking tempting: long series, abundant variables, and a reward for findings.

There are thousands of public series — prices, interest rates, inflation, currency, volume — and each can be sliced into nearly infinite windows, lags and transformations. Whoever hunts for a statistical coincidence in that space will find one, with mathematical certainty. And the incentives all point the same way: a "winning" backtest sells a course, raises a fund, earns a headline. A failed backtest sells nothing — so nobody publishes it.

The aggregate result is an environment where most announced patterns were never truly tested: they were selected. The academic finance literature has discussed the phenomenon for years under names such as data snooping and publication bias; market news coverage, almost never.

Nor is the problem exclusive to course sellers. Academia itself carries it: one of the most cited papers in metascience, published by John Ioannidis in 2005, argues that a large share of published findings across many fields would not survive an honest repetition — precisely because publication incentives reward the significant result, not the true one. In finance, the argument's local version earned its own name: researchers speak of a "factor zoo", hundreds of published return patterns of which a dwindling fraction holds up when tested outside the original sample.

How the house protects itself

The defense against p-hacking is not statistical sophistication — it is process discipline, and it starts before any calculation: the hypothesis is written before the test. What will be measured, on which series, over which period, under which criterion of success — all declared first. After that, the result is whatever it is. The detail of the applied protocol, test by test, stays on the house's bench; the principle is public and fits in one sentence: whoever sets the bar after the jump has not measured the jump.

The most honest proof that the process works is the uncomfortable material it lets through. One of the house's public working papers — the tactical study of the Ânima Index, DOI 10.5281/zenodo.21327608 — ends by concluding that the contrarian edge it set out to test does not hold. The study was published anyway, with the negative conclusion on its cover. A result that contradicts its own premise and survives into the final report is the opposite of p-hacking — and is worth more, as a credential, than ten victorious backtests.

Frequently asked questions

Is p-hacking the same as testing several hypotheses?

No. Testing many hypotheses is legitimate and sometimes necessary — provided every attempt is declared and significance is adjusted for the number of tests. P-hacking begins when the failed attempts are omitted and the survivor is presented as the only question asked.

Is p-hacking fraud?

Not always in the legal sense — many cases are born of self-deception, not bad faith. The effect on the reader, however, is the same as fraud: a conclusion the data does not support.

How does a reader spot p-hacking in a market study?

Typical signs: peculiar date windows with no justification, slightly exotic indicators, no out-of-sample test, and no mention of what was tried and failed. A study that only reports successes deserves the question: where are the failures?

Does a low p-value guarantee the effect exists?

No. It only says the result would be rare under pure chance in a single, honest test. If many silent tests came before it, that rarity is an illusion.

---

P-hacking manufactures patterns that do not exist; the next step on this trail examines the symmetric error — seeing cause where there is only coincidence — in Correlation is not causation.

House reading: the discipline described here underpins today's reading, in the Diário, and the patterns that survived it live among the precedents, in the Atlas.

Auditing a specific study against this kind of defect is bench work — the sort the house takes on by request.

This is the Radar’s memory. Today’s reading — regime, 5 lenses and the day’s analogs — is live, free.

Subscribe to Perene Semanal — US$ 29/mo →

See today’s reading →