Radar PereneRadar Perene
← home

Radar Perene / Archive / science

Statistical significance is not economic relevance

◦ Index methodology v2.2 (working papers with DOI). See the methodology.

Science

A powerful enough microscope finds impurities in any glass of water. The discovery is real, measurable, reproducible — and it does not answer the only question that matters to someone who is thirsty: is it drinkable? Statistics applied to markets lives the adult version of that problem. With enough data, almost everything becomes detectable; the issue was never detection, but deciding what, once detected, matters.

Statistical significance says an observed pattern would be unlikely to arise from chance alone. Economic relevance says the pattern, besides existing, is large enough to matter in the real world — after costs, after risk, after everything practice charges. A result can have the first without the second; the confusion between the two sustains much of what is sold as discovery in finance.

The distinction is not a researcher's pedantry. It is the difference between "this effect exists" and "this effect changes any decision" — two claims that everyday language fuses into the word "significant," and that science has kept apart ever since econometrics learned, at its own expense, the price of fusing them.

What significance actually claims

The classic instrument is the p-value: a measure of how surprising the observed pattern would be if there were no effect at all, only chance. When that surprise crosses a conventional threshold, the result earns the title "statistically significant."

Notice what the sentence does not say. It does not say the effect is large. It does not say it is stable. It does not say it survives outside the data where it was found. It says only: this is unlikely to be pure luck. A tiny effect, measured over a long enough history, crosses the threshold with ease — the microscope improved, not the water.

In markets, the problem carries an aggravating factor the literature calls data dredging: whoever tests twenty patterns against the same history should expect, by chance alone, that one of them looks significant. The house covered that mechanism in what is p-hacking; here the consequence suffices — significance, by itself, is a weak filter precisely where the most testing happens.

What relevance additionally charges

The second question is harder because it leaves mathematics and enters economics. A detectable return pattern has to beat three discounts before it matters.

The first is size: an effect of hundredths of a point per month can be real and still smaller than the brokerage fee it would cost to capture it. The second is risk: a detectable average gain can come attached to extreme episodes the record registers only a few times — and the average does not warn you. The third is persistence: published effects have the documented habit of shrinking once they become known; the literature has measured that shrinkage, and it is not small.

The contrast has a celebrated example: the economist Deirdre McCloskey spent decades arguing that econometrics confused "significance" with "importance" — and that the question how big is big? vanished from technically impeccable papers. The financial market is where that question charges interest: whoever acts on a small effect pays real costs for a pattern that may only exist under the microscope.

Where the house met this frontier in practice

This distinction is not theoretical for Radar Perene — it decided the outcome of a study in the house's own series. The Tactical Ânima Index, the house's third working paper, tested whether a tactical reading of the sentiment index sustained a practical short-term advantage. The published conclusion is negative: the text records that the tested advantage does not hold. In numbers: the study is public on Zenodo, with an active DOI and the adverse conclusion in the body itself — not in a hidden erratum.

Publishing a result that overrules one's own hypothesis is the honesty test the frontier between significance and relevance imposes. A pattern can show up in the data; if it does not cross the real world's discounts, the honest record is to say so — and the house preferred to say so with a permanent identifier. The reader need not trust the description: the document is verifiable.

What changes for the reader of market research

The distinction yields a practical reading question. Facing any study — academic or commercial — announcing a "significant" effect, the attentive reader looks for the effect's size in units that make economic sense: points of return, comparable cost, frequency of episodes. When a text displays the p-value and hides the magnitude, the omission is usually the information.

And the ruler works both ways. A large, economically interesting effect that fails to reach significance also deserves suspicion — it may be chance dressed as a thesis. The two questions ("does it exist?" and "does it matter?") only work together; each one alone has already financed plenty of hasty conclusions.

Frequently asked questions

Can a result be economically relevant without being statistically significant?

Yes — and it is the most treacherous case. A large effect measured over few episodes cannot tell talent from luck. Mature practice treats such a result as a hypothesis to monitor, not a finding.

Which p-value is "enough" in finance research?

There is no universal number. Conventional thresholds exist, and part of the recent literature argues that, in a field where thousands of patterns are tested against the same prices, the traditional threshold is far too loose. The workable consensus: the threshold is the beginning of the conversation, never the end.

Is a "small effect" always irrelevant?

No. Scale changes the arithmetic: a tiny effect can matter to someone operating at institutional volume and be invisible after costs to everyone else. Relevance is always relative to someone — one more reason not to delegate it to the p-value, which is relative to no one.

How does the house apply this distinction day to day?

By separating the verbs: what the house's indices detect is described as historical record, never as instruction. When a detectable pattern failed to sustain practical advantage, the corresponding study says so openly.

This discipline of separating "it exists" from "it matters" flows into the house's most visible rule — never turning record into instruction: why we never write "you should"

House readings: today's note, in the Daily · the episodes that test the ruler, in the Atlas.

Running a specific effect through the three discounts — size, risk, persistence — is an exercise the house conducts on request.

This is the Radar’s memory. Today’s reading — regime, 5 lenses and the day’s analogs — is live, free.

Subscribe to Perene Semanal — US$ 29/mo →

See today’s reading →