Radar PereneRadar Perene
← home

Radar Perene / Archive / science

Overfitting in backtests is not about trading, it is about method

◦ Index methodology v2.2 (working papers with DOI). See the methodology.

Science

There is a guaranteed way to build a model that perfectly explains the last twenty years of any market: give it enough parameters. With sufficient freedom, an equation passes through every point on the chart — every crisis, every rally, every hiccup. And there is a second guarantee, less advertised: the more perfect the fit to the past, the worse the behavior on the first day the model has never seen.

Overfitting is when a model learns the noise of the sample instead of the pattern of the phenomenon — it memorizes the answers to the old exam rather than learning the subject. The classic symptom: brilliant performance on the data that trained the model, mediocre performance outside it.

In Brazil, the term circulates almost exclusively as trader jargon — a technical risk for people who program bots, fixed with a platform checkbox. That reading shrinks the concept. Overfitting is a problem of experimental epistemology: it decides whether a result is knowledge or a well-dressed coincidence — and it reaches any quantitative research, from fund strategies to public policy.

The mechanism: too much freedom, too little data

Every finite dataset contains accidents: coincidences that happened once and mean nothing. A sufficiently flexible model cannot tell the accident from the rule — it fits both with equal enthusiasm. Every extra parameter, every additional rule, every accommodated exception is one more degree of freedom for memorizing the past.

The disguise is that overfitting looks like rigor. The model with fifteen rules explains history better than the model with three — on paper, it is "more complete". The difference only shows in the confrontation with new data, which is why the split between training sample and test sample, made before any calculation, is the frontier between research and the curation of coincidences. Whoever tests on the same ground where they fitted has not tested: they have restaged.

In financial series the problem is aggravated by a structural detail: the data is scarce. Twenty years of trading sounds like a lot, but it contains half a dozen distinct economic regimes — and a model that memorized two cycles has not learned "the market"; it has learned two episodes.

The scientific tradition has an old antidote for this excess: parsimony. Between two explanations with the same power, the simpler one is preferred — not for elegance, but for safety: every added complication is one more opportunity to memorize an accident. In quantitative research the operational question is always the same: does this additional rule exist because the phenomenon demands it, or because the training series rewarded it? When the only honest answer is the second, the rule goes.

The house experience: the test that brought down our own findings

The defense against overfitting has a name in the literature and in the house's practice: the robustness test — requiring a result to survive outside the window, under other measures, under other slices, before calling it a result.

One of the house's public working papers shows that ruler applied against ourselves. The study of Brazilian intermarket relationships — public on Zenodo under DOI 10.5281/zenodo.21327663 — started from a set of candidate relations between asset classes that its first version had identified. Put through harder tests, nearly all of them fell; the published version records that only the utilities-linked relation survived the sieve, and documents the correction of the earlier finding in the text itself. The audit protocol worked as designed: what was fine-tuning to the past was identified and discarded before becoming doctrine.

A study that ends smaller than it began is frustrating to write and valuable to read — because what remains has been through honest attempts at destruction. The intermarket concept the house uses today descends from that survivor, not from the original list of candidates.

The ruler for reading any backtest

Faced with an impressive simulated track record, four questions separate method from decoration:

How many parameters does the model have — and is each one justified by an economic reason, or only by the fit? Was the test run on data the model never saw, set aside before the fitting? Does the result survive reasonable variations — another window, another slice, another transaction cost? And how many versions of the strategy were tested before the one in the report?

The fourth question ties this article to the previous one on the trail: an overfitted backtest chosen among many is p-hacking with a chart. The two defects usually travel together.

Frequently asked questions

Are overfitting and p-hacking the same thing?

They are sibling defects on different axes: overfitting is excess fit within one model; p-hacking is selection among many models or tests. A winning course-ware backtest frequently contains both.

So backtests are useless?

No — they are indispensable, as disciplined hypotheses. What invalidates them is undue promotion: treating simulated past performance as evidence about the future without out-of-sample testing and robustness.

How can you tell a model is overfitted without seeing the code?

External signs: many parameters with vague justifications, performance that is "too good" and too regular, no declared out-of-sample test, and results that depend on one specific date window.

Do simple models never overfit?

They overfit less, having less freedom — but simplicity does not immunize: a simple model chosen among hundreds of simple models inherits the defect through the selection route.

---

The studies cited on this trail share one detail: each can be found, in the exact version mentioned, through a permanent code — and that is the next topic: What is a DOI.

House reading: what survived the tests feeds today's reading, in the Diário; the episodes that served as evidence are among the precedents, in the Atlas.

Putting a specific backtest through this battery of questions is bench work — the kind the house carries out on request.

This is the Radar’s memory. Today’s reading — regime, 5 lenses and the day’s analogs — is live, free.

Subscribe to Perene Semanal — US$ 29/mo →

See today’s reading →