Radar PereneRadar Perene
← home

Radar Perene / Archive / science

Sample size in financial series: why \"20 years of data\" can be too little

◦ Index methodology v2.2 (working papers with DOI). See the methodology.

Science

"Twenty years of data" sounds like abundance. Said in a presentation, the number silences objections: two decades, thousands of trading sessions, what more could anyone want? The phrase deceives because it swaps units without warning. Twenty years of daily closes are about five thousand rows in a spreadsheet — and, for many of the questions that matter, half a dozen genuine episodes.

In financial series, the sample size that supports a conclusion is not the number of observations in the file, but the number of independent episodes of the phenomenon under study — and it is usually far smaller than the calendar suggests. A study of crises with twenty years of daily data does not have five thousand crisis observations; it has two or three.

The illusion of frequency

Raising the frequency of the data multiplies rows, not information. Neighboring market days tell almost the same story: Tuesday's close carries nearly everything Monday's carried. Trading monthly data for daily data makes the spreadsheet grow twentyfold and the knowledge grow almost nothing — the new observations are not independent of the old ones, and statistical conclusions feed on independence, not on volume.

The mistake has practical consequences: studies that treat five thousand sessions as five thousand independent pieces of evidence declare a precision they do not possess. The real margin of error, computed over the episodes that genuinely do not repeat each other, is often many times larger than the printed one.

The right unit is the episode

The question defines the unit. Whoever studies an index's behavior in recessions has as many observations as the series has recessions — in the Brazil of the last two decades, a handful. Whoever studies turns in the interest rate cycle counts cycles, not central bank meetings. Whoever studies panics counts panics. Under that ruler, financial series shrink humiliatingly: the file looks like an ocean and contains, for each specific question, a jar.

That is why "twenty years" can be too little and, worse, too little in a specific way: the two available decades may contain only one kind of environment. A Brazilian series starting in 2003 spends nearly its whole length under structurally falling interest rates; what it teaches about rising-rate environments is close to nothing — and the study that does not declare this is generalizing from what it never saw. The neighbors of this trap already have their own articles in this wing: survivorship bias, which shrinks the sample without notice, and look-ahead bias, which contaminates it from the future.

The house's N, declared before the conclusion

This is not a sermon about other people's errors — it is the ruler that limits this house itself, and it is public. Radar Perene's monthly archive, the backbone of the Atlas essays, is long by Brazilian market standards and still small under the right unit. In numbers: it holds 194 monthly readings, from April 2010 to May 2026 — fewer than two hundred points. As episodes, that archive contains a limited collection of distinct environments: the great scares and the great euphorias of the period can be counted on one's fingers.

The editorial consequence is visible in every essay: the collection describes what the archive recorded and resists promising what it cannot support. When a pattern shows up in few episodes, the text says "few" — because with a small N, the boundary between pattern and coincidence is exactly what is in dispute. Sizing what a series can answer, before asking, is a fixed step of the bench; the full protocol of that step does not fit a public article, but the principle fits one sentence: the N is declared before the conclusion, never after.

What to do when the sample is small

The honest answer is not to give up — it is to lower the ambition of the sentence. Small samples support description ("in the archive's four episodes, this happened three times") and do not support general law. They support hypotheses for the future to test and do not support certainties about the next episode. The difference looks subtle and is the whole difference: the text that describes ages well; the text that generalizes from half a dozen cases ages like prophecy.

Frequently asked questions

Is more data always better?

More independent episodes, yes. More rows of the same story, no — higher frequency grows the file without growing knowledge in the same measure.

Doesn't daily data solve the problem?

Not when the question concerns slow phenomena — regimes, cycles, crises. Daily frequency helps with daily-frequency questions; it does not manufacture crisis episodes the calendar did not contain.

Is there a minimum sample size?

Not as a universal number. The minimum depends on the question, the size of the effect sought and the noise of the series. The practical question the house asks comes first: how many independent episodes of the phenomenon does this series contain? If the answer is "three", no technique turns three into thirty.

Why do so many studies advertise decades of data?

Because the big number impresses and is rarely audited. The reader who converts "twenty years of daily data" into "how many episodes of the phenomenon?" disarms most of the rhetorical effect — and is exactly the kind of reader this trail wants to form.

---

Continue the trail: The costs and timelines of publishing a scientific article: what rarely gets explained

House reading: the daily record is in the Diário; the archive's episodes, told one by one, in the Atlas.

Sizing a sample for a concrete question is work the house already does on request — and it yields more in front of the real question than in the general case.

Read also: Look-ahead bias: the error that lets a model \"predict the past\"

This is the Radar’s memory. Today’s reading — regime, 5 lenses and the day’s analogs — is live, free.

Subscribe to Perene Semanal — US$ 29/mo →

See today’s reading →