Radar Perene / Archive / science
Small sample, big conclusion: the risk of generalizing from scarce data in Brazil
◦ Index methodology v2.2 (working papers with DOI). See the methodology.
Science
There is a statistical irony that haunts anyone researching the Brazilian market: the smaller the market, the larger the conclusions usually published about it. Where data abounds, the researcher is forced into modesty — any strong claim runs into thousands of documented counterexamples. Where data is scarce, the strong claim meets little resistance, and the scarcity that should demand caution ends up financing boldness.
The central risk of researching small markets is not reaching wrong conclusions — it is reaching conclusions larger than the sample supports. Every missing observation should shrink the reach of the final sentence; in practice, sentences of the same size get published over samples ten times smaller, and the reader is rarely told the difference.
This house operates inside that constraint every day, and two of its studies — one on market breadth, another on sector baskets — were born precisely from the friction between imported formulas and the local sample. The numbers of those studies stay on the bench; the methodological lesson the two left behind is the subject here.
The Brazilian scarcity is threefold
The first scarcity is of assets. The universe of Brazilian stocks with meaningful liquidity has grown markedly since 2000, but it remains a small fraction of what the source literature considers normal — dozens of names where the founding papers counted on thousands. Every technique that depends on averages across many assets, on portfolios within portfolios, on populated sectors, operates here with a fraction of the material.
The second scarcity is of time. Reliable, comparable series of the modern Brazilian market begin, for many purposes, after monetary stabilization — a few decades, crossed by structural changes that make it risky to treat the whole period as one thing. Track 1 examined the general case in sample size in financial series; the Brazilian version of the problem is the general case squared.
The third scarcity is the sneakiest: independent observations. In a concentrated market, the few existing assets frequently move together — the same flows, the same indices, at times the same controlling shareholders. A hundred correlated observations are not worth a hundred observations; they are worth something between ten and a hundred, and no one knows exactly how much. The nominal sample fattens the confidence; the effective sample does not keep up.
What the two studies taught the house
The lesson the two works left fits into a rule of conduct: in Brazil, the question "how much data supports this sentence?" must be asked before the research, not after. In the breadth study, the imported formula assumed a universe of assets that Brazil only began to offer — partially — in recent years; applying it to the past meant measuring with a ruler that changed size in the middle of the series. In the baskets study, sector concentration made collective-looking groups behave like half a dozen companies, shrinking the effective sample without altering the nominal one — the mechanism detailed in sector baskets in a concentrated market.
In both cases, the symmetric temptation was present: to publish the big conclusion the small sample seemed to allow. And in both, the discipline that prevailed was the same — sizing the final sentence by what the data can bear, not by what the question deserves. Some conclusions came out smaller than planned. None had to be unsaid later.
How to recognize the oversized conclusion
The reader has no access to anyone's bench, but does have access to the proportion between the declared sample and the published sentence. Some mismatches are visible to the naked eye: studies testing "the Brazilian market" in windows when it had a handful of liquid assets; "sectoral" regularities extracted from sectors with two representatives; "historical" patterns resting on three occurrences. None of these studies is necessarily dishonest — most simply inherited the sentence size from the American literature, without inheriting the sample that justified it.
The heuristic the house applies to others is the one it applies to itself: look, in the text, for the moment when the author declares what the sample does not allow. When that paragraph exists, the rest earns credit. When it does not, the big conclusion over the small sample is usually right above it.
Frequently asked questions
Does a small sample invalidate all research on Brazil?
No — it forces resizing. Narrower questions, conditional claims and declared uncertainty produce legitimate research with little data. What the small sample invalidates is the automatic borrowing of American-sized conclusions.
Why does the house not publish the numbers of the two studies cited?
One of the works is still maturing and the other underpins a methodology in use — in both, what is transferable to the reader is the methodological lesson, published here; the operational detail is bench work, as is house rule for undeposited material.
Will more years of data solve the problem over time?
Partially. The universe has grown and the series lengthens every year, but structural changes keep slicing the past into barely comparable regimes, and concentration keeps reducing the effective sample. The problem changes scale; it does not disappear.
Is there a universal minimum sample size?
No — the minimum depends on what one wants to claim. The transferable rule is one of proportion: the strength of the conclusion has to fit inside the effective sample, and Brazilian samples tend to be smaller than they look.
---
Continue the trail: What the American finance literature assumes that Brazil does not have →
House reading: today's reading is in the Diário; the episodes where the archive says "few similar cases", in the Atlas.
Sizing what a specific sample can bear to claim — before running the test — is the kind of assessment the house does in conversation.
Characters: Method
This is the Radar’s memory. Today’s reading — regime, 5 lenses and the day’s analogs — is live, free.