Radar PereneRadar Perene
← home

Radar Perene / Archive / science

Reproducibility in finance: why almost no one tests it

◦ Index methodology v2.2 (working papers with DOI). See the methodology.

Science

There is a simple test that separates research from well-formatted opinion: hand the study to a stranger and ask them to arrive at the same number. No phone call to the author, no secret file, no "trust me." In finance, that test almost never happens — and when it does, the outcome tends to be uncomfortable.

Reproducibility is a third party's ability to obtain the same result as a study using the same data and the same method as the original work. It differs from replicability, which asks a different question: whether the finding holds up when tested on new data. A study can be reproducible and still be wrong; but a study that is not reproducible cannot even begin to be evaluated.

What does reproducing a study actually involve?

Reproducing is not rereading. It is redoing. The stranger in the test needs three things: the exact data the author used, a complete description of what was done with it, and a way to verify that nothing changed along the way. If any of the three is missing, reproduction turns into archaeology — the reader guessing which version of the series was downloaded, which window was trimmed, which observation was silently dropped.

In economics and finance, this detail is less innocent than it looks. Series get revised by the agencies that publish them, commercial vendors correct their histories without notice, and the same acronym can name two different series in two sources. Anyone who has tried to rebuild a table from a published paper knows: the distance between "the CPI data" and that CPI file, downloaded on that day, is where most reproductions die.

Why does finance test so little?

Three reasons pile up, and none of them is a scandal — they are incentives.

The first is proprietary data. Much of market research runs on licensed databases the author cannot redistribute. The study describes the method, but the ordinary reader has no way to buy the same access and check. The second is code: for decades, journals accepted papers without requiring the programs that generated the tables, and a method described in prose is never as precise as the script that executed it. The third is the most human: reproducing someone else's work consumes weeks and yields little prestige. Academic careers reward the new finding, not the audit of the old one.

The cost of that combination surfaced when someone decided to check. A large-scale replication effort published in 2020, which redid hundreds of "anomalies" documented in the asset-pricing literature, found that most of them did not survive stricter statistical criteria and extended samples. Not because the original authors were dishonest — but because no one had redone the math before.

How a small house practices this

The usual objection is that reproducibility is a luxury for large universities, with staff and funding. This house's experience suggests the opposite: the smaller the structure, the cheaper the discipline — provided it enters the design from the start, not as a patch at the end.

Radar Perene's working papers are published on Zenodo, the open repository operated by CERN, each with its own DOI. Alongside the first two studies in the series sits what matters most for this article: the complete data package behind them. In numbers: 16 public CSV files, each accompanied by its SHA-256 checksum — the cryptographic fingerprint that lets any reader verify, byte by byte, that the file downloaded today is the same one that generated the study's tables (DOIs 10.5281/zenodo.21325743 and 10.5281/zenodo.21327595). One study in the series, on the Brazilian macro regime score (10.5281/zenodo.21402939), went further and published the indicator's own weights — a deliberate exception, chosen precisely because in that case transparency was worth more than reserve.

None of this requires a budget. It requires a decision made early: that the study will be judged by what a stranger can redo, not by what the author claims to have done. How each test is designed and audited internally is another conversation, longer than a public article allows — but the principle fits in one sentence: a cited number is a checkable number.

What reproducibility does not solve

Honesty about the limit is due. A perfectly reproducible study can still have picked the convenient window, tested twenty hypotheses and reported one, or mistaken coincidence for pattern. Reproducibility audits the execution, not the intent. The defense against that other problem has a name of its own — pre-analysis registration, the hypothesis declared before the model runs — and it is the subject of the next article in this track.

Frequently asked questions

Are reproducibility and replicability the same thing?

No. Reproducibility: same data, same method, same result. Replicability: new data, same method, compatible result. The first audits the execution; the second, the validity of the finding.

Can a study built on proprietary data be reproducible?

Partially. The author can publish the code, the exact list of series and the transformations applied, so that anyone holding the same license can redo everything. Less than ideal, but far more than the current standard.

What is a checksum like SHA-256 for?

It is a sequence computed from the file's contents. If a single byte changes, the sequence changes. Publishing it with the data lets the reader verify they downloaded exactly the file used in the study.

Can anyone check Radar Perene's results?

The published working papers and their data packages sit in an open repository, with DOIs and checksums, no registration and no payment. What is claimed there can be redone by anyone willing to put in the time.

---

Continue the track: Pre-analysis registration: declaring the hypothesis before running the model

House readings: what the indices record today is in the Diário; the historical episodes those data support, in the Atlas.

Reproducing a specific study, step by step, is work of another depth — the kind the house discusses in conversation, not in an article.

This is the Radar’s memory. Today’s reading — regime, 5 lenses and the day’s analogs — is live, free.

Subscribe to Perene Semanal — US$ 29/mo →

See today’s reading →