Gatun Capital

How we test

SEARCHED HERE NOTHING TOUCHED HERE older half of the data more recent half

Two periods, strictly separated

The first half of the data is where we search and adjust. The second stays untouched and passes judgement at the end. A rule that only looks good in the first half is discarded — there, everything you went looking for looks good.

discarded A SINGLE PEAK adopted A PLATEAU

A plateau, not a peak

A setting is adopted only if its neighbours are good too. A single good value between two bad ones is chance — and chance does not repeat.

QUARTER END not yet known REPORTING DATE used from here on

Only what was known at the time

Company figures count from the day they were reported — not from the quarter end they refer to. Weeks lie in between. Anyone who skips those weeks is testing with knowledge from the future.

TODAY'S COMPANIES ONLY ALL THAT EVER EXISTED red = disappeared later

Companies that vanished still count

Testing only the companies that still exist today leaves out every one that went bankrupt or was taken over. Among smaller companies that is two thirds of them. With us they are included.

How to read an evaluation

The terms asset managers, funds and institutional investors use when they talk about results — and what can be overlooked in each one. We use them exactly as they stand here.

Show the ten terms
Return p. a.
(final value / initial value) ^ (1 / years) − 1

The geometric annual return: the constant rate that turns the initial amount into the final one. It is the only figure that can be compared across periods of different length.

Not to be confused with the average of the yearly figures. Plus 50 % and minus 50 % average to zero, but in fact leave a quarter less.
Volatility
standard deviation of monthly returns × √12

How widely the monthly results scatter around their mean, annualised. The usual measure of how restless a curve is.

Says nothing about direction. A path can be calm and still fall steadily — and a good month raises volatility just as much as a bad one.
Sharpe ratio
(return − risk-free rate) / volatility

How much return is left per unit of volatility once you subtract the rate obtainable without any risk. The standard measure when comparing two approaches of different risk.

It penalises upward swings as much as downward ones, and it can be flattered — for instance by rarely priced positions whose quotes look smoother than they are.
Sortino ratio
excess return / deviation of the losing months only

The same idea as Sharpe, but only downward movement counts. Anyone running an approach with rare large gains is not punished for them here.

It rests on fewer observations than Sharpe and is therefore less reliable — over short periods it says almost nothing.
Maximum drawdown
lowest point / previous high − 1

The largest loss from peak to trough that would have had to be endured. In practice the single most important figure: it decides whether anyone stays invested.

It depends on the measurement interval. Computed monthly it comes out smaller than daily, because the worst day inside the month never becomes visible.
Recovery time
months from the high to the next high

How long it took to make a decline back. Losing a fifth is one thing if it is recovered after eight months, and quite another if it stands for four years.

Rarely quoted, although it is the real reason investors give up.
Beta and correlation
slope and co-movement relative to the market as a whole

How much of a performance simply comes from the market having risen. A result only becomes interesting once that part is subtracted.

Neither figure is stable. In a crash almost all shares move together — correlations rise precisely when one would like to rely on them.
Hit rate
share of months with a positive result

How often things went up. Easy to grasp, and therefore often quoted.

Worthless on its own. Nine small gains and one large loss make nine positive months out of ten — and still a bad year.
Turnover
volume traded / average holdings

How much of the portfolio is moved in a year. It drives costs directly: every reshuffle pays the spread.

A backtest without costs looks better the higher the turnover. Comparing two approaches without deducting costs is therefore worthless.
Capacity
the amount of money at which one's own order moves the price

Every approach has a ceiling. Invest more than a share trades in a day and you drive up your own entry when buying and push down the proceeds when selling.

Almost always left out of backtests, because there one trades at the closing price in any quantity one likes.

What cannot be recalculated, we do not claim.

Four measurements

Above is what we watch for when testing. Here is how large the effect is — counted from public market data, not from our results. Anyone with the same sources arrives at the same figures.

Show the four measurements

How many companies disappear

US common stocks listed at the end of 19987 400
of those, no longer tradable today6 117
listed at the end of 20055 624
of those, no longer tradable today3 970
share that disappeared since 199882.7 %

Four out of five companies that existed then no longer exist today — taken over, merged, bankrupt. Anyone basing a thirty-year test on today's roster of companies tests the survivors alone and leaves out precisely the cases that went wrong. Such a test is not imprecise; it is systematically too favourable.

Own count over the complete price history (US common stocks, 43.6 million daily prices). A company counts as disappeared when its price series ends more than 200 days before the last trading day. Prices as of 10 Sep 2026.

When company figures actually become available

quarterly reports evaluated681 370
median between quarter end and report44 days
three quarters of reports within65 days
nine tenths of reports within90 days
later than 30 days622 624

At the quarter end nobody knows what happened during the quarter. On average six weeks pass before the report; for one in ten it is more than three months. A backtest that uses company figures as of the quarter end trades on knowledge that existed nowhere at that time — in well over nine cases out of ten. We count from the day of publication.

Own count over all quarterly reports in the data set: the gap between the reference date and the publication date, values above 400 days excluded.

How deep the market falls, and for how long

from Aug 2000, low −45 %75 months
from Oct 2007, low −51 %53 months
from Dec 2021, low −24 %24 months
longest recovery75 months

Three declines of more than a fifth in barely thirty years. The longest stood for over six years before the old level was regained. A return without those two figures — how deep and how long — does not describe what an investor actually had to sit through. That is why both stand next to every return figure we report.

Month-end closing prices of the S&P 500 index fund including dividends, Dec 1997 to today. Depths rounded to whole percent. Publicly verifiable.

A single share is not the market

index with dividends, 2016 to 2025+294.9 %
titles listed at the start of 20164 582
of those, still tradable at the end of 20252 479
median title over the same ten years+41.7 %
better than the index625 of 4 582
lost more than nine tenths666 of 4 582

The index quadrupled; the median share did not even gain half. Only about one title in seven beat the index, and roughly as many lost almost everything. An average across shares therefore says little; the distribution says everything. Whoever selects must be measured against that distribution — and against the titles that were no longer there at the end.

Own count: all US common stocks with a price at the start of January 2016, total return to the last available price, 31 Dec 2025 at the latest. Delisted titles remain included at their final price.

Models a result has to stand up to

A figure on its own says nothing. It only becomes a statement once you check how much of it would have come about without any skill. Tools for that have existed for decades, and we apply them to ourselves.

Show the five models
Market model and multi-factor modelsSince the work of Sharpe (1964) and Fama and French (1993) it has been known that a large part of any equity return can be explained by a few common influences — above all the market itself, alongside company size and valuation. Anyone presenting a performance record has to show what is left once those explanations are subtracted. Everything else is a market that went up.
Cost modelAny backtest that buys and sells at the closing price is useless. A sound model deducts the spread between bid and ask, assumes a delay between signal and execution, and allows for the fact that one's own order moves the price. For smaller companies this is not a detail but the difference between a result and none at all.
Multiple testing and corrected metricsTry a hundred variants and you will find a good one by chance alone. Bailey and López de Prado proposed corrections for this — a Sharpe ratio deflated by the number of trials, and an estimate of the probability that a result is pure overfitting. Anyone who does not state how many trials they ran cannot defend their figure.
Randomisation testThe order of the results can be reshuffled again and again and the backtest replayed a thousand times over. That shows how wide the range of possible paths is and where the actual result sits within it. Close to the edge of the range means a great deal of luck was involved.
Behaviour in crises, considered separatelyAn average over thirty years conceals what happened in 2000, 2008, 2020 and 2022. Those stretches are reported separately, because they answer the question people actually ask: what happens when things go badly.