Deflated Sharpe Ratio (DSR)

A test statistic that adjusts an observed Sharpe ratio for the number of trials conducted, the length of the track record, and the non-normality of returns, estimating the probability that the true Sharpe ratio exceeds zero.

A Sharpe ratio reported in isolation is an incomplete sentence. The same observed value can be strong evidence of skill or the guaranteed by-product of a large search, and the number alone cannot tell you which. The deflated Sharpe ratio, proposed by Bailey and Lรณpez de Prado, completes the sentence by conditioning the statistic on how it was obtained.

The logic rests on an uncomfortable result from extreme value theory: under the null hypothesis of zero skill, the expected maximum Sharpe ratio across N independent trials is not zero โ€” it grows with N. A research process that tries many configurations and reports the best one will therefore produce impressive-looking Sharpe ratios out of pure noise, reliably and reproducibly. The DSR takes the observed Sharpe of the selected strategy and asks whether it exceeds what the best of N skill-less trials would be expected to achieve, additionally correcting for track-record length and for the skewness and excess kurtosis of returns โ€” both of which make naive Sharpe inference optimistic, and both of which are endemic in trading strategies with asymmetric payoffs.

The output is a probability that the true Sharpe ratio is positive, rather than a raw ratio to be admired. A strategy that clears a conventional significance bar on its headline Sharpe can fail decisively once its sibling trials are counted.

Like the probability of backtest overfitting, the DSR is honest only if N is honest. Understate the number of trials โ€” by forgetting them, or by never recording them โ€” and the deflation is too gentle, laundering luck into apparent significance. The corrective machinery of multiple-testing statistics presumes the research hygiene of a kill ledger; without one, deflation degenerates into decoration on an overfit backtest.

← All glossary terms