22 / 97
Deflated Sharpe Ratio (DSR)
Written DSR throughout.
Working definition
A test statistic that adjusts an observed Sharpe ratio for the number of trials conducted, the length of the track record, and the non-normality of returns, estimating the probability that the true Sharpe ratio exceeds zero.
A Sharpe ratio reported in isolation is an incomplete sentence. The same observed value can be strong evidence of skill or the guaranteed by-product of a large search, and the number alone cannot tell you which. The deflated Sharpe ratio, proposed by Bailey and López de Prado, completes the sentence by conditioning the statistic on how it was obtained.
The logic rests on an uncomfortable result from extreme value theory: under the null hypothesis of zero skill, the expected maximum Sharpe ratio across N independent trials is not zero — it grows with N. A research process that tries many configurations and reports the best one will therefore produce impressive-looking Sharpe ratios out of pure noise, reliably and reproducibly. The DSR takes the observed Sharpe of the selected strategy and asks whether it exceeds what the best of N skill-less trials would be expected to achieve, additionally correcting for track-record length and for the skewness and excess kurtosis of returns — both of which make naive Sharpe inference optimistic, and both of which are endemic in trading strategies with asymmetric payoffs.
The output is a probability that the true Sharpe ratio is positive, rather than a raw ratio to be admired. A strategy that clears a conventional significance bar on its headline Sharpe can fail decisively once its sibling trials are counted.
Like the probability of backtest overfitting, the DSR is honest only if N is honest. Understate the number of trials — by forgetting them, or by never recording them — and the deflation is too gentle, laundering luck into apparent significance. The corrective machinery of multiple-testing statistics presumes the research hygiene of a kill ledger; without one, deflation degenerates into decoration on an overfit backtest.
The observed ratio this corrects — with the error bars a short track record puts around it — is computed by the Sharpe ratio calculator.
Commonly confused with
Neighbouring concepts that get used interchangeably, and the distinction that actually separates them.
- Sharpe ratio
The raw ratio describes a return series. The deflated one conditions that description on how the series was selected — how many trials, how long the record, how non-normal the returns. Two numbers computed from the same data, answering different questions, and only one of them is a test.
- Probability of backtest overfitting
Both correct for multiple testing and they report different things. PBO estimates the chance the selected configuration underperforms out of sample; DSR estimates the probability the true Sharpe ratio exceeds zero. Neither subsumes the other, and both need the same honest trial count.
- An annualised Sharpe ratio
Annualising is a scaling convention that assumes independent, identically distributed returns. Deflating is a correction for selection and for non-normality. Annualising a ratio does not deflate it, and a large annualised figure from a wide search is exactly what the deflation exists to catch.
- A significance threshold
The DSR output is a probability that the true Sharpe is positive, not a ratio to be compared against a conventional bar. A strategy can clear a familiar significance threshold on its headline Sharpe and fail decisively once its sibling trials are counted.
How to measure it in your own data
A definition you cannot test is a definition you have to take on trust. This is the shortest honest route from the concept to a number you computed yourself.
- Records you need
The return series with its length, the skewness and excess kurtosis of those returns, and N — the number of trials the reported strategy was selected from. N is the input that decides the answer and the one nobody has unless they decided in advance to keep it.
- What you compute
Compare the observed Sharpe of the selected strategy against what the best of N skill-less trials would be expected to achieve, correcting additionally for track-record length and for the skewness and excess kurtosis of the returns.
- What the answer tells you
The uncomfortable result underneath is from extreme value theory: under a null of zero skill, the expected maximum Sharpe across N independent trials is not zero and grows with N. So a process that tries many configurations and reports the best produces impressive ratios out of pure noise, reliably and reproducibly. Understate N and the deflation is too gentle, which launders luck into apparent significance — deflation on an unrecorded search is decoration rather than correction.
If this has already cost you
A headline Sharpe ratio can be corrected for the trials it was selected from, the length of the record, and the shape of the returns.
- Overfit Assay“Is my backtest real, or did I fit it to noise?”Will not establish: Whether the strategy will be profitable. A backtest that survives the battery is a backtest that was not obviously fitted — it is not a forecast, and the report says so on its first page.
Intake is not open yet, so none of these can be commissioned today. They are listed here so you know the measurement exists and what it would and would not settle — the launch list hears first.
Work it out yourself
Free calculators that take this concept as an input. Each shows its working, so the number it gives you can be checked rather than taken on trust.
Questions and answers
What does the deflated Sharpe ratio actually output?
A probability that the true Sharpe ratio is positive, rather than an adjusted ratio to be admired alongside the original. That change of type is the point: the question stops being "how high is it" and becomes "how likely is it that anything is there".
Why does the number of trials change the answer so much?
Because of a result from extreme value theory that is hard to argue with. Under the null hypothesis of zero skill, the expected maximum Sharpe across N independent trials rises with N. The best of a large search will therefore look good whether or not any real effect exists, and the deflation asks whether the observed value exceeds what that search would have produced anyway.
How is DSR different from the probability of backtest overfitting?
DSR and PBO correct for the same disease and report different symptoms. PBO estimates how likely the selected configuration is to underperform out of sample; DSR estimates the probability that the true Sharpe ratio is above zero. Using both is reasonable; treating either as a substitute for an honest trial count is not.
What happens if I do not know how many trials I ran?
Then the deflation cannot be trusted in the direction that matters. Understating N makes the correction too gentle and turns luck into apparent significance — a result that looks more rigorous than the raw Sharpe while being no more trustworthy. The corrective machinery presumes the research hygiene of a recorded trial count; without one, it is decoration on an overfit backtest.
Related terms
Derived from the links this entry makes and the entries that link back to it.
Where the term is used
Instrument pages whose published copy uses this term. Each page states what it measures and what it does not establish.
In the research
Deflated Sharpe Ratio (DSR) comes up in two research notes on this site.