Backtest Honesty instruments
The family that asks whether the evidence for a strategy would survive hostile review — and attacks the evidence layer rather than the strategy layer, because that is where backtests actually fail.
How we measureHow the Backtest Honesty suite measuresWhat this family is for
Three instruments for three places the evidence can lie. The Overfit Auditor bounds how much of a reported result could be an artifact of the search that produced it: trial accounting, probability of backtest overfitting via combinatorially symmetric cross-validation, deflated performance statistics adjusted for trials, track length and non-normal returns, and walk-forward degradation curves published whole rather than at their best window.
Data Forensics points the same fidelity battery at the dataset underneath — carry-forward and merge artifacts, per-side staleness, gap structure, bar provenance — because a clean statistic computed on a lying dataset is still a lie. The Reproducible Verdict Kernel renders the verdicts themselves: stationary-block-bootstrap confidence intervals, HAC standard errors, Benjamini–Hochberg false-discovery-rate masking, and byte-reproducible output from a zero-dependency core.
They are a family because you cannot audit a search process on data you have not audited. None of them can certify that a strategy works — no statistic can. They bound the ways the evidence can be an illusion.
The 3 instruments in this suite
Each card links the full specification. Prices shown on instrument pages are anticipated pre-launch figures, and no measurement from this suite has been published yet.