How the Backtest Honesty suite measures

The per-instrument battery for Backtest Honesty, the receipts each measurement will carry, and the honest state of what is published today. Looking for what you can buy in this family instead? That is the Backtest Honesty suite hub.

What this suite measures, and why

The Backtest Honesty suite measures whether the evidence for a strategy would survive hostile review. It attacks the two places backtests actually fail: the search process that produced the equity curve, and the dataset underneath it. One instrument bounds how much of a reported result could be an artifact of how many things were tried; the other audits whether the data itself is telling the truth.

Most backtests are not wrong at the strategy layer. They are wrong at the evidence layer — untracked trials, silent data defects, the best window presented as the typical one. A clean statistic computed on a lying dataset is still a lie, which is why this suite audits both.

The measurement battery, per instrument

Each instrument’s battery is summarized from its own published specification. Every dimension follows the six-stage lifecycle defined on the methodology page — pre-registered, content-hashed, deterministic, and shipped with its can-fail proof. Written in the future tense because nothing has been measured yet.

The fidelity battery pointed at a dataset instead of a feed: whether the data under a clean-looking backtest is itself telling the truth.

  • Carry-forward and merge artifacts
  • Per-side staleness inside the dataset
  • Gap structure — what is missing, and where
  • Bar provenance — whether bars are what they claim to be

Honest limit: It audits the data, not the strategy. A clean dataset does not validate the backtest run on it.

Full instrument page

Statistical bounds on how much of a backtest’s reported performance could be an artifact of the search that produced it.

  • Trial accounting — how many configurations were evaluated before this one was selected
  • Probability of backtest overfitting, via combinatorially symmetric cross-validation
  • Deflated performance statistics, adjusted for trials, track length, and non-normal returns
  • Walk-forward degradation curves, published whole — not the best window

Honest limit: It cannot certify that a strategy works — no statistic can. It bounds the ways the evidence can be an illusion, and every output ships with its can-fail proof.

Full instrument page

The event-study verdict engine: effect estimates with stationary-block-bootstrap confidence intervals, HAC standard errors, and Benjamini–Hochberg false-discovery-rate masking — zero-dependency, with byte-reproducible verdicts.

  • Effect estimation with stationary-block-bootstrap confidence intervals
  • HAC standard errors
  • Benjamini–Hochberg FDR masking across the verdict set
  • Keystone self-test — a planted known effect must be recovered while forty placebos are killed
  • Byte-reproducibility — the same inputs yield the identical verdict, byte for byte

Honest limit: It renders verdicts on the event studies it is given; it does not generate hypotheses, and a recovered effect is a statistical finding, not a trading recommendation.

Full instrument page

The receipts this suite will publish

Every measurement from the Backtest Honesty suite carries the same receipt chain, as defined in the full receipt taxonomy. Four of the six receipt types apply from the first measurement onward:

Pre-registration

Each instrument’s battery is written and content-hashed before its first data collection. Any post-hoc change to the methodology would change the hash, and would be visible.

Content hash

Data, code, and results are SHA-256 hashed, so a published result from this suite is recomputable: declared outputs from declared computation on declared inputs.

Can-fail proof

Every test in the battery ships with a demonstration that it could have failed — a known defect injected and detected. A test that cannot fail proves nothing.

Kill entries

Hypotheses registered for this suite and killed by the data are published with the same visibility as registrations. No silent disappearances.

What is published today

MEASUREMENTS PUBLISHED: NOT YET PUBLISHED

Nothing. No measurement from the Backtest Honesty suite has been published, and this page will say so until one has. That order of operations is deliberate: pre-registration means the methodology is public before the data exists, so no result from this suite can ever have shaped the method that produced it.

When the measurement pipeline goes live, this page links, per instrument: the registered methodology document and its hash, point-in-time data manifests, battery results with confidence intervals and honest nulls, the can-fail proof beside every test, and kill entries for whatever does not survive.

Backtest Honesty instrumentsFull methodology