Backtest Honesty instruments
Three instruments for the three places the evidence can lie: the dataset underneath a backtest, the search that produced the result, and the verdict finally rendered on it.
Receipt page · proof before pitchHow the Backtest Honesty suite measuresVersioned methodology · pre-registrations · proof corpus · changelog · recomputation steps- Instruments
- 03
- Published measurements
- NOT YET PUBLISHED
The problem, in the customer’s words
The backtest looked excellent and the live results do not. I cannot tell whether the strategy was wrong, the data was wrong, or I simply kept trying things until something worked.
Not a customer quotation. Nothing on this site is attributed to a customer, and no testimonial appears anywhere on it — this is the problem the family exists for, written the plain way it gets put.
What this family is for
The family that asks whether the evidence for a strategy would survive hostile review — and attacks the evidence layer rather than the strategy layer, because that is where backtests actually fail.
Three instruments for three places the evidence can lie. The Overfit Auditor bounds how much of a reported result could be an artifact of the search that produced it: trial accounting, probability of backtest overfitting via combinatorially symmetric cross-validation, deflated performance statistics adjusted for trials, track length and non-normal returns, and walk-forward degradation curves published whole rather than at their best window.
Data Forensics points the same fidelity battery at the dataset underneath — carry-forward and merge artifacts, per-side staleness, gap structure, bar provenance — because a clean statistic computed on a lying dataset is still a lie. The Reproducible Verdict Kernel renders the verdicts themselves: stationary-block-bootstrap confidence intervals, HAC standard errors, Benjamini–Hochberg false-discovery-rate masking, and byte-reproducible output from a zero-dependency core.
They are a family because you cannot audit a search process on data you have not audited. None of them can certify that a strategy works — no statistic can. They bound the ways the evidence can be an illusion.
The 3 instruments in this suite
Every card states both halves: what the instrument measures, and what it does not establish. The figure on each card is an anticipated pre-launch indication, priced out in full in the table below, and no measurement from this suite has been published yet.
Mathematical bounds on how hard you tortured the data.
Anticipated£119/mo · £1,190/yr
- Measures
- Trial accounting, the probability of backtest overfitting via combinatorially symmetric cross-validation, deflated performance statistics adjusted for trials, track length and non-normal returns, and the walk-forward degradation curve published whole rather than at its best window.
- Does not establish
- It does not tell you a strategy works. No statistic can. It bounds the ways in which the evidence for a strategy can be an illusion — that is all of it, and that is the honest maximum on offer.
Your stats can be clean and your data still lying.
Anticipated£129/mo · £1,290/yr
- Measures
- The fidelity battery pointed at a dataset rather than a live feed: carry-forward and merge artifacts, per-side staleness, gap structure, and bar provenance — with the honest null published wherever the dataset cannot support a dimension.
- Does not establish
- It does not validate your strategy, does not repair data, and does not establish vendor intent. A dataset that passes every check has established exactly one thing — that it is what it claims to be — and nothing about whether a rule fitted to it survives contact with the future.
Same inputs, same verdict, byte for byte — or it is not a verdict.
Anticipated£249/mo · £2,490/yr
- Measures
- Event-study verdicts with stationary-block-bootstrap confidence intervals, HAC standard errors, Benjamini–Hochberg false-discovery-rate masking across the family, and byte-reproducible output from a zero-dependency core with its can-fail proof built in.
- Does not establish
- It does not find you an edge. It adjudicates evidence for effects you already claim, generates no candidates of its own, and mechanically returns “no effect survives” on a family you were fond of. That answer is not a malfunction.
What the Backtest Honesty suite is anticipated to cost
Every figure below is an anticipated pre-launch indication read from the instrument catalogue — not a confirmed price, not a quote and not an offer. Nothing on this site is purchasable: there is no checkout, no cart and no payment link on any page. The launch list is the only thing open.
| Instrument | Billed monthly | Billed annually |
|---|---|---|
| P2 Overfit Auditor | £119/mo | £1,190/yr |
| P10 Data Forensicsdataset certifications quoted per engagement | £129/mo | £1,290/yr |
| P18 Reproducible Verdict Kernelenterprise from £15,000/yr | £249/mo | £2,490/yr |
The family arithmetic
What licensing every instrument in this family separately would come to. The lines stay apart where their units differ — a per-seat licence, a flat account licence and a scoped engagement cannot be added into one number, and a single tidier figure would be a false one.
- Per account, billed monthly
- £497/moP2 + P10 + P18, summed at their monthly figures.
- Per account, billed annually
- £4,970/yr£497 × 10 months = £4,970. Twelve months of service, ten months paid.
Check the arithmetic
Across P2 + P10 + P18, billed monthly for a year: £497 × 12 = £5,964. The same instruments billed annually: £4,970. The gap is exactly two months — £994.
A three-year term pays eight months per year of service: 8 × 3 = 24 months. At £497/mo across P2 + P10 + P18, that is £11,928 for three years of service. Two years of monthly billing is £5,964 × 2 = £11,928. The same number: three years of service for what two years of monthly billing costs.
What a suite licence would change
Nothing yet, and I will not pretend otherwise. The totals above are the plain arithmetic sum of licensing each instrument on its own — that sum is a fact you can add up yourself. I have not directed a suite or bundle discount, so none appears here. Publishing an invented one would be the exact class of fabrication this site exists to argue against, and a figure I would then have to withdraw is worth less than an honest blank.
What has been directed is the commitment ladder below and, for the instruments licensed per seat, the seat bands underneath it. Both are anticipated and both await ratification. Multi-entity and multi-network groups are licensed by negotiation rather than by seat count: flat-rate instruments scale by entity, never by seat, and a negotiated number is not a list price so none is published. The full model sits on the pricing page, and the terms themselves on the licensing page.
Longer terms: one month less per year, per step
The months paid for each year of service fall from twelve on monthly billing to eight on a three-year term.
| Term | Months paid per year | What it does to the annual list |
|---|---|---|
| Monthly billing | 12 | Twelve months paid for twelve months of service. Costs 20% more than the annual list — that uplift is the whole difference. |
| One-year term | 10 | Ten months paid for twelve months of service. This is the annual list in the table above. |
| Two-year term | 9 | Nine months paid per year of service — a tenth off the annual list, with the figure held for the term. |
| Three-year term | 8 | Eight months paid per year of service — a fifth off the annual list, with the figure held for the term. |
Standing policy — read beside every figure above
- No promise of profit, ever. Hadal makes no performance claims and carries no implied edge. Past measurements describe instrument behaviour — never future returns.
- You own risk management. Hadal cannot control it and does not insure it. Good tools do not fix bad discipline — and this site says so.
- Analytical tools, for discretionary use. Nothing here is investment advice or a recommendation to trade. Every decision, and every outcome, is yours.
The same statement stands in the footer of every page.Terms Privacy
How the instruments connect
Input, then which instrument runs, then what it returns, then what the next one consumes. This diagram describes the workflow only. No step on it is a result, and this family has published none.
- 01
Input: the strategy, its parameter history, its trial count, and the dataset
The trial count is an input, not an output. Every downstream statistic conditions on how many configurations were evaluated before this one was selected.
- 02
P10 Data Forensics audits the substrate first
Returns carry-forward and merge artifacts, per-side staleness, gap structure and bar provenance, with an honest null where the dataset cannot support a dimension. What the next step consumes is a dataset whose defects are named.
- 03
P2 Overfit Auditor bounds the search that produced the result
Consumes the audited dataset and the declared trial count, and returns the probability of backtest overfitting, deflated statistics, and the walk-forward degradation curve published whole.
- 04
P18 Reproducible Verdict Kernel renders the verdict
Consumes the events and series that reached this stage and returns interval estimates, HAC standard errors and FDR-masked findings — byte-reproducible by anyone holding the same inputs.
A failed substrate check re-enters at step 01 rather than proceeding: a clean statistic computed on a lying dataset is still a lie.
The Backtest Honesty workflow, end to end: what enters the family, which instrument acts at each stage, and what the next stage consumes.
A process diagram establishes nothing about any market. It states how the instruments are wired to each other — not that any of them has measured anything.
The diagram is the shape; the suite-workflow documentation is the order you actually run them in, including which stages can be skipped and which cannot. It is published before the product, so the intended sequence can be held against whatever ships.
One worked example
NOT YET PUBLISHED
No worked example exists for this suite, because no measurement from it has been published. Rather than stage one from invented data, the five slots a worked example will carry are named in the empty figure beneath this note, with their contents stated and every slot empty. Nothing here is an example run, and no figure, date or digest appears in it.
- Inputs
The equity curve, the parameter history, the declared trial count, and the dataset the backtest was actually run over.
NOT YET PUBLISHED- The run
Which estimators ran at which settings, and the decision rules fixed by pre-registration before any statistic was computed.
NOT YET PUBLISHED- Artifact digest
The digest of the input bundle and of the verdict sheet, so a reviewer can regenerate the same verdict byte for byte.
NOT YET PUBLISHED- The finding
What survived and what did not: the deflated statistic beside the reported one, and the walk-forward degradation curve published whole.
NOT YET PUBLISHED- The limits
That a surviving statistic is a bound on illusion, not evidence a strategy works. No statistic can carry the second claim.
NOT YET PUBLISHED
The five slots a Backtest Honesty worked example will carry. Each row states what the slot will contain; none of them contains it yet.
An empty structure establishes nothing. It shows the shape of the evidence this suite intends to publish, and the absence of that evidence today.
When the first example publishes it will appear here with its artifact downloadable and its digest on the Backtest Honesty receipt page, where the steps for re-deriving it are already written out.