P2Backtest Honesty

Overfit Auditor

Mathematical bounds on how hard you tortured the data.

Anticipated priceAnticipated · not ratified · not an offer

£119/mo

Annual term
£1,190/yrTwelve months of service for the price of ten.
Billed monthly
£1,428/yr£119/mo × 12. Monthly billing costs twenty per cent more than the annual term — the uplift is on monthly, it is not a discount on annual.
The difference
£238/yrTwo months in twelve. Paying annually saves two months, not twenty per cent — the two are different numbers and only one of them is true.
Open to academic access

Backtest overfitting is an open research literature, not a trade secret. The instrument implements the published tests, so an academic using it is checking my arithmetic — which is the point. Anticipated academic rate: 50% of the annual list price — £595/yr. The academic rate applies instead of the seat and term bands, never on top of them. Stacked, the three would land below the cost of serving the account, and a rate I cannot honour is worse than one I never offered.

Academic access — who qualifies and on what terms

Nothing on this site is on sale. There is no checkout, no cart and no payment link on any page. Prices publish pre-launch so they can be read, compared and checked rather than requested — seat bands, group licences and multi-year terms are set out in full below.

Standing policy — read this beside the figure

  • No promise of profit, ever. Hadal makes no performance claims and carries no implied edge. Past measurements describe instrument behaviour — never future returns.
  • You own risk management. Hadal cannot control it and does not insure it. Good tools do not fix bad discipline — and this site says so.
  • Analytical tools, for discretionary use. Nothing here is investment advice or a recommendation to trade. Every decision, and every outcome, is yours.

The same statement stands in the footer of every page.Terms Privacy

SpecificationValue
Catalogue no.P2
SuiteBacktest Honesty
MethodologyHow the Backtest Honesty suite measures
AvailabilityPre-launch — not on sale
Anticipated price£119/mo · £1,190/yr
DeliveryMarketplace SKU · hub subscription
Published measurementsNOT YET PUBLISHED
Provenance artifactNOT YET PUBLISHED

Every backtest is a claim, and most of them are confessions. The Overfit Auditor puts a number on the difference.

The instrument takes a strategy’s backtest — the equity curve, the parameter history, and crucially the number of trials it took to find it — and computes the statistics the marketing screenshot leaves out: the probability that the selected configuration is an artifact of the search rather than a property of the market, and how much of the reported performance survives once you account for how many things were tried before this one worked.

What it measures

The auditor measures the strategy rather than the data underneath it, taking the trial count as an input you supply.

  • Trial accounting. How many configurations were evaluated, explicitly or implicitly, before this one was selected. Every downstream statistic conditions on this number, and almost nobody records it. The instrument makes it unavoidable.
  • Probability of backtest overfitting. Via combinatorially symmetric cross-validation: how often does the in-sample winner underperform the median of its rivals out-of-sample?
  • Deflated performance statistics. The reported Sharpe ratio adjusted for trials, track-record length, and the non-normality of returns — the version of the number a referee would accept.
  • Walk-forward degradation. How performance decays from calibration window to unseen window, rolled through time — the degradation curve published, not the single best window.

A reading from my own book

Pointed at this estate’s own research, the auditor currently returns PBO 0.16 — the probability that the in-sample-best configuration underperforms the median out-of-sample. Below the coin-flip line, which is what “mostly holds up out-of-sample” means and all it means.

I publish the number for two reasons. It is unremarkable, and an auditor that only ever reports alarming figures is selling alarm rather than measurement. And it is mine — a PBO computed on someone else’s book proves nothing about the instrument, whereas one computed on my own can be checked against everything else I publish about that book, including the parts that are unflattering.

Read it beside the companion reading from the Epistemic Harness: on the same estate, 30 nominal tests collapse to ~6.24 effective independent ones. A low PBO and an inflated trial count are not in tension — they are two different ways the same book can mislead you, which is why the two instruments are separate. The estate’s first end-to-end edge hunt — twelve Bonferroni-clean cells and not one tradable edge among them — is written up in the edge that was real and worth half a pip.

What it does not do

The Overfit Auditor does not tell you a strategy works. No statistic can. It bounds the ways in which the evidence for a strategy can be an illusion — that is all, and that is the honest maximum on offer. Every output ships with its can-fail proof and its computation receipt, recomputable from the same inputs.

Who it is for

Anyone about to risk capital on a curve they fitted: independent quants, prop-firm evaluees who paid for an evaluation seat, and desks reviewing external track records. If your backtest cannot survive this instrument, the market will run the same test with your money.

The ship gate

No instrument is sold until it does what this page says it does. Where a page is written in the future tense, that tense is a statement about timing rather than a hedge about capability: the instrument is not finished, so it is not listed as available, not priced as available, and not sold. It waits.

Nothing described in this catalogue is a placeholder that will quietly disappear. An instrument that turns out to be wrong gets a kill-ledger entry, not a deletion — which is the only version of that promise anyone can check.

Commercial terms

Published in full, pre-launch, so they can be read and checked rather than requested. Every figure is an anticipated indication I have set and not yet ratified, and every derived figure is the arithmetic of the one above it — shown, not asserted. Nothing here is purchasable: there is no checkout on this site.

Why Overfit Auditor is priced the way it is

What the figure buys
The licence covers the audit battery run against the backtests you submit: trial accounting, the probability of backtest overfitting via combinatorially symmetric cross-validation, deflated performance statistics, and the walk-forward degradation curve rolled through time rather than reported at its best window. Every output carries its can-fail proof and its computation receipt, so a reviewer can recompute it from the same inputs instead of accepting it on trust. The edge of the scope is where the page puts it: the instrument will not tell you a strategy works, because no statistic can, and it audits the strategy rather than the data underneath it. It also conditions on what you record — the trial count is an input you supply, and the instrument's work is to make that number unavoidable, not to guess it on your behalf.
Why it is priced this way
The rate is flat and recurring against the strategy book of one licensed entity rather than counted per seat, because this instrument is not a surface a person occupies: its outputs are recomputable records, and a whole review circle can read them from the same file. Seat bands belong to instruments a person works within, and applying one here would be pricing readers instead of the thing being audited. Recurrence falls out of the arithmetic itself — every re-fit, every added parameter and every further month of track record changes the trial count, and a trial count that has moved invalidates every statistic conditioned on it. An audit that was true once is a snapshot of a search that has since carried on.
What the alternative costs
The methods are public. Backtest-overfitting probability, combinatorially symmetric cross-validation and the deflation of a Sharpe ratio for trials, track length and non-normality are all in the literature, and a competent quant with the papers and the time can implement them; none of the mathematics here is secret. What that path actually costs is the implementation, the validation of the implementation, and the part that is hardest to hold — the discipline of recording trial counts honestly while you are still searching, rather than reconstructing them afterwards when the incentive runs the other way. A platform's own backtest report is not a substitute either, because it describes the configuration that won and not the search that produced it, so it cannot price a multiplicity it never observed; and doing nothing simply leaves the same test to be run by the market against live capital, on its schedule rather than yours.

Every figure on this page is an anticipated indication awaiting ratification, and nothing here is purchasable. The reasoning above is published for the same reason the arithmetic below is: a price you can interrogate is worth more than a price you have to accept.

The two-SKU split

One route. This instrument has no marketplace equivalent, and the cell says so rather than sitting empty — an empty cell reads as an omission, and this is a fact about how the instrument is sold.
RouteWhat it isAnticipated
Marketplace SKUNone. Buyers do not browse trading-platform marketplaces for an instrument of this shape, so there is no platform SKU for the hub price to sit under. It sells from the hub or not at all.NO MARKETPLACE ROUTE
Hub subscription (this site)The deepening: hosted runs, the published methodology behind them, and the content-hashed artifacts that let a stranger re-derive the result. This is the tier the figure on this page prices.£119/mo · £1,190/yr
The price rule

The at-or-below-marketplace cap binds the instruments that ship through a platform. This one does not, so the cap has nothing to bind. What holds in its place is simpler and checkable: the figure above is the only figure. There is no separate, higher list price it is discounted from, nothing on this site has ever been sold at any price, and so there is no “was” anywhere to strike through.

The commitment ladder

One month less paid per year of service, per step, taken off this instrument's own anticipated annual rate of £1,190/yr. The rate is held at the figure you sign for the whole term, so a multi-year commitment fixes the price as well as reducing it.
TermMonths paid per yearWhat it meansAnticipated
Monthly billing12£119/mo × 12. Twenty per cent more than the annual term — that is the uplift for paying monthly, not a discount for paying annually.£1,428/yr
1-year term10Twelve months of service for the price of ten. This is the annual rate every band below is taken off.£1,190/yr
2-year term9Ten per cent off the annual rate, held at that figure for the whole term.£1,071/yr
3-year term8Twenty per cent off the annual rate, held at that figure for the whole term.£952/yr
Three years for what two years of monthly billing costs

Eight months paid per year, across three years, is 24 months paid for 36 months of service — one year in three carries no charge. On this instrument the three-year term totals £2,856. That is 24 × £119 = £2,856, and two years billed monthly is £1,428 × 2 = £2,856. The same money. It buys three years instead of two, and the arithmetic is on the page so you can check it rather than take it.

Group licences, across entities

Licensing here is per account or per desk, not per seat, so a seat band does not apply and none is offered — applying one would be a category error dressed as a discount. The scaling axis here is entities: where the same instrument is run by more than one legal entity, desk, fund or network inside a group, the licence is negotiated as a single group licence rather than replicated entity by entity. Multi-network clients are exactly the case this exists for, and the commitment ladder above applies to a group licence on the same terms it applies to a single one.

Commercial routes

Pricing scales on entities and on term — one negotiated group licence across desks, funds and legal entities, never a seat band.

Support

  • TierPriority at this instrument’s base contract — ticket, prioritised, first response targeted at one business day. Support tier follows the annual contract value, not the price of a single unit — more seats, a suite licence or a group agreement raise the contract value and can raise the tier with it.

Trial mechanics

The trial runs on the hub tier — thirty days, a full natural proof cycle, disclosed in plain words before you start it and cancellable in one step; none of it is live yet. Trial mechanics in full, by delivery class.

Read before you commit

The documentation is published ahead of the product on purpose — intended behaviour is only a commitment if it exists first. Start with installation and first run, then the limits: the conditions under which this instrument refuses to produce a number are the part worth reading before you pay. The full centre is at /docs/, and the support model states what a ticket does and does not cover.

Questions and answers

Answered from what this instrument publishes about itself. Nothing below is attributed to a customer, because there are none yet.

Will it tell me whether my strategy works?

No, and no statistic can. It bounds the ways in which the evidence for a strategy can be an illusion — that is all, and that is the honest maximum on offer.

Why does the number of trials matter so much?

Because every downstream statistic conditions on it. The auditor accounts for the configurations evaluated explicitly or implicitly before this one was selected, and makes that count unavoidable rather than assumed.

What is a deflated performance statistic?

The reported Sharpe ratio adjusted for the number of trials, the length of the track record and the non-normality of returns — the version of the number a referee would accept.

Can I buy this instrument today?

No. Nothing on this site is on sale — there is no checkout, no card capture, and no product account to create. Every figure on this page is an anticipated indication I have set so it can be read and compared, not an offer, and final pricing awaits my ratification. The launch list is the only thing you can join today.

Change log

NOT YET PUBLISHED

Overfit Auditor has not shipped, so there is nothing to record. When it does, every version lands here — dated, append-only, written by a person, and including the changes that removed a capability rather than added one.

Where this sits

Overfit Auditor is one of the instruments in the Backtest Honesty suite. How that suite measures — the per-instrument battery, and the receipts each measurement will carry — is set out in the Backtest Honesty methodology, part of the site-wide measurement methodology.

Also in the Backtest Honesty suite

Overfit Auditor shares the Backtest Honesty suite with two other instruments.

Research behind this instrument

Overfit Auditor draws on thirteen research notes on this site.

---