# How many trades prove a trading edge?

> There is no universal number: the trades required scale with the square of your edge's dispersion-to-size ratio. How to compute your own, from your own log.

- Canonical: https://hadalinstruments.com/research/how-many-trades-prove-a-trading-edge/
- Published: 2026-08-03
- Author: Hadal Research
- Answers the question: "how many trades to prove a trading edge"
- Coins the term: **The Luck Horizon** — The number of trades below which a record cannot be distinguished from a lucky zero-edge process: not a universal figure but a function of the edge's size relative to its per-trade dispersion — fine edges push the horizon into the thousands, and a record shorter than its horizon is testimony, not evidence.

---
## The short answer

There is no universal number, and everyone quoting one is answering a different strategy's question. The trades required to distinguish an edge from luck scale with the **square** of your edge's dispersion-to-size ratio: coarse edges show themselves in hundreds of trades, fine edges hide for thousands, and the same record that proves one strategy is statistically silent about another. I call the crossover **the luck horizon** — the record length at which your results stop being consistent-with-luck — and the useful fact is that yours is computable, in minutes, from the log you already keep.

The term has an honest lineage worth crediting: poker's working vocabulary has long carried a "luck horizon" — the point where luck's share of results shrinks within tolerance — because poker players confront variance without the comfort of a market to blame. Trading's version differs only in that almost nobody computes it.

  **WHAT THIS IS — AND WHAT IS NOT PUBLISHED.** This article is method — the arithmetic of sample sufficiency applied to a trading record. **It makes no claim about any strategy's profitability, including yours**, and computing a horizon neither blesses nor condemns the edge it measures: it states how much record would constitute evidence. Status of any measured claim: `NOT YET COMPUTED`.

## Prerequisite Knowledge
You need your own trade log with results in risk units — R-multiples, or profit and loss you can divide by the risk taken per trade. Journals export this; a spreadsheet reconstructs it from any statement. Nothing else: the computation below uses only the mean and spread of your own per-trade results.

## What the question actually asks

"Does my strategy work?" is, statistically, "is the average of my per-trade results distinguishable from zero, given how much they vary?" — the oldest question in statistics wearing a trading account. Each trade contributes one noisy observation of your edge. The signal is the mean result per trade; the noise is the spread around it; and the record's power to separate them grows only with the square root of its length. That square root is the whole story: it is why the early curve feels so persuasive and proves so little, and why doubling your evidence requires quadrupling your record.

## The arithmetic, in words

Take your log's mean result per trade, in R. Take the standard deviation of those same results. The ratio of the first to the second is your recorded edge's signal-to-noise per trade — call it the trade quality. Your record's overall evidence is that quality times the square root of the number of trades; the conventional threshold of taking-seriously sits near two. Rearranged, the horizon is: **four divided by the trade quality squared** — the number of trades at which a real edge of your recorded size would typically clear the threshold.

A trade quality of one fifth puts the horizon near a hundred trades. A quality of one twentieth — closer to most honest retail records — puts it near sixteen hundred. Same formula, different lives.

Two warnings the formula carries within itself. It assumes the trades are comparable draws from one process — a strategy revised mid-record restarts its own clock, and correlated positions count as fewer [effective observations](/glossary/effective-sample-size/) than the row count suggests. And it prices only the recorded past: an edge measured over one regime says nothing yet about the next, which is a different and harder question than sufficiency.

## Why the win rate keeps fooling people

The win rate is the most visible statistic a record produces and very nearly the least informative. It ignores what wins and losses are worth, so it cannot see the account-killing shape — frequent small wins, rare large losses — that produces proud win rates and negative expectancies simultaneously. The calculators that put honest intervals on win rates measure the wrong object well. Put the interval on the thing that pays: mean R per trade. That is the number the horizon is computed from, and the number the whole question was secretly about.

## Your own horizon in ten minutes

From your log: mean R, standard deviation of R, their ratio, then four over the ratio squared. Compare the result to your record's length.

The [sample-size check](/tools/sample-size-check/) does that arithmetic on a pasted log and refuses where the log is too short to measure its own spread. If the record is longer, your edge has cleared the ordinary bar of evidence — for the period and process it was recorded under. If it is shorter — the common case — the honest statement is not "the strategy fails" but "the record is not yet evidence", which is a different sentence with different consequences: it prices further testing instead of further conviction. The [risk-of-ruin](/glossary/risk-of-ruin/) entry carries the survival mathematics that decides whether you can afford the remaining distance; the [Sharpe ratio](/glossary/sharpe-ratio/) entry connects this per-trade arithmetic to the annualised statistic the industry quotes.

## The Observable Mechanism
Everything here is computable from your own log in a spreadsheet, and the claim structure is falsifiable at home: simulate a zero-edge process at your trade frequency and watch how often it produces records resembling yours below the horizon. The mathematics is not mine — it is the ordinary theory of sampling error — and the only thing this house adds is the insistence that it be applied before conviction, not after.

## What This Does Not Establish (The Limits)
This article establishes sufficiency — how much record constitutes evidence — not validity of any particular record, and not prediction. A record past its horizon establishes that the RECORDED process had an edge over the RECORDED period; regime dependence, selection among strategies (the multiple-testing problem the [overfitting article](/research/how-to-test-a-backtest-for-overfitting/) treats), and every form of survivorship remain untouched by sample size alone. The horizon is the floor of seriousness, not the ceiling of proof.

## Where this leads

The desk instruments hold this discipline as running machinery: the [Prop-Evaluee Risk Guardian](/instruments/prop-evaluee-risk-guardian/) computes survival arithmetic against live rule sets, and the [Overfit Auditor](/instruments/overfit-auditor/) handles the harder sibling question — what a record means when it was selected from among many attempts. The vocabulary lives in the [lexicon](/lexicon/); the horizon joins it with this article as its receipt.
---

## Claims examined

### Claim 01 — canonical: https://hadalinstruments.com/refutations/#claim-bc82c468

> "I'm up big after three months, so the strategy clearly works." — our reading: Unproven

Three profitable months is a fact about the past and a hope about the mechanism. A zero-edge process produces three-month runs like yours routinely — how routinely is computable, and until that computation is made, the run cannot testify about which process produced it. The uncomfortable arithmetic is that the information in a record grows with the number of trades times the SQUARE of the edge-to-dispersion ratio, so a modest edge hides inside its own noise for far longer than intuition expects. The question is not answerable by looking at the equity curve harder.

**What is true:** A trading record becomes evidence at a computable point — when its length exceeds the luck horizon set by the edge's size relative to its per-trade dispersion — and before that point the honest description of any run, however profitable, is: consistent with luck.

Evidence: https://hadalinstruments.com/research/how-many-trades-prove-a-trading-edge/#claim-bc82c468

### Claim 02 — canonical: https://hadalinstruments.com/refutations/#claim-7d976bbe

> "My win rate is over sixty per cent, so I have an edge." — our reading: Misleading

Two things are wrong at once. Small-sample noise first: a sixty-per-cent reading over thirty trades carries an uncertainty band wide enough to include a losing system, and the calculators that put an interval on a win rate — the good ones use a Wilson score interval — are measuring honestly. But they are measuring the wrong object: a win rate says nothing about what winning and losing are WORTH, and a high win rate with small wins and occasional large losses is the classic shape of a negative edge wearing a flattering statistic. The quantity that pays is expectancy per trade — the mean result in R terms — with its own interval, and no win-rate calculator computes it.

**What is true:** The paying quantity of a trading record is expectancy per trade — the average result in risk units, with its uncertainty interval — and a win rate is meaningful only alongside the sizes of wins and losses, never on its own.

Evidence: https://hadalinstruments.com/research/how-many-trades-prove-a-trading-edge/#claim-7d976bbe

### Claim 03 — canonical: https://hadalinstruments.com/refutations/#claim-3adf6a8c

> "A hundred trades is enough to validate a strategy." — our reading: Misleading

A hundred trades resolves a coarse edge and cannot resolve a fine one, which is why every fixed number in this genre is wrong in both directions. A process winning fifty-five in a hundred with even payoffs shows itself in a few hundred trades; a process winning fifty-two in a hundred — which compounded carefully is a real business — needs thousands before it separates from a coin. The trades required scale with the square of dispersion over edge: halve the edge and you need four times the trades. The right output of a validation question is not a pass mark at some round number but the horizon for YOUR recorded edge — and the honest answer for most retail records is that the horizon has not yet been reached.

**What is true:** The number of trades that proves an edge is a function, not a constant: it grows with the square of the ratio of per-trade dispersion to per-trade edge, so each halving of the edge quadruples the record required — and any fixed validation number is therefore wrong for almost every strategy it is applied to.

Evidence: https://hadalinstruments.com/research/how-many-trades-prove-a-trading-edge/#claim-3adf6a8c

## Cite This Article

APA BibTeX HTML

Hadal Research. (2026). How many trades prove a trading edge?. Hadal Research. https://hadalinstruments.com/research/how-many-trades-prove-a-trading-edge/ Version e472963, 2026-09-14.

@misc{hadal_2026_how-many-trades-prove-a-trading-edge,
author = {Hadal Research},
title = {How many trades prove a trading edge?},
year = {2026},
url = {https://hadalinstruments.com/research/how-many-trades-prove-a-trading-edge/},
howpublished = {Hadal Research},
version = {e472963},
note = {Published: 2026-08-03; version dated 2026-09-14}
}

Source: Hadal Research, How many trades prove a trading edge?. <a href='https://hadalinstruments.com/research/how-many-trades-prove-a-trading-edge/' rel='canonical'>Original Research</a>

Copy Citation

**Version e472963** identifies the commit that last changed this page in Hadal's content repository. That repository is not public, so the identifier does not resolve externally — it is published so a citation pins one specific state rather than a moving page. To obtain the exact version cited, use the [press and research route](https://hadalinstruments.com/press/).

## Explore further

### Instruments

- [Overfit Auditor](https://hadalinstruments.com/instruments/overfit-auditor/)
- [Prop-Evaluee Risk Guardian](https://hadalinstruments.com/instruments/prop-evaluee-risk-guardian/)

### Concepts

- [Effective Sample Size](https://hadalinstruments.com/glossary/effective-sample-size/)
- [Risk of Ruin](https://hadalinstruments.com/glossary/risk-of-ruin/)
- [Sharpe Ratio](https://hadalinstruments.com/glossary/sharpe-ratio/)

### Research

- [How do I test a backtest for overfitting?](https://hadalinstruments.com/research/how-to-test-a-backtest-for-overfitting/) Asked as: how to test a backtest for overfitting

[All Hadal research](https://hadalinstruments.com/research/)[This article as plain markdown](https://hadalinstruments.com/research/how-many-trades-prove-a-trading-edge.md)

---

## Raw artifact — NOT PUBLISHED FOR THIS PAGE

No downloadable artifact ships with this page. Eight published measurements do, each content-hashed so a reader can verify the figures independently. Where a measurement is published here without one, that is a gap rather than a policy, and it is stated rather than left to be noticed.

[Measurements that ship their data](https://hadalinstruments.com/research/)
