The edge that was real and worth half a pip
Asked as: can a trading edge be statistically significant but not profitable
My engine's first end-to-end edge hunt found statistically real structure — leans of half to two percent from a coin flip — and none of it is tradable.
Ask whether my engine has found a trading edge, and the honest answer from its first end-to-end hunt is: it found structure that is statistically real and economically negligible — leans of half a percent to two percent from a coin flip, undeniable at fifty thousand observations and worth a fraction of a pip before a spread is paid. I call the shape Significance Without Size: a p-value that is small because the sample is large, sitting beside an effect that cannot pay for its own execution.
This is a post-mortem rather than a study of somebody else’s idea. The hunt was built, run, adversarially checked and reported in a single day, 12 August 2026, and the report it produced is blunt about its own design errors. This note is that report, translated for a reader who does not have the engine.
WHAT THIS IS — AND WHAT IS NOT PUBLISHED. Every figure below is transcribed from the engine’s own handover document of 12 August 2026. The result artifacts — the lattice, the walk-forward ledger, the cell table — are not yet published, so no digest for them appears here and this note carries no dataset block by design: a provenance claim without its artifact is the thing the registers exist to refuse. No strategy is proposed, no pair is recommended, and no figure here is a forecast. Status of the artifacts: NOT YET PUBLISHED.
Prerequisite Knowledge
Two ideas. A first-passage question asks, from a bar’s close, whether price touches a level above before a level below within a fixed number of bars — a direction question with a decided answer, an undecided answer, or no answer. And a cost unit: the engine sizes its targets in multiples of a spread proxy rather than in pips, so that “a target of three to eight units” means the same thing on a wide pair and a tight one. Nothing else is needed to follow the arithmetic; the statistics are the ordinary ones.
What the hunt asked
The atom is one entry bar. From its close, does price touch plus T before minus S within H bars? The engine records four outcomes and never three: up first, down first, both barriers inside one bar — where the order is unknowable at bar grain and is never assumed — and no touch at all, the appetite that did not pay. A walk the tape cannot resolve is excluded and counted.
A cell is a pair, a target band and a volatility state: bands of three to eight, eight to twenty and twenty to fifty cost units, and the regime label frozen at the moment of entry, point in time. Every hourly stamp is an entry. The cell’s statistic is the probability of up-first among the decided walks, with the no-touch and ambiguous fractions printed beside it and never folded into its denominator. Over fifteen and a half years and twenty-eight pairs that gave three hundred lit cells.
Confirmation was forward and blind. Flag dates were simulated every ninety days through the history; the discovery pass at each date saw nothing after it; a mechanical rule flagged a cell when its deviation from a coin flip, scaled by the square root of its sample, reached one, with a floor of a hundred decided walks; and confirmation was measured only on the ninety days that followed. For every flag the ledger records whether the lean kept its side, and how much of it survived.
What it found
Five thousand four hundred and forty-nine flags, of which two thousand nine hundred and eighteen could be confirmed. Broad mechanical flagging fails: unique-cell persistence came out at 0.330, which is to say the typical flagged anomaly mean-reverted once it had been flagged.
A concentrated set survives. Twenty-eight of the eighty-five cells that could be ranked cleared a nominal p of 0.05, and twelve survived a Bonferroni correction across the family. Twenty-two of the twenty-eight were yen and franc crosses in the calm volatility state — one coherent phenomenon, which the report reads as funding and haven drift, rather than twenty-eight discoveries.
That is the part a marketing page would stop at. Twelve cells, Bonferroni-clean, fifteen years, walk-forward. The report did not stop there.
The central finding: put the effect size beside the p-value
The persistence statistic measures the direction of a lean, not its size. When the effect sizes are placed beside the p-values, the picture inverts:
| Cell | Up-first probability | Decided walks | Target | Gross expectancy per decided walk |
|---|---|---|---|---|
| CADJPY, band 1, calm | 0.503 | 57,724 | 31 pips | +0.2 pips |
| GBPCHF, band 1, calm | 0.490 | 69,194 | 26 pips | +0.5 pips |
| AUDJPY, band 2, calm | 0.487 | 23,168 | 62 pips | +1.6 pips |
| GBPAUD, band 2, calm | 0.518 | 28,684 | 119 pips | +4.3 pips |
Deviations of half a percent to two percent from a coin flip. The p-values are tiny because the samples are fifty thousand and more, not because the effect is large. Half a pip of gross expectancy does not survive a spread, let alone a commission. The report’s own sentence for this is the one I would carve above the door: a p-value without an effect size beside it is a dangerous instrument — and the engine’s edge report now refuses to print one without the other.
Four reasons the design could not have found tradable edge
The report lists them in order of damage, and the first is the one it calls the biggest single design error.
Symmetric barriers force a coin-flip baseline. With the target equal to the stop, a market that is close to a fair game gives up-first about half the time by construction. The design can therefore only ever detect tiny deviations from a half — and it has converted a question about magnitude into a question about direction, which is the one question this corpus has already answered “no” to at every resolution the engine has measured.
A touch probability is not a profit and loss. Whether price reaches one barrier before another says nothing about how far it travels, how long the capital is tied up, or what the paths that never touch were doing meanwhile. Two cells with the same up-first probability can have opposite expectancies once path and time are priced.
No cost model was armed. A cost-floor law exists in the engine — a target must be at least three times the loaded cost — and it renders unarmed, because the commission line item does not yet exist. Everything in the table above is pre-cost, and at these effect sizes cost is not a haircut. It is the verdict.
The conditioning was shallow. One state family, one horizon of twenty-four bars, one grain at a time, while the engine holds around twenty other measured conditioners that have never been crossed with the lattice.
The part that makes the null worth keeping
A null result earns its keep by having been able to come out the other way, and this one was. The machinery flagged, confirmed forward, corrected for the family, and then declined the broad claim it could easily have made from twelve Bonferroni-clean cells. Any graded claim in the engine goes through a registered-trial ledger — one pre-declared endpoint per hypothesis, a draw budget derived rather than chosen, a jointly drawn null, a negative control — and the two hypotheses this hunt drafted are waiting on that ledger rather than being announced.
The report’s diagnosis fits in one paragraph, and I will give it in the report’s own shape rather than soften it. The lattice asked which barrier gets touched first at equal distance — a direction question with a mechanically even baseline — on a market where the engine has already shown direction to be unforecastable at every resolution it has measured. It validated the answer with a statistic that is blind to magnitude, on samples large enough to make a half-percent deviation look like a one-in-ten-billion discovery, with no cost model armed. The machinery is sound: it found real structure, refused the broad claim, and reported honestly. It was pointed at the one question the corpus keeps answering no.
The Observable Mechanism
The arithmetic that condemns the finding is checkable in anyone’s own data at spreadsheet scale, without the engine. Take any pattern’s win rate, the sample it came from, and the size of a typical win: the lean from a coin flip times the target is the gross expectancy per attempt, and the spread is the first thing it must clear. If the lean is half a percent and the target is thirty pips, the attempt is worth a fifth of a pip before anything is paid — and that sentence is true at any sample size. The engine’s run itself is not yet published; the reasoning that rejects it is.
What This Does Not Establish (The Limits)
This note establishes what one design, on one corpus, found and why the finding is not an edge.
It does not establish that no edge exists in these markets; it does not establish that the yen-and-franc calm-state lean is worthless under a different question — asymmetric barriers, expectancy in place of touch probability, a cost model armed, all of which sit on the engine’s own forward path; and it establishes nothing about any pair beyond the recorded period. The artifacts are not published, so a third party cannot yet recompute any number on this page; until they are, the correct weight to give the figures is the weight of a stated record, not of a verified one. Nothing here is trading advice, and no pair, band or state is a recommendation.
Where this leads
The report’s forward path is engine work, listed in order of expected value and none of it a claim: barriers made asymmetric, so the question becomes payoff geometry rather than direction; barriers scaled from the one quantity the engine can actually forecast, which is volatility; the cost floor armed, so the lattice can refuse an untradable cell by itself; the conditioners crossed instead of taken one at a time; expectancy and the full excursion distribution in place of a touch probability; the relative-value tier run through the same forward pipeline, because it is non-directional by construction; and only then a grade, through the registered-trial machinery that is built and planted-truth tested.
Two disciplines on this site exist because of exactly this shape of result. The Overfit Auditor bounds what a record means when it was selected from among many attempts — which is what eighty-five ranked cells are. The Epistemic Harness keeps the trial count honest while the search is still running, which is the number every correction above conditions on. And the Reproducible Verdict Kernel is the reason the caveat at the top of this note is temporary: when the hunt’s artifacts publish, their digests join the registers, this note gains its dataset block, and the numbers above become recomputable rather than stated. Until then the kill ledger is where the engine’s dead hypotheses are counted, and a half-pip edge is, for now, one more entry waiting to be written.
Claims examined
Claim 01§ claim-cf685939
If a pattern is statistically significant, the edge is real and I can trade it.
This is the finding of the hunt below, stated as the belief it corrects. Twenty-eight of eighty-five cells cleared p < 0.05 and twelve survived a Bonferroni correction across the family, on samples of twenty to seventy thousand decided walks each. Put the effect sizes beside them and the flagship survivor is a first-touch probability of 0.503 against a baseline of 0.5 — worth about a fifth of a pip per decided walk, gross, on a pair whose spread is many times that. The significance was real. The edge was not, because the quantity that pays was never in the statistic.
Claim 02§ claim-eaacc2f2
More data makes a small edge safer to trade.
The hunt ran on fifteen and a half years of hourly entries across twenty-eight pairs, which is why its p-values reached the order of one in ten billion: the interval around each lean had narrowed to almost nothing. What the sample cannot do is move the lean itself. A deviation of half a percent from a coin flip is half a percent whether it is measured on fifty walks or fifty thousand. The large sample converts 'probably noise' into 'certainly real and certainly negligible', which is a different sentence and not a better trade.
Claim 03§ claim-faf19b2f
A pattern that persisted in the past will keep paying.
The confirmation stage flagged cells every ninety days through the history, blind to what came after, and then asked whether each flag kept its side over the following ninety days. Across the broad set it did not: unique-cell persistence came out at 0.330, meaning the typical flagged anomaly leaned the other way after it was flagged. The exception was concentrated and coherent — twenty-two of the twenty-eight surviving cells were yen and franc crosses in the calm volatility state, one phenomenon rather than twenty-eight discoveries — and even that survivor is the half-pip lean above.
Each claim above has a permanent address — the § link — whose canonical home is the refutation index, where it carries its variant phrasings and the true proposition stated on its own feet; this article is the evidence behind it. If a claim's text ever changes, it becomes a new claim at a new address, and the old one stops resolving rather than silently meaning something else.
Explore further
Concepts
Research
- Does a wick rejection mean anything?Asked as:
does a wick rejection actually mean anything
- The break that cleared neither barAsked as:
how do you know a detected regime change is real
- How do you verify a trading track record?Asked as:
how do you verify a trading track record