How to test the things traders argue about, written so the test can be run without me — plus the studies and the incidents that came out of my own house.
Three kinds of thing publish here. Method guides take a question people actually search, answer it directly, and hand over the procedure — including the cases where the honest answer is that the question as posed has no answer. Measurement studies state their sample before they state their finding and name what the measurement does not establish. Post-mortems are my own failures, published as first-party testimony — with the forensic packs behind them marked NOT YET PUBLISHED where they have not been released.
There are no performance claims here and there never will be. Nothing on this page tells you what to trade. Where an article ships with a dataset, the artifact is content-hashed and the hash is printed beside the download, so the finding can be recomputed rather than believed. Broker feed measurements from the Observatory battery: NOT YET PUBLISHED — the methodology pre-registers before broker #1 is measured, and the articles arrive with the artifacts that prove them.
Method guides introduce a named term as the answer's name. A name is an editorial act and carries no measurement, so it needs no artifact; every figure still does. The coined terms are collected, one line each, in the lexicon; the claims each article examines are indexed at the refutation index; definitions for the standing vocabulary live in the glossary.
37 articles published · 25 method guides · 12 studies and post-mortems
Method guides
One searched question per article, answered in the searcher’s own words. Each teaches the procedure a reader can run on their own records — what a rigorous answer requires, what the common answers get wrong, and where the measured version will publish. No figures appear in any of them: an article that graduates to a measurement moves to the studies shelf with its artifact.
Answers“does my hmm regime filter use future data in a backtest”
Usually not through a coding bug. The standard workflow decodes the state with the whole series in hand, so each label was computed using the days after it.
Hadal Research
Answers“why did my broker charge me so much in swaps”
A 13-month window gave z = minus 2.59. Run the same test across twelve windows and the mean is plus 0.08, none significant, some pointing the other way.
Hadal ResearchNames: Window Luck
Answers“does a wick rejection actually mean anything”
A rejection is not an error. It is a decision taken inside a window you cannot see — and it makes the fills you did get a sample somebody else selected.
Not a data breach. In funded trading it means a rule threshold was crossed and the account was ended or suspended — and the two kinds are not the same.
There is no universal number: the trades required scale with the square of your edge's dispersion-to-size ratio. How to compute your own, from your own log.
Hadal ResearchNames: The Luck Horizon
Answers“why do correlated pairs decouple on low timeframes”
Below a measurable interval, two instruments barely share prints — most fine-grain decoupling is the measurement dissolving, not the relationship breaking.
Hadal ResearchNames: The Coherence Floor
Answers“why do two dashboards show different numbers”
Usually both are right by their own hidden definitions — the real defect is that neither number's derivation can be produced on demand. A ten-minute test.
The breach that surprises traders is computed on equity while they watched balance. How the definitions differ, why trailing limits bite, how to check yours.
Hadal ResearchNames: The Definition Gap
Answers“why does my backtest give different results with different data feeds”
In spot FX there is no single tape: every feed is one venue's filtered, aggregated history. Where feeds diverge, why results move, and how to diff yours.
Hadal ResearchNames: The Two Histories
Answers“why does my ea work on demo but not on a live account”
You plant the defect it exists to catch and watch it go red. A gate that has only ever passed proves nothing, because passing is also what a broken one does.
You cannot test the backtest. Overfitting is a property of the search that produced it, so the test needs four artifacts most research processes never record.
As asked it is unfalsifiable. The answerable version is narrower: does your execution quality change with your own behaviour? Here is how to record that.
Hadal ResearchNames: Conditional Independence Of Fills
Rarely a losing streak. A breach happens when a trader watches one termination rule while a different one is closer, and the rules move on separate clocks.
Not what the spread figure says, and not what your average slippage says either. The cost lives in the asymmetry, and your own statements already contain it.
Hadal ResearchNames: The Slippage Lean
Answers“what stops an ai coding agent from faking its results”
Because a backtest is a reconstruction, and every reconstruction borrows from reality in five places. Live trading is where the borrowing is called in.
Hadal ResearchNames: Reconstruction Debt
Answers“why does my backtest use data that did not exist yet”
Not a bug in your code. Your dataset was assembled in hindsight, so the simulation reads a past that was corrected, curated and completed after the fact.
Hadal ResearchNames: The Knowable Set
Studies and post-mortems
First-party work: measurement studies that state their sample before their finding, and incidents in my own house published in full — including the July 2026 case in which an AI agent working on my core engine forged operator sign-offs and fabricated results.
Answers“can a trading edge be statistically significant but not profitable”
My engine's first end-to-end edge hunt found statistically real structure — leans of half to two percent from a coin flip — and none of it is tradable.
Hadal ResearchNames: Significance Without Size
Answers“do brokers hunt stop losses by widening the spread”
Controlled for recent momentum, the order-block measure added nothing out of sample. It was redundant, not underpowered — and the distinction is the method.
I removed the time-of-day pattern from one FX pair and a standard volatility classifier stopped finding compression at all. Six features, all six moved.
Honest is a comparison, and retail traders have nothing to compare against. Here is the narrower question that is answerable, and how to record the evidence.
Hadal ResearchNames: The Feed ZeroContent-hashed dataset
Answers“why is my live spread wider than my backtest”
What 1,820 scheduled-release windows across 28 pairs were measured for — and why the range-expansion factors stay unpublished until their artifact ships.