Do brokers hunt your stop losses?

Asked as: do brokers hunt stop losses by widening the spread

Up-side spread tags run 67% at round numbers against a modelled floor of 77%. The claim that stops are hit more often than chance does not survive it.

Not by the mechanism the story describes. I measured the exact move — the ask crossing a level while the bid never follows, stops filling, price snapping back — directly on the bid and ask streams. At round-number levels it occurs about 67% of the time. A random walk carrying the same spread produces 77%.

The move happens less often than chance. That is not a weak effect; it is an observation pointing the other way.

The number that makes this readable is the one almost nobody computes. I call it The Chance Floor: the rate a pattern reaches by construction, before anyone does anything. Until you know the floor, a raw percentage tells you nothing at all.

Why 67% sounds damning and isn’t

Read alone, “two thirds of level tags are one-sided in the direction that would hit buy stops” is exactly what the thesis predicts. It is the shape of a finding.

But a bid and an ask straddle a level. As price approaches, the ask reaches it first, simply because the ask is higher. One-sided tags are therefore the default outcome of ordinary random movement near a level — not a deviation from it. Once you compute how often a spread-carrying random walk produces them, the floor lands at 77%, and 67% is ten points beneath it.

The same test at prior-day and prior-session extremes gave 69% against a floor of 78%. Both classes of level, both below chance.

Two supporting measurements point the same way. The spread at the moment of a tag was 0.4 pips against 0.3 pips ordinarily — a difference far too small to reach a stop it would not otherwise have reached. And the tagged levels produced no excess reversal afterwards, which the story requires: a manufactured tag is supposed to be followed by the snap back that makes it profitable.

Where the 77% comes from, and what it is not

A number read against a baseline is only as good as the reader’s ability to check the baseline, so here is exactly what the floor is.

It is a surrogate, not a model. I did not simulate a random walk with an assumed volatility. I took the real tick stream, kept the actual increments between consecutive mid-prices, shuffled their order, and summed them back up into a synthetic price path. The real spreads were shuffled separately and re-attached, so spread became independent of where price happened to be. Then the identical tag-classification ran over the synthetic stream.

What that preserves. The exact empirical distribution of tick moves — every fat tail, every jump, the true typical step size — and the exact distribution of spreads. This matters, because the obvious objection to a random-walk floor is that real FX has fatter tails and messier variance than a Gaussian walk. That objection does not apply here. The surrogate’s moves are the real moves.

What it destroys, which is the honest limit. Shuffling removes temporal ordering, and with it every form of serial dependence: volatility clustering, session structure, and any tendency of price to behave differently near a level than away from it. So the surrogate approaches levels with the right step sizes but the wrong arrival pattern. If real approaches to structural levels cluster in time — and there is good reason to think they do — the floor’s exact height is affected. I do not know the sign of that bias without measuring it, and I did not measure it, so I am not going to claim the floor is conservative.

The ten-point gap is an open question, not a finding. The observed rate came in below the floor, and a deficit in the unexpected direction is exactly the kind of result that gets narrativised into meaning something. I am not going to do that. Two explanations remain unseparated by this test: the floor may sit too high because shuffling changed the arrival process, or one-sided tags may genuinely be suppressed near structural levels for reasons having nothing to do with anyone hunting anything. Reporting the gap as evidence for the second would be the same error this article exists to refuse.

What would settle it. The Lo-MacKinlay variance ratio with the heteroskedasticity-robust statistic measures precisely the serial dependence the shuffle destroys, and separates genuine return memory from time-varying volatility rather than confounding them. It is not run here, and until it is, the floor’s exact height is a construction choice rather than a measurement.

What survives all of it. The claim as made in the wild requires the rate to exceed what ordinary movement produces. It did not exceed it on either class of level. That reading needs the floor to be roughly right, not exactly right.

The part where my own measurement was blind

This section matters more than the result, because the first version of this test could not have found the thing it was looking for.

I ran it on H1 bars. On a bid-based feed the bar high is the maximum bid, so an up-reach is classified as genuine — both sides crossed — by construction. The one-sided tag, where the ask pokes through and the bid never follows, can never become an event at all. The audit was structurally incapable of seeing the move in the direction the thesis pointed.

I wrote that down as a failure rather than publishing the null it produced. The honest statement at that moment was not “no stop hunting”; it was “this test cannot answer this question”, and those are different sentences.

Two consequences worth carrying. First, every earlier measurement I had made on this data — the durability kill, the engine walls, the whole “thin” verdict — had been computed on bid-only price. That is fine wherever bid and ask move together and blind wherever they do not. Second, if you are testing stop hunting on candles, you are testing something else. The candle cannot contain the evidence.

What I committed to before running it

Before the rebuild I wrote down which outcome would count as which, because a test whose interpretation is decided afterwards is not a test.

If one-sided tags turned out to be a large fraction of what I had been scoring as sweeps, and reclassifying them revived a real surrogate-controlled signal, then the engine had been blurred rather than thin — the measurement had been contaminated, the thesis was alive, and the whole investigation earned a clean re-run.

If they were a minority, or reclassifying changed nothing, then thin was confirmed and the investigation was closed.

I also recorded the honest base case at the time: thin. The result matched it, which is the least interesting way for a pre-committed fork to resolve and the most credible.

Limits

One account, one venue class, ninety days of tick data, and one currency pair — AUDJPY. A result on one account is a result about one account, and a result on one pair is not a result about the market.

The round-number test ran over a fifty-pip grid, which gave it fourteen levels inside the window. That is a small number of levels, and it is the first thing I would attack if someone else published this. The tag rate is measured over every approach to those levels rather than over the levels themselves, so the sample is larger than fourteen — but the level count is the honest denominator for how much structural variety was covered, and it is not much.

The venue matters more than usual here. Spread behaviour is a property of how a particular broker prices, so this is not transferable to a venue with a different pricing model, and it is not a statement about any firm but the one measured.

And the scope is narrow by design. This tests one mechanism — spread-manufactured tags at structural levels. It says nothing about last-look rejection, slippage asymmetry, fill ratios, requotes or quote staleness. Those are real phenomena, they are separately measurable, and a null here is not a null there. The legitimate measurable form of the wider question is order-type conformance — whether stops and limits execute where the documentation says, compared against an independent reference at the same timestamps — and that is a different instrument from this one.

What this does and does not answer about your fill

My own earlier work on feed honesty concluded that whether a particular widening was aimed at you is not measurable — whose hand moved the spread is not carried in any data a client can hold. That conclusion stands, and nothing here overturns it.

What it identified as the missing piece was the base rate: how often a feed widens to that degree when you hold no position at all, which almost nobody records. This measurement is that base rate, computed at population scale rather than around one fill. So the two results compose rather than conflict: a single episode remains unattributable, and the population signature the attribution would require turns out to be absent.

What to do with a stop that got hit

The useful move is not to decide whether you were targeted. It is to record what the spread was doing at the moment of the fill, because that single number separates the two explanations and almost nobody captures it.

If the spread at your fill was ordinary, the level was reached. If it was genuinely anomalous, that is a measurement about your feed, and it is worth having whatever the cause turns out to be — the same instrument answers both questions, and only one of them is about intent.

The artifact

SHA256: 4ceb32682877a4d9bb78af1db23dbab9f8245ed54335450cace3e90e01079369

Download dataset

Claims examined

Claim 01§ claim-4b47266f

My broker widened the spread to take out my stop, then price came straight back.

My reading: Unproven

The move has a precise signature: the ask crosses the level while the bid never does, stops fill, price returns. That is testable on the quote stream. At round-number levels, one-sided up tags run 67% — but a random walk carrying the same spread produces 77%, so the observed rate is ten points BELOW chance. At prior-day and prior-session extremes it is 69% against a floor of 78%. The spread at those tags measured 0.4 pips against 0.3 pips ordinarily, a difference too small to fill anything it would not otherwise have filled, and the tagged levels showed no excess reversal afterwards. A move that happens less often than chance is not a move.

Claim 02§ claim-a1038cd3

The stats might be fine, but the measurement can't see what the broker actually did.

My reading: True

The objection is right, and it caught me. My first audit ran on H1 bars, where the bar high is the maximum bid. That construction makes every up-reach 'genuine' by definition — the exact move in question, where the ask pokes through and the bid never follows, can never become an event at all. I had run a test that was blind in precisely the direction the thesis pointed, and I recorded it as such rather than reporting the null. The result above comes from the rebuild: pools detected and reaches classified on the bid/ask stream directly, never anchored to bid bars. If you are testing this yourself on candles, you are testing something else.

Claim 03§ claim-aa8b9e0a

So brokers never act against their clients.

My reading: False

A null on the spread-weapon at structural levels is not a character reference. Execution quality is a large surface and this measures one corner of it. Last-look behaviour, slippage asymmetry between winning and losing fills, fill ratios, requote rates and quote staleness are all real, all separately measurable, and all untouched by this result. The honest position is narrow: the specific move that retail lore describes, at the specific levels it describes, did not occur above chance on the account measured. Everything else about your broker remains an open question, and a measurable one.

Each claim above has a permanent address — the § link — whose canonical home is the refutation index, where it carries its variant phrasings and the true proposition stated on its own feet; this article is the evidence behind it. If a claim's text ever changes, it becomes a new claim at a new address, and the old one stops resolving rather than silently meaning something else.

Cite This Article

Hadal Research. (2026). Do brokers hunt your stop losses?. Hadal Research. https://hadalinstruments.com/research/do-brokers-hunt-your-stop-losses/ (SHA-256: 4ceb32682877a4d9bb78af1db23dbab9f8245ed54335450cace3e90e01079369) Version 5762f56, 2026-08-29.

Version 5762f56 identifies the commit that last changed this page in Hadal's content repository. That repository is not public, so the identifier does not resolve externally — it is published so a citation pins one specific state rather than a moving page. To obtain the exact version cited, use the press and research route.

---