Does COT positioning predict reversals?
Asked as: does cot positioning predict reversals
Crowded positioning extremes precede fewer liquidations, not more. The interval excludes zero, p = 0.0066, and the sign is backwards.
No. On my data the relationship runs the other way, and the interesting part is that it runs the other way significantly.
The AUD leg returned a trigger effect of −0.0477, bootstrap interval [−0.0801, −0.0137], p = 0.0066. The interval excludes zero. The sign says crowded positioning extremes were followed by liquidation events at 3.1% against the mid-range controls — less often, not more.
I call that shape Wrong-Way Significance: an effect that clears every bar a reader checks and points at the opposite of the hypothesis it was registered to test.
Why this is the first thing I publish with its data attached
Every measurement on this site is held to three standards: the methodology is pre-registered before the first tick, the artifacts are content-hashed, and the results are recomputable. This is the first page where all three are true at once, which is why it goes first rather than something with a better headline.
Pre-registered, and the ordering is provable rather than asserted. The registration records that at the moment it was written the positioning data had never been contacted — the directory did not exist, and the single prior fetch attempt had returned an HTTP 503 and written nothing. The fetch log agrees: every attempt on that date failed. So the hypothesis, its thresholds, its horizon grid and its decision tree were fixed at a point when looking was not possible.
That matters more than it sounds. A pre-registration is worth exactly the provability of its ordering, and most are worth nothing on that test.
Content-hashed. The dataset below carries nineteen input hashes, one per source CSV, and every one of those CSVs is committed. The compute step opens no socket.
Recomputable — and this was measured rather than assumed. The claim most sites would make here is “you could re-derive this”. I re-ran the driver into a clean temporary directory and compared the output field by field against the published artifact: fourteen of fourteen result keys byte-identical, none differing. So you can download the file, check its hash against the one published here, and re-derive it from the committed inputs — and I have already done that myself and reported what came back. Nothing in the chain asks you to trust me.
What was actually measured
Seven currency books, 201 aligned weekly release stamps, latest 2026-06-30, drawn from the CFTC’s Traders-in-Financial-Futures report. Leveraged-fund net positioning is joined to each leg’s weekly price and keyed to the Friday release stamp, giving 821 release-keyed weekly observations per leg after the 2010 price join.
The frozen defaults: a 156-week percentile window, a 2× range threshold, deciles at 0.9 and 0.1, horizons of 2, 4 and 6 weeks, a 1,000-iteration block-bootstrap surrogate, four folds.
The surrogate deserves its own line, because it is where most tests of this kind are weakest. It resamples the weekly changes in a stationary block bootstrap, relabels the extremes on the surrogate while holding the real outcomes fixed, and re-runs the test. An independent-and-identically-distributed shuffle would have been easier and would have been anti-conservative — positioning is autocorrelated, and destroying that structure inflates significance. I had learned that on an earlier hypothesis and wrote the harder surrogate into this registration before running it.
The result, leg by leg
AUD: not live. Effect −0.0477, interval [−0.0801, −0.0137], p = 0.0066, and it fails the relabel surrogate. Windowing is durable. The NZD leg was registered as a uniformity control rather than a target, and the AUD effect survives it — the difference between the two legs is −0.092, so what is being seen is AUD-specific rather than generic antipodean risk flow.
JPY: a preliminary unconditioned null. Effect +0.0442, interval [−0.0152, +0.1131], p = 0.18. This is the unconditioned result, and the registered verdict for JPY requires an intervention-and-rate regime split whose feed does not exist yet. So it is reported as preliminary and not as a finding, which is the state it is actually in.
Verdict, from the tree written in advance: the single-leg positioning fuel gauge does not carry the crowding-to-liquidation signal. That was the outcome the registration declared it expected, which is the least exciting way for a pre-registered test to resolve and the most credible.
What this does not establish
The null gates rather than kills. The registered thesis is a conjunction — crowding and dealer positioning and thin liquidity — and this measures the first term alone. A null here means Layer 1 does not on its own justify the spend on Layers 2 and 3. It does not mean the conjunction is refuted, and I would have had to say that either way, because the tree deciding it was fixed first.
The JPY leg is preliminary for the reason above. The AUD result is one currency on one weekly cadence. And positioning data is a weekly position snapshot, not a price series — the artifact says so on its own face, in the banner it ships with.
The download
The measurement is the file, not this page. Its hash is published beside it; if your copy hashes differently, one of me has a problem worth knowing about.
If you recompute it and disagree with me, send the recomputation. A correction with a receipt attached gets published as a correction, with the original left legible beside it.
The artifact
SHA256: 5dd9ef789a81d9889935189729a666ccdbef7ae4faf6d3d4f264f25bc577760a
Download datasetClaims examined
Claim 01§ claim-89e77e34
When the leveraged funds are all on one side, the reversal is coming.
The AUD leg returned a trigger effect of −0.0477 with a bootstrap interval of [−0.0801, −0.0137] and p = 0.0066. That interval excludes zero, which is normally where a reader stops. The sign is the problem: it is negative, meaning crowded extremes were followed by liquidation at 3.1% against the mid-range controls rather than more often. It then failed the block-bootstrap relabel surrogate that was frozen before the run, so it is not a live signal in either direction. A significant effect pointing the wrong way is not a weak version of the belief. It is evidence against it.
Claim 02§ claim-393e90f4
You just didn't have enough data.
The objection is fair against most nulls and weak here, because the parameters could not have been tuned to produce this outcome — they were frozen before the data existed. The defaults were a 156-week percentile, a 2× range threshold, deciles at 0.9 and 0.1, a horizon grid of 2/4/6 weeks, a 1,000-iteration block-bootstrap surrogate and four folds. What the sample size does bound is the JPY leg, whose registered verdict needs an intervention-and-rate regime split that has no feed yet — so JPY is reported as a preliminary unconditioned null rather than as a finding.
Claim 03§ claim-d6615f4e
So the whole positioning thesis is dead.
The registered thesis is a conjunction — crowding, dealer positioning and thin liquidity together — and this measures the first of those alone. My own verdict language is that it gates rather than kills: Layer 1 on its own does not justify the spend on Layers 2 and 3. Reading a one-layer null as a three-layer refutation would be the same error as reading a one-layer hit as confirmation, and I would have had to say so either way, because the tree that decides it was written down first.
Each claim above has a permanent address — the § link — whose canonical home is the refutation index, where it carries its variant phrasings and the true proposition stated on its own feet; this article is the evidence behind it. If a claim's text ever changes, it becomes a new claim at a new address, and the old one stops resolving rather than silently meaning something else.
Explore further
Instruments
Research
- The break that cleared neither barAsked as:
how do you know a detected regime change is real
- Is my volatility regime just telling me the time?Asked as:
is my volatility indicator just measuring time of day
- How do you verify a trading track record?Asked as:
how do you verify a trading track record