Receipt page · Desk Discipline
How the Desk Discipline suite measures
This is a receipt page. Dossier §8.2 fixes what one contains, and the five sections below are that anatomy in order: the versioned methodology, the pre-registrations, the proof corpus, the changelog, and the invitation to recompute any published number without asking me for anything. Nothing has been measured yet. Three of the five sections are therefore empty, and they are empty in a way you can inspect — the structure is printed, the slots are named, and each one says NOT YET PUBLISHED rather than showing you something plausible. Looking for what this family will contain instead? That is the Desk Discipline suite hub.
Receipt anatomy · dossier §8.2Desk Discipline suiteReceipt version v1.0
- Instruments in scope
- 3
- Metrics defined
- 12
- Pre-registrations posted
- NOT YET PUBLISHED
- Artifacts published
- NOT YET PUBLISHED
- Measurements published
- NOT YET PUBLISHED
- 01Methodology, versionedProse and formal definitions. Every metric carries its name, its definition, the estimator that computes it, its known failure modes, and what it does not establish. The page carries a version, and the version is printed on it.DEFINED · v1.0
- 02Pre-registrationsDeclarations made and dated before any measurement runs: decision thresholds, exclusion rules, completeness bars. Immutable once posted — an amendment appends beneath the original with its reason, and never overwrites it.NOT YET PUBLISHED
- 03Proof corpusThe published artifacts themselves: dataset descriptors, content hashes, download links, and the version tag of the code that generated each one. A result whose inputs are not downloadable is not a receipt.NOT YET PUBLISHED
- 04ChangelogAppend-only, dated, written by a person. What changed on this page, when, and why. A correction is a new entry naming the entry it corrects — entries are never edited away.1 ENTRY
- 05Independent recomputationThe exact numbered steps a stranger follows to re-derive a published number without asking me for anything. Where the recomputation disagrees, the disagreement is the finding.PROCEDURE STATED
Methodology, versioned
version v1.0last changed 2026-08-0112 metrics · 3 instruments
The Desk Discipline suite measures the operational envelope of a live trading desk: distance to the rules that end accounts, the behaviour of the current book under replayed historical shocks, and whether the execution path itself leaks time and money by more than noise can explain.
Accounts are ended by arithmetic and structure at least as often as by bad ideas — a daily-loss limit crossed by one oversized position, a correlated cluster held as if it were five separate risks, fills that are systematically worse than quotes. This suite instruments those failure modes directly, and has no opinion about anyone’s entries.
Every metric below carries five fields, because four of them are the fields a methodology usually leaves out. Definition states what the quantity is, closely enough that a stranger could implement it. Estimator states how it is computed and which parameters must be fixed before the data is read. Known failure modes states where the estimator breaks — written now, while nothing depends on it, rather than conceded later under questioning.Does not establish states the claim the metric will not support, however natural the reading. The last field is the one that costs something to publish, which is why it is published.
Measurement of structural latency disadvantage in an execution path — whether fills are worse than quotes by more than noise can explain. The instrument is specified; it has not been run against any published sample.
P22.1Quote-to-fill drift
- Definition
- The signed price movement between the moment an order was committed and the moment it filled, kept as a distribution rather than recalled as an anecdote.
- Estimator
- Per-order signed difference between the quote at commit and the fill price, accumulated into a distribution and tested for asymmetry against a symmetric null — random slippage is symmetric, structural disadvantage is not.
- Known failure modes
- Commit timestamps come from the subscriber’s platform and fill timestamps from the broker. Unless clock offset is measured, the drift silently contains it, so offset estimation is part of the procedure rather than an assumption inside it.
- Does not establish
- Asymmetric drift does not establish front-running. It establishes that fills in this sample were worse than quotes by more than a symmetric null explains.
P22.2Adverse-selection signature
- Definition
- Whether the market moves systematically against the subscriber in the interval around their own orders, conditioned on order type, session and size.
- Estimator
- Conditional distribution of post-order price movement over a declared window, computed per condition cell, with every cell publishing its own sample size.
- Known failure modes
- Conditioning on many cells with a fixed sample makes some cells tiny. Cells below the registered completeness bar publish the honest null instead of a number that looks like a finding.
- Does not establish
- A signature does not identify a counterparty and does not establish that anyone acted on the subscriber’s order.
P22.3Rejection and requote clustering
- Definition
- When and how orders are rejected or requoted, and whether those events cluster in the conditions where an informed intermediary would benefit.
- Estimator
- Event-rate comparison across declared conditions against the rate expected if rejections were independent of those conditions.
- Known failure modes
- Rejections are logged by the platform, and platforms differ in what they record. An unlogged rejection is invisible to the measurement, so the coverage gap is stated with the result.
- Does not establish
- Clustering does not establish deliberate handling. Routing, session load and the subscriber’s own infrastructure are enumerated beside every finding.
Standing limit. It will not name a culprit. Measured disadvantage has innocent explanations, and the instrument enumerates them beside every finding rather than in a footnote.
Full instrument page →
Risk instrumentation against the actual rules of a prop-firm evaluation — the arithmetic boundary that ends accounts before the market does.
P3.1Halt distance
- Definition
- Distance to every rule that can end an evaluation — daily loss, maximum drawdown, trailing thresholds — expressed as the adverse movement, at current sizing, that would breach it.
- Estimator
- Rule-by-rule computation against the account’s current positions and the evaluation’s published rule set; the reported distance is the minimum across rules, and the binding rule is named beside it.
- Known failure modes
- It is correct only if the rule set is transcribed correctly. Firms differ on how trailing thresholds ratchet and on whether unrealised profit counts, and a mis-transcribed rule produces a confidently wrong distance.
- Does not establish
- Halt distance does not establish that a breach will not occur. A gap through the level breaches it regardless of the distance measured a second earlier.
P3.2Correlation-adjusted heat
- Definition
- Effective exposure across the book once correlation between positions is accounted for — what five tickets in correlated instruments actually constitute.
- Estimator
- Exposure aggregated under a correlation matrix estimated over a lookback declared in advance, with the lookback printed beside the figure.
- Known failure modes
- Correlations estimated on calm data understate stress correlations, which is exactly when the figure matters. The estimate publishes with its lookback so a reader can see what it was fitted on.
- Does not establish
- It does not establish what the book will do in a shock. That question belongs to the Stress Harness, which replays recorded events instead of estimating a matrix.
P3.3Risk-of-ruin surface
- Definition
- The probability terrain of breaching the evaluation’s rules before passing them, given the current sizing grid.
- Estimator
- Computed over the declared sizing grid and rule set with the return-distribution assumption stated on the surface itself, and recomputed as the account moves.
- Known failure modes
- The surface is only as good as its return-distribution assumption, and the assumption most likely to be wrong is the one about the tail. Every surface publishes its assumptions rather than embedding them out of sight.
- Does not establish
- A low modelled ruin probability does not establish safety. It states what the model says under assumptions a reader can inspect and reject.
P3.4Permitted sizing grid
- Definition
- The position sizes the rules actually permit from the account’s current state, precomputed so a decision under pressure is a lookup rather than an estimate.
- Estimator
- Direct evaluation of the rule set against current equity, open exposure and remaining allowance — the boundary the rules impose, not a recommendation.
- Known failure modes
- A grid computed against a stale account state is worse than no grid at all, so the staleness of the input is reported alongside it.
- Does not establish
- A permitted size is not a suggested size. The instrument has no view on what to trade, or whether to.
Standing limit. It has no opinion on entries or strategy. It polices the boundary between the strategy and the rules the evaluee agreed to.
Full instrument page →
Historical shock replay against the current book: what named events — the SNB floor break, COVID, carry unwinds, the gilts episode — would do to today’s positions.
P8.1Named shock replay
- Definition
- The effect on the current book of a recorded historical shock, applied at current sizing and in the correlation structure the event itself produced.
- Estimator
- The recorded price path of the named event is applied to current positions. Joint moves come from the event’s own cross-instrument path, never from a correlation matrix estimated on quiet data.
- Known failure modes
- A replay is bounded by what the archive contains. Instruments that did not exist at the event, or whose history is thin, cannot be replayed and are reported as unreplayable rather than proxied by a lookalike.
- Does not establish
- Surviving a replay does not establish that the book is safe. The next shock is not in the library — by construction, the events that break books are the ones the sample did not contain.
P8.2Gap-through fill accounting
- Definition
- How an order is filled when the historical path moved past its level without ever printing there.
- Estimator
- Fills are placed at the next price for which a print exists in the record, never at the order’s own level. That single rule is what separates a stress test from a reassurance.
- Known failure modes
- Print records at the event are themselves incomplete. Where the record is sparse the fill is placed conservatively and the sparsity is flagged with the result rather than smoothed over.
- Does not establish
- It does not establish what a venue would have done. Requotes, spread widening and the order in which a broker liquidates are that broker’s behaviour, not a price series.
P8.3Vacuum-fill accounting
- Definition
- The treatment of intervals in which no quotes existed at all: the fill is marked at what the next real print permits and labelled a vacuum fill.
- Estimator
- Intervals with no prints are identified from the record, and the resulting fills are reported separately from fills that occurred against a live book.
- Known failure modes
- The boundary of a vacuum depends on which venue’s record is used; a different archive yields a different boundary, so the archive is named with the result.
- Does not establish
- A vacuum fill does not establish what the subscriber’s own broker would have done in that interval.
P8.4Ruin flag
- Definition
- A state recorded where a scenario closes the account out, published instead of a return figure — because a book that is closed out does not have a return.
- Estimator
- A scenario is flagged when equity crosses the supplied closeout condition at any point along the replayed path. The test is path-dependent, not end-of-path.
- Known failure modes
- The flag depends entirely on the closeout condition supplied; a margin regime transcribed incorrectly moves the flag in either direction.
- Does not establish
- The absence of a ruin flag does not establish survivability under any real event.
P8.5Position-level attribution
- Definition
- Which position carried the damage in each scenario, decomposed rather than reported as one book-level number.
- Estimator
- Per-position profit and loss along the replayed path, summed to the book total so the decomposition is exact rather than approximate.
- Known failure modes
- Attribution is interpretable only where positions are independent legs. A hedged structure attributes damage to one leg and relief to another, and must be read as a whole or not at all.
- Does not establish
- Attribution does not establish which position to cut. It states where the damage fell in this scenario.
Standing limit. It replays recorded history. It does not forecast the next shock, and a book that survives every replay is not thereby safe.
Full instrument page →
Versioning rule: this page is v1.0. A change to any definition, estimator, failure mode or limit above increments the version and appends an entry to the changelog in §04 naming what changed. Definitions are never edited silently, because a definition that can move after a result is published is not a definition — it is a degree of freedom.
Pre-registrations
A pre-registration is a declaration made and dated before the measurement runs: the thresholds that will decide, the rules that will exclude, and the sample bar below which the honest null publishes instead of a number. Its entire value comes from its ordering. Posted before the answer is known it is a constraint; posted afterwards it is a description of a result, which is a different and much cheaper object wearing the same clothes.
Pre-registration record · Desk DisciplineNOT YET PUBLISHED
No pre-registration has been posted for the Desk Discipline suite. Not one that is pending review, not one that is drafted and unhashed — none. This block is the structure a registration will occupy, printed empty on purpose, because the alternative is a page that describes a discipline while quietly implying it has already been exercised.
A pre-registration is worth exactly the provability of its ordering. It has to be posted, dated and content-hashed while the answer is still unknown; posted afterwards it is a description of a result, which is a different and much cheaper object. So the first registration cannot be backdated into this slot, and the slot stays visibly empty until one is posted in the only way that counts.
The field schema of a pre-registration record for the Desk Discipline suite: each field, what it will hold, and its current value. Every value reads NOT YET PUBLISHED because no registration exists.| Field | What it will hold | Value |
|---|
registration_ref | The permanent identifier this registration is cited by. | NOT YET PUBLISHED |
|---|
scope | The instruments and the measurement window the declaration binds. | NOT YET PUBLISHED |
|---|
declared_utc | When the declaration was posted — necessarily before any data was touched. | NOT YET PUBLISHED |
|---|
first_data_utc | When collection began. This must fall after the line above, and the ordering is the evidence. | NOT YET PUBLISHED |
|---|
thresholds | Every decision threshold, fixed while the answer was still unknown. | NOT YET PUBLISHED |
|---|
exclusion_rules | What will be dropped from the sample, and on what stated grounds. | NOT YET PUBLISHED |
|---|
completeness_bar | The minimum sample below which the honest null publishes instead of a number. | NOT YET PUBLISHED |
|---|
document_sha256 | The content hash of the registered document itself. Any later edit changes it, visibly. | NOT YET PUBLISHED |
|---|
amendments | Appended corrections, each with its own date and reason. The original text stays. | NOT YET PUBLISHED |
|---|
What a Desk Discipline registration must fix in advance. The lists below are classes of declaration, not declarations. They name the decisions that have to be made before the data is touched, because each one is a decision that could otherwise be made afterwards, in the direction that flatters the result. No value below has been registered.
Thresholds
- The correlation lookback used to compute effective exposure, fixed before it is computed.
- The closeout condition under which a replayed scenario is flagged as ruin.
- The units of adverse movement in which halt distance is expressed.
Exclusion rules
- How positions with insufficient history are treated inside the correlation estimate.
- How a fill is placed where the recorded path is sparse, and when an interval is declared a vacuum.
- Which historical scenarios are in the replay library — named before any book is run against them.
Completeness bars
- The minimum overlapping history required before two instruments enter the correlation estimate.
- The minimum print density required before a scenario is replayed rather than marked unreplayable.
- The minimum position count below which attribution publishes as a book-level figure only.
Immutability, stated before it is tested. Once a registration is posted it is not edited. If it is wrong, an amendment is appended beneath it carrying its own date and the reason for the change, and the original text stays where it is, readable, above the correction. A registration that quietly improved after the data arrived would be indistinguishable from one that was right all along — which is precisely why the append rule is written here, now, while there is nothing yet to be tempted by.
Amendment rule, stated in advance: a posted registration is never edited. An amendment is appended beneath the original carrying its own date and its reason, and the original text stays above it, readable. This page will show both.
Proof corpus
The corpus is the set of artifacts a published measurement ships with — not a description of them, the artifacts themselves, downloadable, each with the digest that proves you received the bytes I measured and the code tag that produced them. A result whose inputs cannot be downloaded is not a receipt; it is an assertion with better typography.
The corpus for this suite is empty. Every row below is a slot, and every slot is NOT YET PUBLISHED. The table is printed anyway, because a reader should be able to see the exact shape of what will arrive — and because a page that described a corpus without showing how empty it currently is would be making the claim it exists to refuse.
Five columns: artifact, contents, content hash, code tag, download. Scroll sideways if they do not all fit.
Proof corpus for the Desk Discipline suite: the seven artifact classes a published receipt carries, what each will contain, and its current state. Every content hash, code tag and download reads NOT YET PUBLISHED, because no artifact from this suite's proof corpus has been published yet.| Artifact | What it will contain | Content hash | Code tag | Download |
|---|
| Registered methodology document | The versioned document these definitions are taken from, in the exact form it was registered — estimators, parameters, and the limits stated above. | SHA-256NOT YET PUBLISHED | NOT YET PUBLISHED | NOT YET PUBLISHED |
|---|
| Pre-registration record | The dated declaration: thresholds, exclusion rules and completeness bars, plus any amendments appended beneath the original with their reasons. | SHA-256NOT YET PUBLISHED | NOT YET PUBLISHED | NOT YET PUBLISHED |
|---|
| Rule transcription, position snapshot and replay archive | The evaluation rule set as transcribed, the book as held at the measurement instant, and the recorded price paths for every named scenario. | SHA-256NOT YET PUBLISHED | NOT YET PUBLISHED | NOT YET PUBLISHED |
|---|
| Result set | Per-metric results with intervals and effective sample sizes, and the honest nulls wherever a sample could not support a metric. | SHA-256NOT YET PUBLISHED | NOT YET PUBLISHED | NOT YET PUBLISHED |
|---|
| Can-fail proof transcript | For every test in the battery: the planted defect, the refusal that was expected, and the outcome that was observed. | SHA-256NOT YET PUBLISHED | NOT YET PUBLISHED | NOT YET PUBLISHED |
|---|
| Kill-ledger extract | Hypotheses registered against this suite and killed by the data, each with the run that killed it. Published with the same visibility as a registration. | SHA-256NOT YET PUBLISHED | NOT YET PUBLISHED | NOT YET PUBLISHED |
|---|
| Generation code | The tagged commit that produced the result set, with its build receipt. Named here because a result whose code version is unstated cannot be re-run. | SHA-256NOT YET PUBLISHED | NOT YET PUBLISHED | NOT YET PUBLISHED |
|---|
Corpus state, per instrument
Per instrument, so that the emptiness cannot hide behind a suite-level summary.
Changelog
Append-only, dated, written by a person. It records changes to this page — it is not a measurement log, and it will not become one. A correction is a new entry that names the entry it corrects; nothing here is ever edited away, because a changelog you can rewrite is a marketing surface with a monospace font.
2026-08-01 · v1.0
Receipt page established for the Desk Discipline suite, carrying all five parts of the §8.2 anatomy: the versioned methodology with a formal definition, estimator, failure modes and non-claim for each of the 12 metrics its 3 instruments measure; the pre-registration structure with no registration in it; the proof-corpus table with no artifact in it; this changelog; and the independent-recomputation procedure. Supersedes the earlier per-suite methodology summary at this URL, which carried the battery outline without the receipt anatomy. No measurement, pre-registration, artifact or hash accompanies this version — every receipt slot below is empty as a matter of fact, not of omission.
Independent recomputation
Independent recomputation is what makes the rest of the page checkable rather than merely well-written. It is the exact sequence a stranger follows to re-derive a published number from this suite using only artifacts I published — no account, no request, no conversation with me at any point.
Today the procedure terminates at step 1, because no measurement from the Desk Discipline suite has been published and there is no receipt to open. The steps are written now, in the specific form they will take for this suite, precisely so that they exist before the first result does and cannot afterwards be shaped to fit one.
Open the receipt and take its four identifiers.
Every published measurement links a receipt carrying four: the registration reference, the methodology-document digest, the input-manifest digest, and the generation-code tag. If any one is missing, stop — the result is not recomputable and should not be treated as though it were, including by me.
Verify the methodology document against its digest.
Download it, hash it, compare. A mismatch means the method you are about to apply is not the method that was registered, and everything after this step would be measuring a different thing.
Check the ordering before you check anything else.
The registration timestamp must precede the first-data timestamp on the manifest. If it does not, the registration is a description of a result rather than a constraint on one, and no statistic downstream can repair that.
Download the rule transcription, the position snapshot and the replay archive, and verify every digest.
The rule transcription is a first-class artifact here. Most disagreements about a desk number turn out to be disagreements about a rule, and this is the step where that surfaces.
Apply the registered exclusion rules yourself.
The position-inclusion rule, the correlation lookback, and the treatment of positions with insufficient history. A lookback chosen after seeing the answer is precisely the failure this step exists to prevent.
Check out the harness at the code tag on the receipt and run it against the verified snapshot.
A replay is deterministic given the archive: the same path, the same fills, the same flags, every time.
Recompute the halt distances, the correlation-adjusted exposure, and per scenario the fill sequence and the ruin flag.
Fills must match exactly. The fill rule is arithmetic over the recorded path — it is not a model, and it has no tolerance band.
Run the can-fail proof.
Replace one scenario path with a synthetic path that gaps through every stop, and confirm the harness reports gap-through fills and a ruin flag rather than filling politely at the stop level.
If your number differs, the difference is the finding.
Send it with your inputs and the version you ran. A confirmed discrepancy publishes as a correction appended beside the original — and the original stays exactly where it is, unedited, because the error is the part of the record that proves the discipline is real.
The point of publishing this before there is anything to check: a recomputation procedure written after a result is a procedure written by someone who already knows which steps would be inconvenient.