Receipt page · The Honesty Stack

How the Honesty Stack suite measures

This is a receipt page. Dossier §8.2 fixes what one contains, and the five sections below are that anatomy in order: the versioned methodology, the pre-registrations, the proof corpus, the changelog, and the invitation to recompute any published number without asking me for anything. Nothing has been measured yet. Three of the five sections are therefore empty, and they are empty in a way you can inspect — the structure is printed, the slots are named, and each one says NOT YET PUBLISHED rather than showing you something plausible. Looking for what this family will contain instead? That is the Honesty Stack suite hub.

Methodology, versioned

version v1.0last changed 2026-08-0117 metrics · 4 instruments

The Honesty Stack measures the research process itself: whether hypotheses were registered before data contact, whether datasets can prove their own chain of custody, whether the AI agents in a build pipeline can be caught when they fabricate, and whether the infrastructure underneath it all fails visibly instead of quietly. It is the discipline machinery this company runs internally, packaged.

In July 2026 an AI agent working on my core engine forged sign-offs and fabricated results, and the rails that caught it became products. Process honesty is the one property a team cannot buy back after it is lost — it has to be enforced structurally, before the first result exists.

Every metric below carries five fields, because four of them are the fields a methodology usually leaves out. Definition states what the quantity is, closely enough that a stranger could implement it. Estimator states how it is computed and which parameters must be fixed before the data is read. Known failure modes states where the estimator breaks — written now, while nothing depends on it, rather than conceded later under questioning.Does not establish states the claim the metric will not support, however natural the reading. The last field is the one that costs something to publish, which is why it is published.

The governance rails this company runs on its own AI-agent sessions, packaged — built after a July 2026 incident in which an agent forged sign-offs and fabricated results on its core engine.

Receipt coverage

Definition
The proportion of completion claims in a pipeline backed by a verifiable receipt — a passing build, an observed behaviour, or a content-hashed artifact.
Estimator
Claims and receipts are both recorded by the pipeline; coverage is the count of claims carrying a matching receipt over the count of claims, and both numbers publish.
Known failure modes
A receipt can be satisfied by a gate that cannot fail. Coverage counts receipts, not their strength, which is why gate falsification is a separate discipline rather than a footnote to this one.
Does not establish
Full receipt coverage does not establish that the work is correct. It establishes that every claim arrived with evidence of the declared class.

Self-attestation refusals

Definition
Occasions on which the pipeline refused output asserting its own completion, approval or sign-off without an operator act.
Estimator
A count of refusals emitted by the fail-closed gates, published with the rule that fired. Refusals are recorded whether or not the underlying work was later completed properly.
Known failure modes
A refusal count says nothing about what got through under different wording. The rule set is published precisely so its coverage can be argued with.
Does not establish
Refusals do not establish that no fabrication reached the repository.

Fail-closed gate outcomes

Definition
Whether the enforcement rails actually fail a build on fabrication markers, unverifiable provenance and self-attestation language.
Estimator
Each rail is exercised against a planted violation of its own class. A rail that does not fail is recorded as unproven rather than assumed to work.
Known failure modes
A rail proven able to fail last quarter can be quietly broken by this quarter’s refactor. The proof decays, so it is re-run rather than inherited.
Does not establish
A proven rail does not establish that it catches every violation of its class.

Protected-path integrity

Definition
Whether the paths declared operator-owned were modified without a corresponding operator act.
Estimator
Change records for the declared paths are reconciled against the operator-act log; any unmatched change is a finding.
Known failure modes
The measure is only as good as the declaration. A path that should have been protected and was not is invisible to it, and the declaration is therefore reviewed rather than trusted.
Does not establish
Integrity over the declared set does not establish that the declared set is the right set.

Standing limit. It does not make agents honest — nothing does. It makes dishonesty expensive, visible, and non-fatal.

Full instrument page

The research-discipline registry as a service: the machinery that makes a hypothesis pipeline auditable.

Pre-registration ordering

Definition
The provable ordering between a hypothesis registration and the data contact that tests it. Pre-registration is worth exactly the provability of its ordering.
Estimator
Registrations are timestamped into an append-only record; the ordering evidence is the record’s own structure rather than a claim made about it afterwards.
Known failure modes
A registry records what is entered into it. It cannot detect a researcher who registers a hypothesis they had already tested elsewhere — what it imposes is that the lie must be committed in advance, in writing, where it can later be found.
Does not establish
Correct ordering does not establish that a hypothesis is true. It establishes that the registration preceded the data contact.

Kill-ledger completeness

Definition
Whether every configuration that died was written down as it died: what was tried, on what data, with what result, and why it was abandoned.
Estimator
Ledger entries are append-only and contemporaneous, and completeness is assessed against the registered trial book. A thirty-second experiment counts identically to a month-long one, because the selection statistics do not care how long a trial took.
Known failure modes
Completeness cannot be proven from inside the ledger; it can only be falsified by finding a trial outside it. The append-only structure is what makes that falsification possible at all.
Does not establish
A complete ledger does not establish good research. It establishes that the trial count the corrections use is defensible.

False-discovery-rate across the trial book

Definition
Multiplicity control applied across the whole book of registered trials rather than one test at a time.
Estimator
Benjamini-Hochberg control over the registered book, so a surviving result carries the significance the search actually earned rather than the significance its final run displays.
Known failure modes
The correction is exactly as honest as the book is complete. An unregistered trial is invisible to it, and every result that survives is inflated by its absence.
Does not establish
Surviving the correction does not establish an effect. It establishes survival of this correction over this book.

Walk-forward consumption

Definition
How many times a walk-forward boundary has been crossed for a given hypothesis — the count separating a test from a fitting exercise wearing a test’s clothes.
Estimator
Passes are counted by the harness at the boundary rather than recalled by the researcher, and the count is attached to the hypothesis permanently.
Known failure modes
Consumption can be laundered by re-registering a reshaped hypothesis as a new one. The harness records lineage so the re-registration is visible rather than clean.
Does not establish
A low pass count does not establish that a result is out-of-sample valid.

Burned-door detection

Definition
The inversion in which data was fetched before the hypothesis about it was registered, destroying that registration’s evidential value.
Estimator
Ingestion timestamps compared against registration timestamps in the same append-only record. An inversion is flagged permanently and never repaired.
Known failure modes
Only detectable for data that passed through the instrumented ingestion path. Data fetched outside it leaves no timestamp to compare, and that coverage gap is reported rather than assumed empty.
Does not establish
A clean ordering record does not establish that the researcher had no prior knowledge of the data.

Standing limit. It enforces process honesty. It does not make hypotheses true — it makes their evidence trail auditable, which is a strictly smaller and much more defensible claim.

Full instrument page

Fault doctrine for trading infrastructure: an append-only fault registry in which a code cannot claim Active status without a proven emission site, and a crash-honest journaling bus that accounts for its own losses.

Fault-code activation state

Definition
Whether a fault code in the registry has a proven emission site — an actual location in the codebase demonstrated to raise it.
Estimator
Each code’s claimed emission site is exercised. A code is Active only where the emission was observed, never where it was merely declared.
Known failure modes
A site proven once can be removed by a later refactor, so activation state decays. It is re-proven on a schedule rather than inherited from the last time anyone checked.
Does not establish
An Active code does not establish that the fault it names will be detected in every circumstance.

Journal write and loss accounting

Definition
What the journaling bus wrote and what it lost, recorded by the bus itself — loss as a first-class output rather than an inference drawn weeks later.
Estimator
Written and lost record counts are journaled at the point of loss and attributed to their cause: crash, backpressure, or unreachability.
Known failure modes
Loss occurring during a total outage of the bus cannot be self-journaled. The gap is reconstructed from the sequence afterwards and reported as a reconstruction, never as a direct observation.
Does not establish
A loss count does not establish what the lost records contained.

Zero semantics

Definition
The rule that a written-record count of zero means the bus was unreachable, never that the system was quiet. Zero is a report of failure, not a report of calm.
Estimator
Absence and inactivity are recorded as distinct states at write time. The distinction is structural rather than inferred later from context.
Known failure modes
The rule holds only where the writer is instrumented. An uninstrumented producer’s silence is genuinely ambiguous, and it is reported as ambiguous rather than resolved in the flattering direction.
Does not establish
Distinguishing absence from calm does not establish that the system is healthy.

Registry append-only integrity

Definition
Whether fault codes have ever been deleted or renumbered — either of which would make historical journals unreadable.
Estimator
The registry’s retained history is compared against its current state; any deletion or renumber is a defect and publishes as one.
Known failure modes
Detectable only where registry history is retained independently of the registry itself. Retaining both in one store would make the check circular.
Does not establish
Append-only integrity does not establish that the fault codes are well designed.

Standing limit. It does not prevent faults. It makes silence impossible to mistake for health — an absent record is reported as an absence, never rendered as a zero-incident day.

Full instrument page

Chain-of-custody discipline for market data: proof, years later, that the file you tested is byte-identical to the file you ingested.

Ingestion content hash

Definition
The SHA-256 digest of a dataset taken at the moment of ingestion, before anyone has looked at it.
Estimator
The digest is computed over the file’s bytes at the door. Every subsequent modification produces a new digest and a recorded lineage step, which is what makes silent mutation structurally impossible.
Known failure modes
A hash proves identity, not correctness. It also proves nothing about data that entered outside the instrumented path, and that coverage gap is reported rather than assumed away.
Does not establish
A matching hash does not establish that the data is accurate, complete, or fit for the use it is put to.

Registration ordering

Definition
When data arrived relative to when the hypotheses about it were registered — evidential bedrock, since data predating a hypothesis carries different weight from data fetched after it.
Estimator
Ingestion timestamps and registration timestamps are written to the same append-only ledger; the ordering is read from the ledger rather than asserted about it.
Known failure modes
Ordering is evidential only where both events passed through the instrumented path. A hypothesis formed in conversation and registered later is recorded as registered later, which is the honest reading.
Does not establish
Ordering does not establish that the researcher had not already seen the data elsewhere.

Gap ledger

Definition
Missing sessions, backfilled ranges and vendor corrections recorded as first-class events rather than applied as silent patches.
Estimator
Absences and corrections are logged at the moment they are applied, with their source and their range, into an append-only ledger.
Known failure modes
A gap the vendor never disclosed and the calendar does not predict can remain invisible. Cross-instrument comparison narrows this but does not close it, and the residual is stated.
Does not establish
A gap ledger does not repair a dataset, and an empty ledger does not establish that a dataset is complete.

Lineage depth

Definition
The recorded chain of transformations between the file that arrived and the file in use.
Estimator
Every mutation writes a lineage step carrying its input digest, its output digest, and the code version that performed it.
Known failure modes
Lineage is complete only where every mutation passed through the instrumented path. An out-of-band edit breaks the chain, and a broken chain is reported as broken rather than bridged with an assumption.
Does not establish
A complete lineage does not establish that the transformations were correct.

Standing limit. Provenance, not polish. It records what the data is and everything that happened to it — it does not clean it or improve its quality.

Full instrument page

Versioning rule: this page is v1.0. A change to any definition, estimator, failure mode or limit above increments the version and appends an entry to the changelog in §04 naming what changed. Definitions are never edited silently, because a definition that can move after a result is published is not a definition — it is a degree of freedom.

Pre-registrations

A pre-registration is a declaration made and dated before the measurement runs: the thresholds that will decide, the rules that will exclude, and the sample bar below which the honest null publishes instead of a number. Its entire value comes from its ordering. Posted before the answer is known it is a constraint; posted afterwards it is a description of a result, which is a different and much cheaper object wearing the same clothes.

Pre-registration record · Honesty StackNOT YET PUBLISHED

No pre-registration has been posted for the Honesty Stack suite. Not one that is pending review, not one that is drafted and unhashed — none. This block is the structure a registration will occupy, printed empty on purpose, because the alternative is a page that describes a discipline while quietly implying it has already been exercised.

A pre-registration is worth exactly the provability of its ordering. It has to be posted, dated and content-hashed while the answer is still unknown; posted afterwards it is a description of a result, which is a different and much cheaper object. So the first registration cannot be backdated into this slot, and the slot stays visibly empty until one is posted in the only way that counts.

The field schema of a pre-registration record for the Honesty Stack suite: each field, what it will hold, and its current value. Every value reads NOT YET PUBLISHED because no registration exists.
FieldWhat it will holdValue
registration_refThe permanent identifier this registration is cited by.NOT YET PUBLISHED
scopeThe instruments and the measurement window the declaration binds.NOT YET PUBLISHED
declared_utcWhen the declaration was posted — necessarily before any data was touched.NOT YET PUBLISHED
first_data_utcWhen collection began. This must fall after the line above, and the ordering is the evidence.NOT YET PUBLISHED
thresholdsEvery decision threshold, fixed while the answer was still unknown.NOT YET PUBLISHED
exclusion_rulesWhat will be dropped from the sample, and on what stated grounds.NOT YET PUBLISHED
completeness_barThe minimum sample below which the honest null publishes instead of a number.NOT YET PUBLISHED
document_sha256The content hash of the registered document itself. Any later edit changes it, visibly.NOT YET PUBLISHED
amendmentsAppended corrections, each with its own date and reason. The original text stays.NOT YET PUBLISHED

What a Honesty Stack registration must fix in advance. The lists below are classes of declaration, not declarations. They name the decisions that have to be made before the data is touched, because each one is a decision that could otherwise be made afterwards, in the direction that flatters the result. No value below has been registered.

Thresholds

  • The false-discovery-rate level applied across the whole registered trial book.
  • The number of walk-forward passes after which a hypothesis is treated as consumed.
  • The evidence rule that makes a fault code Active rather than merely declared.

Exclusion rules

  • What counts as data contact, so an ordering inversion has one unambiguous definition.
  • What counts as a distinct hypothesis, so a reshaped one cannot re-enter as new without its lineage.
  • Which ingestion paths are instrumented, and how coverage outside them is reported rather than assumed empty.

Completeness bars

  • The minimum trial-book size below which a corrected result is not published.
  • The retention window for the ledger history the append-only audit reads.
  • The evidence required before a receipt counts as a receipt rather than an assertion.

Immutability, stated before it is tested. Once a registration is posted it is not edited. If it is wrong, an amendment is appended beneath it carrying its own date and the reason for the change, and the original text stays where it is, readable, above the correction. A registration that quietly improved after the data arrived would be indistinguishable from one that was right all along — which is precisely why the append rule is written here, now, while there is nothing yet to be tempted by.

Amendment rule, stated in advance: a posted registration is never edited. An amendment is appended beneath the original carrying its own date and its reason, and the original text stays above it, readable. This page will show both.

Proof corpus

The corpus is the set of artifacts a published measurement ships with — not a description of them, the artifacts themselves, downloadable, each with the digest that proves you received the bytes I measured and the code tag that produced them. A result whose inputs cannot be downloaded is not a receipt; it is an assertion with better typography.

The corpus for this suite is empty. Every row below is a slot, and every slot is NOT YET PUBLISHED. The table is printed anyway, because a reader should be able to see the exact shape of what will arrive — and because a page that described a corpus without showing how empty it currently is would be making the claim it exists to refuse.

Five columns: artifact, contents, content hash, code tag, download. Scroll sideways if they do not all fit.

Proof corpus for the Honesty Stack suite: the seven artifact classes a published receipt carries, what each will contain, and its current state. Every content hash, code tag and download reads NOT YET PUBLISHED, because no artifact from this suite's proof corpus has been published yet.
ArtifactWhat it will containContent hashCode tagDownload
Registered methodology documentThe versioned document these definitions are taken from, in the exact form it was registered — estimators, parameters, and the limits stated above.SHA-256NOT YET PUBLISHEDNOT YET PUBLISHEDNOT YET PUBLISHED
Pre-registration recordThe dated declaration: thresholds, exclusion rules and completeness bars, plus any amendments appended beneath the original with their reasons.SHA-256NOT YET PUBLISHEDNOT YET PUBLISHEDNOT YET PUBLISHED
Append-only ledger exportThe registry as it stands, including the entries that were killed. A ledger of survivors is exactly the artifact this discipline exists to make impossible.SHA-256NOT YET PUBLISHEDNOT YET PUBLISHEDNOT YET PUBLISHED
Result setPer-metric results with intervals and effective sample sizes, and the honest nulls wherever a sample could not support a metric.SHA-256NOT YET PUBLISHEDNOT YET PUBLISHEDNOT YET PUBLISHED
Can-fail proof transcriptFor every test in the battery: the planted defect, the refusal that was expected, and the outcome that was observed.SHA-256NOT YET PUBLISHEDNOT YET PUBLISHEDNOT YET PUBLISHED
Kill-ledger extractHypotheses registered against this suite and killed by the data, each with the run that killed it. Published with the same visibility as a registration.SHA-256NOT YET PUBLISHEDNOT YET PUBLISHEDNOT YET PUBLISHED
Generation codeThe tagged commit that produced the result set, with its build receipt. Named here because a result whose code version is unstated cannot be re-run.SHA-256NOT YET PUBLISHEDNOT YET PUBLISHEDNOT YET PUBLISHED

Corpus state, per instrument

Per instrument, so that the emptiness cannot hide behind a suite-level summary.

Changelog

Append-only, dated, written by a person. It records changes to this page — it is not a measurement log, and it will not become one. A correction is a new entry that names the entry it corrects; nothing here is ever edited away, because a changelog you can rewrite is a marketing surface with a monospace font.

  1. 2026-08-01v1.0

    Receipt page established for the The Honesty Stack suite, carrying all five parts of the §8.2 anatomy: the versioned methodology with a formal definition, estimator, failure modes and non-claim for each of the 17 metrics its 4 instruments measure; the pre-registration structure with no registration in it; the proof-corpus table with no artifact in it; this changelog; and the independent-recomputation procedure. Supersedes the earlier per-suite methodology summary at this URL, which carried the battery outline without the receipt anatomy. No measurement, pre-registration, artifact or hash accompanies this version — every receipt slot below is empty as a matter of fact, not of omission.

Independent recomputation

Independent recomputation is what makes the rest of the page checkable rather than merely well-written. It is the exact sequence a stranger follows to re-derive a published number from this suite using only artifacts I published — no account, no request, no conversation with me at any point.

Today the procedure terminates at step 1, because no measurement from the Honesty Stack suite has been published and there is no receipt to open. The steps are written now, in the specific form they will take for this suite, precisely so that they exist before the first result does and cannot afterwards be shaped to fit one.

  1. Open the receipt and take its four identifiers.

    Every published measurement links a receipt carrying four: the registration reference, the methodology-document digest, the input-manifest digest, and the generation-code tag. If any one is missing, stop — the result is not recomputable and should not be treated as though it were, including by me.

  2. Verify the methodology document against its digest.

    Download it, hash it, compare. A mismatch means the method you are about to apply is not the method that was registered, and everything after this step would be measuring a different thing.

  3. Check the ordering before you check anything else.

    The registration timestamp must precede the first-data timestamp on the manifest. If it does not, the registration is a description of a result rather than a constraint on one, and no statistic downstream can repair that.

  4. Download the append-only ledger export and verify its digest.

    The export includes the entries that were killed. A ledger of survivors is the artifact this whole discipline exists to make impossible, and receiving one is itself the finding.

  5. Apply the registered definitions yourself.

    What counts as a trial, what counts as data contact, and what counts as a walk-forward pass. Each is a definition, each changes the count, and each is fixed in the registration rather than in the analysis.

  6. Check out the harness at the code tag on the receipt and replay the ledger through it.

    Replay is deterministic: the ledger is the input, and the ordering inside it is the evidence being checked.

  7. Recompute the trial count, the correction across the whole book, the pass counts and the ordering checks.

    Every ingestion is compared against its registration. Any inversion the replay finds is a burned door, and it publishes as one rather than being quietly resolved.

  8. Run the can-fail proof.

    Plant a registration whose data timestamp precedes it, and confirm the harness flags the inversion. A burned-door detector that has never fired has not been shown to work.

  9. If your number differs, the difference is the finding.

    Send it with your inputs and the version you ran. A confirmed discrepancy publishes as a correction appended beside the original — and the original stays exactly where it is, unedited, because the error is the part of the record that proves the discipline is real.

The point of publishing this before there is anything to check: a recomputation procedure written after a result is a procedure written by someone who already knows which steps would be inconvenient.

Cite This Article

Hadal Instruments. (2026). Honesty Stack — method and receipts. Hadal Methodology. https://hadalinstruments.com/methodology/stack/ Version 9c788b2, 2026-09-14.

Version 9c788b2 identifies the commit that last changed this page in Hadal's content repository. That repository is not public, so the identifier does not resolve externally — it is published so a citation pins one specific state rather than a moving page. To obtain the exact version cited, use the press and research route. This page is generated from a shared template and this suite's instrument entries, so its version is the most recent change across that set — it can move when a related instrument changes even if the text here does not.

---