Receipt page · The Honesty Stack
How the Honesty Stack suite measures
This is a receipt page. Dossier §8.2 fixes what one contains, and the five sections below are that anatomy in order: the versioned methodology, the pre-registrations, the proof corpus, the changelog, and the invitation to recompute any published number without asking me for anything. Nothing has been measured yet. Three of the five sections are therefore empty, and they are empty in a way you can inspect — the structure is printed, the slots are named, and each one says NOT YET PUBLISHED rather than showing you something plausible. Looking for what this family will contain instead? That is the Honesty Stack suite hub.
Receipt anatomy · dossier §8.2Honesty Stack suiteReceipt version v1.0
- Instruments in scope
- 4
- Metrics defined
- 17
- Pre-registrations posted
- NOT YET PUBLISHED
- Artifacts published
- NOT YET PUBLISHED
- Measurements published
- NOT YET PUBLISHED
- 01Methodology, versionedProse and formal definitions. Every metric carries its name, its definition, the estimator that computes it, its known failure modes, and what it does not establish. The page carries a version, and the version is printed on it.DEFINED · v1.0
- 02Pre-registrationsDeclarations made and dated before any measurement runs: decision thresholds, exclusion rules, completeness bars. Immutable once posted — an amendment appends beneath the original with its reason, and never overwrites it.NOT YET PUBLISHED
- 03Proof corpusThe published artifacts themselves: dataset descriptors, content hashes, download links, and the version tag of the code that generated each one. A result whose inputs are not downloadable is not a receipt.NOT YET PUBLISHED
- 04ChangelogAppend-only, dated, written by a person. What changed on this page, when, and why. A correction is a new entry naming the entry it corrects — entries are never edited away.1 ENTRY
- 05Independent recomputationThe exact numbered steps a stranger follows to re-derive a published number without asking me for anything. Where the recomputation disagrees, the disagreement is the finding.PROCEDURE STATED
Methodology, versioned
version v1.0last changed 2026-08-0117 metrics · 4 instruments
The Honesty Stack measures the research process itself: whether hypotheses were registered before data contact, whether datasets can prove their own chain of custody, whether the AI agents in a build pipeline can be caught when they fabricate, and whether the infrastructure underneath it all fails visibly instead of quietly. It is the discipline machinery this company runs internally, packaged.
In July 2026 an AI agent working on my core engine forged sign-offs and fabricated results, and the rails that caught it became products. Process honesty is the one property a team cannot buy back after it is lost — it has to be enforced structurally, before the first result exists.
Every metric below carries five fields, because four of them are the fields a methodology usually leaves out. Definition states what the quantity is, closely enough that a stranger could implement it. Estimator states how it is computed and which parameters must be fixed before the data is read. Known failure modes states where the estimator breaks — written now, while nothing depends on it, rather than conceded later under questioning.Does not establish states the claim the metric will not support, however natural the reading. The last field is the one that costs something to publish, which is why it is published.
The governance rails this company runs on its own AI-agent sessions, packaged — built after a July 2026 incident in which an agent forged sign-offs and fabricated results on its core engine.
P5.1Receipt coverage
- Definition
- The proportion of completion claims in a pipeline backed by a verifiable receipt — a passing build, an observed behaviour, or a content-hashed artifact.
- Estimator
- Claims and receipts are both recorded by the pipeline; coverage is the count of claims carrying a matching receipt over the count of claims, and both numbers publish.
- Known failure modes
- A receipt can be satisfied by a gate that cannot fail. Coverage counts receipts, not their strength, which is why gate falsification is a separate discipline rather than a footnote to this one.
- Does not establish
- Full receipt coverage does not establish that the work is correct. It establishes that every claim arrived with evidence of the declared class.
P5.2Self-attestation refusals
- Definition
- Occasions on which the pipeline refused output asserting its own completion, approval or sign-off without an operator act.
- Estimator
- A count of refusals emitted by the fail-closed gates, published with the rule that fired. Refusals are recorded whether or not the underlying work was later completed properly.
- Known failure modes
- A refusal count says nothing about what got through under different wording. The rule set is published precisely so its coverage can be argued with.
- Does not establish
- Refusals do not establish that no fabrication reached the repository.
P5.3Fail-closed gate outcomes
- Definition
- Whether the enforcement rails actually fail a build on fabrication markers, unverifiable provenance and self-attestation language.
- Estimator
- Each rail is exercised against a planted violation of its own class. A rail that does not fail is recorded as unproven rather than assumed to work.
- Known failure modes
- A rail proven able to fail last quarter can be quietly broken by this quarter’s refactor. The proof decays, so it is re-run rather than inherited.
- Does not establish
- A proven rail does not establish that it catches every violation of its class.
P5.4Protected-path integrity
- Definition
- Whether the paths declared operator-owned were modified without a corresponding operator act.
- Estimator
- Change records for the declared paths are reconciled against the operator-act log; any unmatched change is a finding.
- Known failure modes
- The measure is only as good as the declaration. A path that should have been protected and was not is invisible to it, and the declaration is therefore reviewed rather than trusted.
- Does not establish
- Integrity over the declared set does not establish that the declared set is the right set.
Standing limit. It does not make agents honest — nothing does. It makes dishonesty expensive, visible, and non-fatal.
Full instrument page →
The research-discipline registry as a service: the machinery that makes a hypothesis pipeline auditable.
P4.1Pre-registration ordering
- Definition
- The provable ordering between a hypothesis registration and the data contact that tests it. Pre-registration is worth exactly the provability of its ordering.
- Estimator
- Registrations are timestamped into an append-only record; the ordering evidence is the record’s own structure rather than a claim made about it afterwards.
- Known failure modes
- A registry records what is entered into it. It cannot detect a researcher who registers a hypothesis they had already tested elsewhere — what it imposes is that the lie must be committed in advance, in writing, where it can later be found.
- Does not establish
- Correct ordering does not establish that a hypothesis is true. It establishes that the registration preceded the data contact.
P4.2Kill-ledger completeness
- Definition
- Whether every configuration that died was written down as it died: what was tried, on what data, with what result, and why it was abandoned.
- Estimator
- Ledger entries are append-only and contemporaneous, and completeness is assessed against the registered trial book. A thirty-second experiment counts identically to a month-long one, because the selection statistics do not care how long a trial took.
- Known failure modes
- Completeness cannot be proven from inside the ledger; it can only be falsified by finding a trial outside it. The append-only structure is what makes that falsification possible at all.
- Does not establish
- A complete ledger does not establish good research. It establishes that the trial count the corrections use is defensible.
P4.3False-discovery-rate across the trial book
- Definition
- Multiplicity control applied across the whole book of registered trials rather than one test at a time.
- Estimator
- Benjamini-Hochberg control over the registered book, so a surviving result carries the significance the search actually earned rather than the significance its final run displays.
- Known failure modes
- The correction is exactly as honest as the book is complete. An unregistered trial is invisible to it, and every result that survives is inflated by its absence.
- Does not establish
- Surviving the correction does not establish an effect. It establishes survival of this correction over this book.
P4.4Walk-forward consumption
- Definition
- How many times a walk-forward boundary has been crossed for a given hypothesis — the count separating a test from a fitting exercise wearing a test’s clothes.
- Estimator
- Passes are counted by the harness at the boundary rather than recalled by the researcher, and the count is attached to the hypothesis permanently.
- Known failure modes
- Consumption can be laundered by re-registering a reshaped hypothesis as a new one. The harness records lineage so the re-registration is visible rather than clean.
- Does not establish
- A low pass count does not establish that a result is out-of-sample valid.
P4.5Burned-door detection
- Definition
- The inversion in which data was fetched before the hypothesis about it was registered, destroying that registration’s evidential value.
- Estimator
- Ingestion timestamps compared against registration timestamps in the same append-only record. An inversion is flagged permanently and never repaired.
- Known failure modes
- Only detectable for data that passed through the instrumented ingestion path. Data fetched outside it leaves no timestamp to compare, and that coverage gap is reported rather than assumed empty.
- Does not establish
- A clean ordering record does not establish that the researcher had no prior knowledge of the data.
Standing limit. It enforces process honesty. It does not make hypotheses true — it makes their evidence trail auditable, which is a strictly smaller and much more defensible claim.
Full instrument page →
Fault doctrine for trading infrastructure: an append-only fault registry in which a code cannot claim Active status without a proven emission site, and a crash-honest journaling bus that accounts for its own losses.
P19.1Fault-code activation state
- Definition
- Whether a fault code in the registry has a proven emission site — an actual location in the codebase demonstrated to raise it.
- Estimator
- Each code’s claimed emission site is exercised. A code is Active only where the emission was observed, never where it was merely declared.
- Known failure modes
- A site proven once can be removed by a later refactor, so activation state decays. It is re-proven on a schedule rather than inherited from the last time anyone checked.
- Does not establish
- An Active code does not establish that the fault it names will be detected in every circumstance.
P19.2Journal write and loss accounting
- Definition
- What the journaling bus wrote and what it lost, recorded by the bus itself — loss as a first-class output rather than an inference drawn weeks later.
- Estimator
- Written and lost record counts are journaled at the point of loss and attributed to their cause: crash, backpressure, or unreachability.
- Known failure modes
- Loss occurring during a total outage of the bus cannot be self-journaled. The gap is reconstructed from the sequence afterwards and reported as a reconstruction, never as a direct observation.
- Does not establish
- A loss count does not establish what the lost records contained.
P19.3Zero semantics
- Definition
- The rule that a written-record count of zero means the bus was unreachable, never that the system was quiet. Zero is a report of failure, not a report of calm.
- Estimator
- Absence and inactivity are recorded as distinct states at write time. The distinction is structural rather than inferred later from context.
- Known failure modes
- The rule holds only where the writer is instrumented. An uninstrumented producer’s silence is genuinely ambiguous, and it is reported as ambiguous rather than resolved in the flattering direction.
- Does not establish
- Distinguishing absence from calm does not establish that the system is healthy.
P19.4Registry append-only integrity
- Definition
- Whether fault codes have ever been deleted or renumbered — either of which would make historical journals unreadable.
- Estimator
- The registry’s retained history is compared against its current state; any deletion or renumber is a defect and publishes as one.
- Known failure modes
- Detectable only where registry history is retained independently of the registry itself. Retaining both in one store would make the check circular.
- Does not establish
- Append-only integrity does not establish that the fault codes are well designed.
Standing limit. It does not prevent faults. It makes silence impossible to mistake for health — an absent record is reported as an absence, never rendered as a zero-incident day.
Full instrument page →
Chain-of-custody discipline for market data: proof, years later, that the file you tested is byte-identical to the file you ingested.
P15.1Ingestion content hash
- Definition
- The SHA-256 digest of a dataset taken at the moment of ingestion, before anyone has looked at it.
- Estimator
- The digest is computed over the file’s bytes at the door. Every subsequent modification produces a new digest and a recorded lineage step, which is what makes silent mutation structurally impossible.
- Known failure modes
- A hash proves identity, not correctness. It also proves nothing about data that entered outside the instrumented path, and that coverage gap is reported rather than assumed away.
- Does not establish
- A matching hash does not establish that the data is accurate, complete, or fit for the use it is put to.
P15.2Registration ordering
- Definition
- When data arrived relative to when the hypotheses about it were registered — evidential bedrock, since data predating a hypothesis carries different weight from data fetched after it.
- Estimator
- Ingestion timestamps and registration timestamps are written to the same append-only ledger; the ordering is read from the ledger rather than asserted about it.
- Known failure modes
- Ordering is evidential only where both events passed through the instrumented path. A hypothesis formed in conversation and registered later is recorded as registered later, which is the honest reading.
- Does not establish
- Ordering does not establish that the researcher had not already seen the data elsewhere.
P15.3Gap ledger
- Definition
- Missing sessions, backfilled ranges and vendor corrections recorded as first-class events rather than applied as silent patches.
- Estimator
- Absences and corrections are logged at the moment they are applied, with their source and their range, into an append-only ledger.
- Known failure modes
- A gap the vendor never disclosed and the calendar does not predict can remain invisible. Cross-instrument comparison narrows this but does not close it, and the residual is stated.
- Does not establish
- A gap ledger does not repair a dataset, and an empty ledger does not establish that a dataset is complete.
P15.4Lineage depth
- Definition
- The recorded chain of transformations between the file that arrived and the file in use.
- Estimator
- Every mutation writes a lineage step carrying its input digest, its output digest, and the code version that performed it.
- Known failure modes
- Lineage is complete only where every mutation passed through the instrumented path. An out-of-band edit breaks the chain, and a broken chain is reported as broken rather than bridged with an assumption.
- Does not establish
- A complete lineage does not establish that the transformations were correct.
Standing limit. Provenance, not polish. It records what the data is and everything that happened to it — it does not clean it or improve its quality.
Full instrument page →
Versioning rule: this page is v1.0. A change to any definition, estimator, failure mode or limit above increments the version and appends an entry to the changelog in §04 naming what changed. Definitions are never edited silently, because a definition that can move after a result is published is not a definition — it is a degree of freedom.
Pre-registrations
A pre-registration is a declaration made and dated before the measurement runs: the thresholds that will decide, the rules that will exclude, and the sample bar below which the honest null publishes instead of a number. Its entire value comes from its ordering. Posted before the answer is known it is a constraint; posted afterwards it is a description of a result, which is a different and much cheaper object wearing the same clothes.
Pre-registration record · Honesty StackNOT YET PUBLISHED
No pre-registration has been posted for the Honesty Stack suite. Not one that is pending review, not one that is drafted and unhashed — none. This block is the structure a registration will occupy, printed empty on purpose, because the alternative is a page that describes a discipline while quietly implying it has already been exercised.
A pre-registration is worth exactly the provability of its ordering. It has to be posted, dated and content-hashed while the answer is still unknown; posted afterwards it is a description of a result, which is a different and much cheaper object. So the first registration cannot be backdated into this slot, and the slot stays visibly empty until one is posted in the only way that counts.
The field schema of a pre-registration record for the Honesty Stack suite: each field, what it will hold, and its current value. Every value reads NOT YET PUBLISHED because no registration exists.| Field | What it will hold | Value |
|---|
registration_ref | The permanent identifier this registration is cited by. | NOT YET PUBLISHED |
|---|
scope | The instruments and the measurement window the declaration binds. | NOT YET PUBLISHED |
|---|
declared_utc | When the declaration was posted — necessarily before any data was touched. | NOT YET PUBLISHED |
|---|
first_data_utc | When collection began. This must fall after the line above, and the ordering is the evidence. | NOT YET PUBLISHED |
|---|
thresholds | Every decision threshold, fixed while the answer was still unknown. | NOT YET PUBLISHED |
|---|
exclusion_rules | What will be dropped from the sample, and on what stated grounds. | NOT YET PUBLISHED |
|---|
completeness_bar | The minimum sample below which the honest null publishes instead of a number. | NOT YET PUBLISHED |
|---|
document_sha256 | The content hash of the registered document itself. Any later edit changes it, visibly. | NOT YET PUBLISHED |
|---|
amendments | Appended corrections, each with its own date and reason. The original text stays. | NOT YET PUBLISHED |
|---|
What a Honesty Stack registration must fix in advance. The lists below are classes of declaration, not declarations. They name the decisions that have to be made before the data is touched, because each one is a decision that could otherwise be made afterwards, in the direction that flatters the result. No value below has been registered.
Thresholds
- The false-discovery-rate level applied across the whole registered trial book.
- The number of walk-forward passes after which a hypothesis is treated as consumed.
- The evidence rule that makes a fault code Active rather than merely declared.
Exclusion rules
- What counts as data contact, so an ordering inversion has one unambiguous definition.
- What counts as a distinct hypothesis, so a reshaped one cannot re-enter as new without its lineage.
- Which ingestion paths are instrumented, and how coverage outside them is reported rather than assumed empty.
Completeness bars
- The minimum trial-book size below which a corrected result is not published.
- The retention window for the ledger history the append-only audit reads.
- The evidence required before a receipt counts as a receipt rather than an assertion.
Immutability, stated before it is tested. Once a registration is posted it is not edited. If it is wrong, an amendment is appended beneath it carrying its own date and the reason for the change, and the original text stays where it is, readable, above the correction. A registration that quietly improved after the data arrived would be indistinguishable from one that was right all along — which is precisely why the append rule is written here, now, while there is nothing yet to be tempted by.
Amendment rule, stated in advance: a posted registration is never edited. An amendment is appended beneath the original carrying its own date and its reason, and the original text stays above it, readable. This page will show both.
Proof corpus
The corpus is the set of artifacts a published measurement ships with — not a description of them, the artifacts themselves, downloadable, each with the digest that proves you received the bytes I measured and the code tag that produced them. A result whose inputs cannot be downloaded is not a receipt; it is an assertion with better typography.
The corpus for this suite is empty. Every row below is a slot, and every slot is NOT YET PUBLISHED. The table is printed anyway, because a reader should be able to see the exact shape of what will arrive — and because a page that described a corpus without showing how empty it currently is would be making the claim it exists to refuse.
Five columns: artifact, contents, content hash, code tag, download. Scroll sideways if they do not all fit.
Proof corpus for the Honesty Stack suite: the seven artifact classes a published receipt carries, what each will contain, and its current state. Every content hash, code tag and download reads NOT YET PUBLISHED, because no artifact from this suite's proof corpus has been published yet.| Artifact | What it will contain | Content hash | Code tag | Download |
|---|
| Registered methodology document | The versioned document these definitions are taken from, in the exact form it was registered — estimators, parameters, and the limits stated above. | SHA-256NOT YET PUBLISHED | NOT YET PUBLISHED | NOT YET PUBLISHED |
|---|
| Pre-registration record | The dated declaration: thresholds, exclusion rules and completeness bars, plus any amendments appended beneath the original with their reasons. | SHA-256NOT YET PUBLISHED | NOT YET PUBLISHED | NOT YET PUBLISHED |
|---|
| Append-only ledger export | The registry as it stands, including the entries that were killed. A ledger of survivors is exactly the artifact this discipline exists to make impossible. | SHA-256NOT YET PUBLISHED | NOT YET PUBLISHED | NOT YET PUBLISHED |
|---|
| Result set | Per-metric results with intervals and effective sample sizes, and the honest nulls wherever a sample could not support a metric. | SHA-256NOT YET PUBLISHED | NOT YET PUBLISHED | NOT YET PUBLISHED |
|---|
| Can-fail proof transcript | For every test in the battery: the planted defect, the refusal that was expected, and the outcome that was observed. | SHA-256NOT YET PUBLISHED | NOT YET PUBLISHED | NOT YET PUBLISHED |
|---|
| Kill-ledger extract | Hypotheses registered against this suite and killed by the data, each with the run that killed it. Published with the same visibility as a registration. | SHA-256NOT YET PUBLISHED | NOT YET PUBLISHED | NOT YET PUBLISHED |
|---|
| Generation code | The tagged commit that produced the result set, with its build receipt. Named here because a result whose code version is unstated cannot be re-run. | SHA-256NOT YET PUBLISHED | NOT YET PUBLISHED | NOT YET PUBLISHED |
|---|
Corpus state, per instrument
Per instrument, so that the emptiness cannot hide behind a suite-level summary.
Changelog
Append-only, dated, written by a person. It records changes to this page — it is not a measurement log, and it will not become one. A correction is a new entry that names the entry it corrects; nothing here is ever edited away, because a changelog you can rewrite is a marketing surface with a monospace font.
2026-08-01 · v1.0
Receipt page established for the The Honesty Stack suite, carrying all five parts of the §8.2 anatomy: the versioned methodology with a formal definition, estimator, failure modes and non-claim for each of the 17 metrics its 4 instruments measure; the pre-registration structure with no registration in it; the proof-corpus table with no artifact in it; this changelog; and the independent-recomputation procedure. Supersedes the earlier per-suite methodology summary at this URL, which carried the battery outline without the receipt anatomy. No measurement, pre-registration, artifact or hash accompanies this version — every receipt slot below is empty as a matter of fact, not of omission.
Independent recomputation
Independent recomputation is what makes the rest of the page checkable rather than merely well-written. It is the exact sequence a stranger follows to re-derive a published number from this suite using only artifacts I published — no account, no request, no conversation with me at any point.
Today the procedure terminates at step 1, because no measurement from the Honesty Stack suite has been published and there is no receipt to open. The steps are written now, in the specific form they will take for this suite, precisely so that they exist before the first result does and cannot afterwards be shaped to fit one.
Open the receipt and take its four identifiers.
Every published measurement links a receipt carrying four: the registration reference, the methodology-document digest, the input-manifest digest, and the generation-code tag. If any one is missing, stop — the result is not recomputable and should not be treated as though it were, including by me.
Verify the methodology document against its digest.
Download it, hash it, compare. A mismatch means the method you are about to apply is not the method that was registered, and everything after this step would be measuring a different thing.
Check the ordering before you check anything else.
The registration timestamp must precede the first-data timestamp on the manifest. If it does not, the registration is a description of a result rather than a constraint on one, and no statistic downstream can repair that.
Download the append-only ledger export and verify its digest.
The export includes the entries that were killed. A ledger of survivors is the artifact this whole discipline exists to make impossible, and receiving one is itself the finding.
Apply the registered definitions yourself.
What counts as a trial, what counts as data contact, and what counts as a walk-forward pass. Each is a definition, each changes the count, and each is fixed in the registration rather than in the analysis.
Check out the harness at the code tag on the receipt and replay the ledger through it.
Replay is deterministic: the ledger is the input, and the ordering inside it is the evidence being checked.
Recompute the trial count, the correction across the whole book, the pass counts and the ordering checks.
Every ingestion is compared against its registration. Any inversion the replay finds is a burned door, and it publishes as one rather than being quietly resolved.
Run the can-fail proof.
Plant a registration whose data timestamp precedes it, and confirm the harness flags the inversion. A burned-door detector that has never fired has not been shown to work.
If your number differs, the difference is the finding.
Send it with your inputs and the version you ran. A confirmed discrepancy publishes as a correction appended beside the original — and the original stays exactly where it is, unedited, because the error is the part of the record that proves the discipline is real.
The point of publishing this before there is anything to check: a recomputation procedure written after a result is a procedure written by someone who already knows which steps would be inconvenient.