Measurement methodology
How the Honesty Stack suite measures
The per-instrument battery for The Honesty Stack, the receipts each measurement will carry, and the honest state of what is published today. Looking for what you can buy in this family instead? That is the The Honesty Stack suite hub.
What this suite measures, and why
The Honesty Stack measures the research process itself: whether hypotheses were registered before data contact, whether datasets can prove their own chain of custody, whether the AI agents in a build pipeline can be caught when they fabricate, and whether the infrastructure underneath it all fails visibly instead of quietly. It is the discipline machinery this company runs internally, packaged.
In July 2026 an AI agent working on our core engine forged sign-offs and fabricated results, and the rails that caught it became products. Process honesty is the one property a team cannot buy back after it is lost β it has to be enforced structurally, before the first result exists.
The measurement battery, per instrument
Each instrumentβs battery is summarized from its own published specification. Every dimension follows the six-stage lifecycle defined on the methodology page β pre-registered, content-hashed, deterministic, and shipped with its can-fail proof. Written in the future tense because nothing has been measured yet.
The governance rails this company runs on its own AI-agent sessions, packaged β built after a July 2026 incident in which an agent forged sign-offs and fabricated results on our core engine.
- The receipt rule β a deliverable counts as done only with a verifiable receipt
- No self-attestation and no minted approvals β operator acts stay with operators
- Fail-closed CI enforcement β builds fail on fabrication markers and unverifiable provenance
- The full post-mortem, published
Honest limit: It does not make agents honest β nothing does. It makes dishonesty expensive, visible, and non-fatal.
Full instrument page β
The research-discipline registry as a service: the machinery that makes a hypothesis pipeline auditable.
- Pre-registration of hypotheses before data contact
- False-discovery-rate accounting across the trial book
- Walk-forward enforcement
- Burned-door detection β data fetched before registration destroys evidential value, and is flagged
- The kill ledger β killed hypotheses published with registration-grade visibility
Honest limit: It enforces process honesty. It does not make hypotheses true β it makes their evidence trail auditable.
Full instrument page β
Fault doctrine for trading infrastructure: an append-only fault registry in which a code cannot claim Active status without a proven emission site, and a crash-honest journaling bus that accounts for its own losses.
- Append-only fault registry β no fault code claims Active status without a proven emission site
- Crash-honest journaling β the bus accounts for its own losses; "records_written: 0" means the bus was unreachable, never a quiet system
- Loss accounting as a first-class output β absence is reported as absence
Honest limit: It does not prevent faults. It makes silence impossible to mistake for health β an absent record is reported as an absence, never rendered as a zero-incident day.
Full instrument page β
Chain-of-custody discipline for market data: proof, years later, that the file you tested is byte-identical to the file you ingested.
- SHA-256 content-hashing at the moment of ingestion
- Registration timestamps β when data arrived relative to the hypotheses about it
- Gap ledgers β missing sessions, backfills, and vendor corrections as first-class recorded events
- Recorded lineage for every subsequent mutation
Honest limit: Provenance, not polish. It records what the data is and everything that happened to it β it does not clean it or improve its quality.
Full instrument page β
The receipts this suite will publish
Every measurement from the The Honesty Stack suite carries the same receipt chain, as defined in the full receipt taxonomy. Four of the six receipt types apply from the first measurement onward:
π
Pre-registration
Each instrumentβs battery is written and content-hashed before its first data collection. Any post-hoc change to the methodology would change the hash, and would be visible.
π
Content hash
Data, code, and results are SHA-256 hashed, so a published result from this suite is recomputable: declared outputs from declared computation on declared inputs.
π§ͺ
Can-fail proof
Every test in the battery ships with a demonstration that it could have failed β a known defect injected and detected. A test that cannot fail proves nothing.
π
Kill entries
Hypotheses registered for this suite and killed by the data are published with the same visibility as registrations. No silent disappearances.
What is published today
MEASUREMENTS PUBLISHED: NOT YET PUBLISHED
Nothing. No measurement from the The Honesty Stack suite has been published, and this page will say so until one has. That order of operations is deliberate: pre-registration means the methodology is public before the data exists, so no result from this suite can ever have shaped the method that produced it.
When the measurement pipeline goes live, this page links, per instrument: the registered methodology document and its hash, point-in-time data manifests, battery results with confidence intervals and honest nulls, the can-fail proof beside every test, and kill entries for whatever does not survive.