Point-in-Time Data

Data recorded as it was actually known on each historical date — first-print economic releases, as-of universe membership, unrevised financials — rather than the revised series that exists only in hindsight.

Most historical datasets are the present’s version of the past. Economic series are revised: headline employment and GDP figures are restated in subsequent releases, sometimes drastically, so the “historical” value a standard database shows for a given month is a number that nobody possessed in that month. Corporate financials are restated. Index membership lists are published as they stand today, not as they stood on each date. Prices are retroactively adjusted. A backtest run on such data is simulating a market participant who traded on information from the future — politely, a look-ahead bias; precisely, a fiction.

Point-in-time data is the corrective: every value is stored with the timestamp at which it became knowable, and queries return what was knowable as of the simulated date. The first print of a release, not its final revision — because the first print is what moved the market. The universe as constituted on the day, including the members that later delisted — because their absence is exactly the conditioning that survivorship bias smuggles in. The distinction is institutionalised where the stakes are understood: central-bank archival databases preserve data vintages — the full sequence of what each series looked like at each publication date — precisely because the revised series and the tradable series are different objects.

The discipline is expensive and unglamorous. It multiplies storage, complicates every query, and produces backtests with worse numbers, since the flattering revisions and the buried failures are exactly what it removes. That is the point. A simulation on revised data measures a strategy against a memory of history — tidied, corrected, and survivor-only. A simulation on point-in-time data measures it against something resembling the history that occurred. Maintaining the distinction requires knowing what a dataset contains and when each part of it arrived, which is the concern of tick data provenance extended across every input a strategy consumes.

← All glossary terms