45 / 97
Information Coefficient (IC)
Written IC throughout.
Working definition
The correlation between a forecast and the outcome it predicted — the standard scalar summary of how accurate a signal was, measured across a cross-section or through time.
The information coefficient is the correlation between what a signal said would happen and what happened. Computed across a cross-section at each date, then summarised through time, it answers a narrow question precisely: did the ordering the forecast implied match the ordering the market delivered?
Two variants are in common use. The Pearson IC correlates raw forecast values with raw outcomes, and is sensitive to outliers on both sides — one extreme return can set the number. The rank (Spearman) IC correlates orderings instead, and because most signals are used to rank rather than to size, it usually describes the intended use better.
The statistic reaches portfolio outcomes through the fundamental law of active management, which relates an information ratio to the information coefficient multiplied by the square root of breadth. That relation is a useful piece of intuition and a dangerous one: breadth means independent bets, and a portfolio of five hundred positions driven by one factor has a breadth far closer to one than to five hundred — the effective sample size problem in portfolio clothing.
An IC is also a pre-cost quantity. A signal can order a cross-section correctly and still lose money once execution cost, market impact and financing are charged, particularly where its accuracy concentrates in the fastest-decaying, highest-turnover part of the horizon. And like any correlation estimated from a finite sample it carries a standard error; small coefficients are normal in liquid markets, which makes separating a small real one from zero a question about sample size rather than about the number.
Why it matters
The IC is the cleanest available separation of forecast quality from everything else — sizing, costs, risk model, luck in a single path. That makes it valuable as a diagnostic and misleading as a headline. A signal with a stable, modest coefficient across regimes is a different object from one whose average is identical but earned entirely inside a few months, and only the time series tells them apart.
Commonly confused with
Neighbouring concepts that get used interchangeably, and the distinction that actually separates them.
- Sharpe ratio
The IC grades predictions; the Sharpe ratio grades outcomes of a strategy that was actually run. A signal can order a cross-section correctly and still lose money once costs are charged, which is exactly the gap between the two statistics.
- Pearson versus rank IC
Pearson correlates raw values and is sensitive to outliers on both sides — one extreme return can set the number. Rank correlates orderings. Since most signals are used to rank rather than to size, the rank version usually describes the intended use better.
- Hit rate
A hit rate counts how often the direction was right. An IC measures whether the ordering matched, which is a stronger and more useful claim for a cross-sectional signal — you can be right more often than not and still rank the winners below the losers.
- Breadth
The fundamental law relates an information ratio to the IC times the square root of breadth, and breadth means independent bets. Five hundred positions driven by one factor have breadth far closer to one — the effective sample size problem in portfolio clothing.
How to measure it in your own data
A definition you cannot test is a definition you have to take on trust. This is the shortest honest route from the concept to a number you computed yourself.
- Records you need
Forecasts with the timestamp at which they were made, the realised outcomes over the intended horizon, and the cross-section they were meant to rank. Forecasts reconstructed after the fact do not qualify.
- What you compute
Correlate forecast against outcome across the cross-section at each date, then summarise that series through time. Prefer the rank version unless the raw magnitudes are genuinely being used, and report the time series rather than only its mean.
- What the answer tells you
Small coefficients are normal in liquid markets, so separating a small real one from zero is a question about sample size rather than about the number itself — it carries a standard error like any correlation from a finite sample. And the mean hides the thing that matters: a signal with a stable modest coefficient across regimes is a different object from one whose identical average was earned entirely inside a few months.
Questions and answers
Should I use Pearson or rank IC?
Rank, in most cases. Pearson correlates raw values and one extreme return can dominate the estimate on either side. Since a signal is usually used to order a cross-section rather than to size positions by its raw magnitude, the rank version measures the use to which the forecast is actually being put.
Does a positive IC mean the signal makes money?
No, because an IC is a pre-cost quantity. A signal can order a cross-section correctly and still lose once execution cost, market impact and financing are charged — particularly where its accuracy concentrates in the fastest-decaying, highest-turnover part of the horizon, which is where trading it is most expensive.
What counts as a good IC?
Small coefficients are normal in liquid markets, so the headline value alone settles very little. The question that matters is whether it is distinguishable from zero given your sample, and whether it is stable across regimes — which requires the time series rather than its average.
Why is breadth so often overstated?
Because breadth in the fundamental law means independent bets, and a portfolio counts positions. Five hundred holdings driven by a single factor behave as very nearly one bet, so the square-root-of-breadth term promises a multiple of information ratio that the portfolio cannot deliver.
Related terms
Derived from the links this entry makes and the entries that link back to it.