Skill scorecard · latest scored step
Is the signal real? Information Coefficient across signal dates
One forecast date is one observation. ICIR is mean IC ÷ its standard deviation; the t-stat asks whether mean IC differs from zero. Only non-overlapping dates count, and universes and modes are never pooled.
Quintile returns · latest step
Names sorted into five equal groups by forecast. If the model ranks well, realized returns rise from Q1 to Q5.
Ranking quality by horizon step
Cross-sectional rank correlation between forecast and realized return. This is what a basket strategy depends on, and a constant level bias cannot flatter it. Zero is no relationship.
Interval calibration by horizon step
Share of names whose realized return fell inside the model's 10th–90th percentile band. The target is 0.80.
Ranking quality by size tier · latest step
Where the ranking works and where it does not. Core-eligible names are the pool the overall and sector baskets draw from.
Every elapsed step
Returns and errors in percentage points; skill and correlation are ratios. MPE/MAPE/MSPE are computed on relative price levels, where they are stable.
What these measures are, and what "good" looks like
- Information Coefficient (IC)
- Rank correlation between forecast and realized return across names on one date. The standard measure of a cross-sectional signal. Commonly cited (Grinold & Kahn): about 0.05 is good, 0.10 very good, and much higher is rare enough to suspect an error.
- ICIR and IC t-stat
- Consistency of IC over time. Needs many signal dates. A t-stat above 2 is the conventional bar; Harvey, Liu & Zhu (2016) argue for about 3 once many signals have been tried.
- Out-of-sample R² vs random walk
- 1 − model squared error ÷ squared error of forecasting zero (Campbell & Thompson 2008). Above 0 beats "no change". For individual stocks, even +0.3–0.4% a month is considered strong (Gu, Kelly & Xiu 2020); negative values over a few sessions are normal.
- Hit rate
- Share of names whose forecast had the right sign. 50% is a coin flip.
- Quintile spread
- Mean realized return of the top fifth of forecasts minus the bottom fifth — the classic factor test. Positive and rising steadily from Q1 to Q5 is what a working ranking looks like.
- 80% interval coverage
- Share of realized returns inside the 10th–90th percentile band. Should be close to 80%: much higher means bands too wide, much lower means too narrow.
- Interval score
- Gneiting & Raftery (2007): band width plus ten times any miss. Rewards bands that are narrow and right. Lower is better; compare between forecasts, not against a fixed bar.
- Pinball loss
- The standard loss for quantile forecasts, averaged over the 10th, 50th and 90th percentiles — the loss Chronos is trained on. Lower is better.
Aggregates per universe × mode × basket. Only complete periods are evidence; live rows are marks.