Honest evaluation — embargoed rolling backtests
The metrics on the Machine-learning tab are historical / diagnostic: they use the standard chronological protocol, which lets training labels share advisory events with the test window and inflates apparent performance. The numbers below are deployment-honest: the same evaluation is repeated at many points in time (an expanding-window rolling backtest), and at every fold the training labels are rebuilt using only advisories published before that fold's scoring date (the label embargo). No training label ever depends on an advisory the model is later tested on. Pooled metrics combine every fold's test predictions — hundreds of advisory outcomes across 2023–2025 rather than one two-month window.
Runs marked stored labels use the historical (leaky) labels on the same folds, shown for contrast. Produced by tools/rolling_backtest.py.
No rolling backtest results found. Run tools/rolling_backtest.py and refresh this page.