Every forecast is scored against what actually happened. In public.
Most published “track records” are curve-fit hindsight. CutPoint's models are trained only on data that was publicly knowable at each moment — the join key is feature_available_at, not the observation date — then walked forward through history fold by fold. The model card shows every fold, every prediction, and every miss, including the days the 80% band failed to contain reality.
It's the standard a desk quant would hold an internal model to — published, with the receipts.