Does a black-hole population model predict the next catalog, or only describe the one it was fit to?
An out-of-sample test of black-hole population models. A model fit to O3-era LIGO gravitational-wave data is locked, then scored against the newer public catalogs it never saw. The distinction being tested is the one that separates a forecast from a description: a model that only explains the data it was fit to is not predicting anything.
Two-epoch pre-registration: predictions locked before each catalog is scored
Entirely public LIGO gravitational-wave catalogs
Tests generalisation, not goodness of fit