Every implemented rule · replayed over up to a century

The first answers with power behind them.

Until now no rule in this project had more than about three independent observations — the graded ledger re-made the same call over overlapping windows inside five years. This replay runs every implemented rule over the research universe, back to 1927 where the data goes, spacing its calls so that no two on one instrument share a forecast window, and prices each against the grader's own measured null.

exploratory

Every rule, its lift, and what its evidence could see

Each point is the rule's realised lift over its own null. The bar around it is the smallest lift its independent evidence could distinguish at the family bar — anything inside the bar cannot be told from chance. A point outside the bar, to the right of zero, would be an edge. Hover a row for the arithmetic.

above its null (single test — chance puts ~2.5% here) no edge above the bar below its null (single test) cannot tell yet

Is the null itself honest?

If these rules carry no edge and the grader's chance models are calibrated, the rules' z-scores should look like draws from a standard normal: centred on zero, spread of one, about one in twenty beyond ±1.96. A shifted centre would mean every rule was being priced against the same wrong chance. That would invalidate every row above, so it is measured rather than assumed.