Method
Evaluation and the memorisation control
How runs are scored, and why a solved problem sits in the problem set.
Runs are not scored on whether they solve anything. They are scored on whether the increment is checkable, whether the provenance labels survive inspection, and whether the identified obstruction matches the one the literature documents.
Why a solved problem is in the set
The Poincaré Conjecture was settled by Perelman in 2003 and its proof is thoroughly represented in any plausible training corpus. That makes it useless as a target and invaluable as a control: it is the one problem where a model can succeed entirely by recall, so it is the one place we can measure the gap between reconstruction and recitation.
A run is scored as reconstruction only if it derives the entropy monotonicity functional rather than quoting it, explains why cigar solitons must be excluded, and correctly identifies which steps fail in dimension four. Recitation is common; reconstruction is not. Everything else in the problem set is calibrated against that ratio.
What would count as signal
Not a proof. The realistic positive outcomes are narrower: an obstruction rediscovered without being prompted with it, a correct judgement about which of two dead ends is less dead, or a reformulation that a working mathematician finds worth pursuing. Each is a weak signal individually. The measurement is whether they occur more often than chance across a large number of runs.
Selection effects
Runs are never re-rolled. Publishing only the interesting transcripts would make the archive a selection artifact and destroy its value as a measurement, so the first pass is what gets published — including the large majority that restate the problem and stop.