What happened

Both results were valid inside their own experiments. They were not interchangeable.

The held-out evaluation and the operational replay now have separate claim records, each with its own denominator, threshold context, precision, and recall. The site will never borrow the stronger number while describing the other dataset.