Both results were valid inside their own experiments. They were not interchangeable.
The held-out evaluation and the operational replay now have separate claim records, each
with its own denominator, threshold context, precision, and recall. The site will never
borrow the stronger number while describing the other dataset.
Attached measurements
The denominator stays with the number.
92.06% precision
On an 85,429-row held-out test set, the UPI model reached 92.06% precision and 12.81% recall at a 0.5% alert budget.
85,429 held-out transactions
75.22% replay precision
Across a 22,071-transaction seven-day replay, the UPI system recorded 75.22% precision, 12.13% recall, and no alert-budget violations.
22,071 transactions across 7 days
Source boundary
What this record was checked against
01artifact
Reviewed internally. The repository or artifact is private, so its path is not published.
02artifact
Reviewed internally. The repository or artifact is private, so its path is not published.