Finding 9 · 2026-08-16 (later)
Metrics that ranked ties by input order — a bug our own data happened never to trigger
A constructed test found that EF, BEDROC and PPV broke ties by the order compounds were listed in, not by value. Identical data, reordered, moved EF@10% from the theoretical maximum to zero. AUC never showed it, because AUC already handled ties correctly.
What we expected
Enrichment metrics (EF, BEDROC, PPV) should not depend on the order compounds are listed in.
What happened
On a constructed test — 100 compounds, 10 actives, one constant score — actives listed first gave EF@10% = 10.0 (theoretical maximum); actives listed last gave 0.0. Identical data, and AUC read 0.5 both times because auc_roc already handled ties and hid the problem.
Why it happened
EF, BEDROC and PPV ranked ties by input position rather than breaking them in a way invariant to ordering.
What changed
The team notes their own historical numbers were never actually corrupted by this, because Boltz-2’s probabilities are tie-free — but concludes that is not a defensible property for an evaluator to rely on. Commit 18fc86e. A second, unrelated bias was also found and fixed in the same pass: averaging rank before exponentiating is not the same as the true expectation, validated by Monte Carlo at 0.1157 vs 0.1142 over 400 random rankings.
“‘correct only if the customer’s model emits continuous floats’ is not a property an evaluator can have.”
Bring us the number you are least sure about.
That is usually the one worth thirty minutes.