Finding 3 · 2026-08-14
The shim silently dropped every pocket constraint it was ever sent, for four months
The Boltz-2 shim's YAML builder translated ligands, polymers and properties into the backend's input format — but never pocket constraints. The client log line printed on every call described the constraint being built, not one being applied, and 2,444 real 7AAD scores existed at the moment of the fix, every one of them unconstrained.
What we expected
A pocket constraint computed client-side (remapped from author to sequence numbering, range-checked, logged) should be applied by the Boltz-2 backend it is sent to.
What happened
services/boltz2_shim/main.py:_build_boltz_yaml() translated only polymers, ligands and properties into boltz YAML — grep -in "constraint\|pocket\|contacts" services/boltz2_shim/main.py returned zero hits. Every pocket constraint was silently discarded at the shim boundary, while the client log printed on every call:
INFO Pocket constraint for 7AAD: remapped author residues [766, 862, ...] -> sequence positions [107, 203, ...] (offset=659)
That log line described the payload being built, not a constraint being applied. The shim had been the backend since 2026-04-04. “2,444 real 7AAD scores existed at the moment of the fix, every one unconstrained.”
Why it happened
The log statement confirmed construction of the payload, not receipt or application by the server — no downstream signal ever verified the constraint actually reached the model.
What changed
_build_boltz_yaml now translates the constraint into boltz 2.x's native constraints: [{pocket: {...}}] schema, confirmed live in the shim log (“Applied 1 pocket constraint(s) to boltz input (12 contacts total)”), and the response now carries _pocket_constraint_applied / _n_pocket_contacts so a stored score is auditable after the fact.
An A/B re-score of 3 previously-scored compounds against a measured noise floor (sd=0.0094, n=3 replicates) found the constraint effect real but compound-dependent: deltas of +0.018, +0.064, +0.066, with the latter two ~7σ above the noise floor. The team ruled that pre-fix (unconstrained) and post-fix (constrained) scores must not be pooled, and the screen was stopped deliberately before deploying the fix so no run straddles the boundary.
A related bug in the same file — PAE (prediction confidence) had been inert since May because the shim never forwarded --write_full_pae — was found and fixed in the same pass.
“Every score this backend has ever produced was effectively unconstrained — while the client log printed, on every single call... That log line describes the payload being built, not a constraint being applied.”
Lab notebook, lines 4177, 4187.
This is the standard we would hold your data to.
If that sounds like what your result needs, thirty minutes is enough to find out.