Finding 11 · 2026-08-16
An undocumented hard dependency on a free external API brought down every GPU prediction mid-campaign
Every compound in a knockout experiment failed at a constant 64.6 seconds with an error that wrongly blamed a GPU architecture mismatch. The real cause was that every Boltz-2 prediction fetches a fresh sequence alignment from a free external server, and that server was unreachable — while boltz itself exits 0 on the failure, so no one saw the true error until the shim was patched to log it.
What we expected
The pocket-knockout experiment (and the pipeline generally) should fail, if it fails, for a reason connected to the model or the GPU.
What happened
Every compound in the knockout failed at a constant 64.6s with an error message that blamed a GPU architecture mismatch inherited from an unrelated April incident. That message was wrong. The shim never logged boltz’s actual stderr because boltz exits 0 on this failure, so the operator never saw the real cause until the shim was patched to log it directly:
requests.exceptions.ConnectTimeout: HTTPSConnectionPool(host='api.colabfold.com', port=443): Max retries exceeded Exception: Too many failed attempts for the MSA generation request.
Every Boltz-2 prediction fetches a fresh multiple-sequence alignment (MSA) from api.colabfold.com — the shim passes --use_msa_server on every call — and that server was unreachable from GX10 while general internet access worked fine. This was not knockout-specific: the whole pipeline was down.
Why it happened
Nothing documented this dependency: it did not appear in the run script’s flag annotations, the watchdog, the queue’s health probe, or the campaign plan. The queue’s own reachability gate checked only the shim itself, which stayed healthy and returned 200 throughout. Separately, /cache/msa was already mounted into the shim with 25 cached MSAs sitting in it, unused — the only MSA-related line in main.py was --use_msa_server.
What changed
MSAs are now cached on disk keyed by sha256(sequence)[:16], written atomically so a killed run cannot poison a later prediction, and --use_msa_server is appended only when a fetch is actually needed. The cache was seeded from the 25 orphaned alignments already on disk. Verified with ColabFold still down: “MSA: all 1 chain(s) served from cache — no network fetch… boltz predict completed in 114.8s.” Until this landed, no new target could be onboarded — CDK2, ABL1, EGFR and every other candidate from the expansion review needed an MSA it could not currently fetch.
“A free third-party API is a single point of failure for every GPU prediction this project makes. Nothing documented it.”
Bring us the number you are least sure about.
That is usually the one worth thirty minutes.