Finding 2 · 2026-08-16

The queue silently launched four concurrent copies of one GPU job

The campaign queue's health check is supposed to stop a second GPU job starting while one already runs against a backend that serves one prediction at a time. A step was added to the run list on 2026-08-14 but never added to the check's hand-written list of “a job is running” markers, and by the time anyone noticed, four processes were competing for the GPU.

What we expected

The campaign queue's health check should prevent a second GPU job from starting while one is already running against a backend that serves one prediction at a time.

What happened

confirm_parp1_hits was showing “attempt 2/3” with no failure logged, because there was no failure — the queue had started a second instance while the first was still running, and by the time it was noticed there were four processes (two wrapper/child pairs) competing for the backend.

Why it happened

campaign_queue.sh keeps two coupled lists — STEPS (what to run) and a hand-written needles tuple (what counts as “a GPU job is already running”). confirm_parp1_hits was added to STEPS on 2026-08-14 and never added to the needles, so every tick read the GPU as idle and started another copy. This was the same defect class already patched once for the harness pipeline, patched that time by hand-adding two more needles — the exact fragile pattern that caused this incident.

What changed

The needles are now derived automatically from the step commands by extracting script names, so a step cannot be added without becoming detectable. The fix itself initially had the same bug: the extraction regex scripts/[a-z_/]+\.py had no digits in its character class, so it silently missed confirm_parp1_hits, compute_ecfp4_baseline and audit_boltz2_outputs — the three scripts with numerals in their names, including the exact one that caused the incident. Corrected to [a-z0-9_/]+.

“Any config where correctness depends on a human remembering a second list will eventually be wrong. Derive it.”

Lab notebook, line 5762.

This is the standard we would hold your data to.

If that sounds like what your result needs, thirty minutes is enough to find out.

Book a 30-min call