How I Beat Claude at Protein Design for $677
Anthropic reported 12 binders among 150 designs against TNFα, with an estimated compute cost of $27,100. In my campaign, seven of ten designs bound, and the compute bill was $677.
The campaigns used the same target, contract lab, assay, month, and compute provider. Claude was involved in both. The main difference was not that my pipeline generated better starting backbones. It was able to improve ordinary ones instead of discarding them when their sequences changed.
That distinction matters more on TNFα than it would on an easy target.
Why TNFα normally requires more sampling
TNFα is a homotrimer, and its receptor-binding surface runs along the seam between two subunits. DeepMind described the site as flat and highly polar, with relatively few hydrophobic contacts available for binding.
Before August, four published campaigns had tested 169 designs against human TNFα. AlphaProteo reported 0 of 54, PXDesign 0 of 20, BoltzGen 0 of 35, and Anthropic’s earlier Mythos Preview 0 of 60. Anthropic’s new result, with 12 binders among 150 designs, was the first published non-zero campaign.

One response to a difficult surface is to generate and screen more backbones. That is reasonable: if only a small fraction begin in a useful pose, broader sampling increases the chance of finding one.
My campaign spent relatively little compute that way. Backbone generation accounted for 14% of the total, while sequence optimization accounted for 46%. The pipeline relied entirely on public software: PXDesign, RFdiffusion3, and Complexa for backbones; ProteinMPNN for sequences; public hotspot and clash filters; and Boltz-2 for screening. I used mosaic, Escalante Bio’s JAX framework, to combine the models, loss terms, and optimizer. Mosaic provided the framework, but not the trunk-state injection described below.
The actual compute bills were $27,100 and $677, a 40-fold difference. The hardware was different, so this does not imply a 40-fold difference in productivity per GPU-hour. At Anthropic’s published rate of $4 per H100-hour, my 353 GPU-hours would have cost $1,412, or about one nineteenth of Anthropic’s estimated budget.
The lower cost therefore did not come from a private model or from replacing optimization with a cheap filter. It came from doing more with each backbone after it had been generated.
Why sequence optimization loses good backbones
Hallucination-based binder design usually begins with a random binder sequence. The sequence and target are passed through a structure predictor, gradients update the sequence toward a better score, and the process repeats. Because the initial contact is accidental, standard hallucination gives limited control over where on the target the binder docks or what binding pose it adopts.
But the structure predictor normally starts each iteration with a fresh internal state. Once the sequence changes, the next prediction may place the binder on another part of the target or give it a different fold. The score can improve even though it now describes a different structure from the one that was selected at the beginning.
This limits how far a sequence can be optimized. If the intended pose disappears after a few updates, the pipeline depends on finding a backbone that is already close to a solution. More backbone sampling compensates for the short optimization range.
My implementation still began sequence optimization from a random sequence. What it inherited from the existing target–binder complex was not the sequence, but Boltz-2’s trunk representation. I saved that representation before optimization, injected it into the first prediction as its initial recycling state, and then passed the updated trunk state from one step to the next. The model therefore retained information about the binder’s position and structure while the sequence changed.
The production run used 115 optimization steps, divided into phases of 50, 50, and 15. By the end, the optimized sequences differed from their original candidate sequences at about half of their positions, 57% in these runs, while the overall pose and topology remained similar to the starting design.
The method does not make a poor backbone good by itself. It gives the optimizer more room to search for a sequence that supports the selected interface. A plausible starting backbone can therefore be refined instead of being replaced by another round of generation.
The same experimental order included a direct comparison. Ten designs used the same generators, seeds, hotspots, and selection procedure but omitted this optimization stage. Both groups were tested on the same plate and day. One of ten designs bound in the comparison arm, compared with seven of ten in the optimized arm.
This result supports the idea that preserving the pose made the existing backbones more useful. It also introduces a separate concern: a predictor may remain confident simply because its own internal state is being passed from one iteration to the next.
Preventing the loop from confirming itself
Confidence inside a recurrent optimization loop is not an independent measure of molecular quality. The predictor may preserve a pose without producing a sequence that another model would place in the same structure.
I used ProteinMPNN to constrain this part of the search. Mosaic includes an inverse-folding loss that rewards sequences compatible with the current backbone. Its default weight was 10. I initially expected that weight to restrict the search too strongly, so I lowered it to about 1.5.
The optimization itself appeared normal. Its loss decreased and the runs completed. The problem became visible only during downstream evaluation: independent co-folding engines did not identify candidates worth ordering.
When I restored the ProteinMPNN weight to 10, candidates again passed the downstream evaluation. ProteinMPNN can also be exploited by an optimizer, so this is not an independent experimental validation. In these runs, however, the stronger inverse-folding constraint produced sequences whose predicted structures were more consistent across separate co-folding engines.
The two parts of the optimization serve different purposes. Injecting the trunk representation into successive steps keeps the binder near the original position and structure as its sequence changes. The ProteinMPNN term favors sequences that remain compatible with that structure. Without the first, optimization cannot move far from the starting sequence before the pose changes. Without the second, it can preserve a structure that other predictors do not recover.
This is a modern implementation of an older design principle. In the 2003 Top7 paper, Brian Kuhlman and David Baker cycled between sequence design and backbone optimization because backbones generated without considering side-chain packing did not necessarily have low-energy sequences. Their solution was to optimize sequence and structure together.
What the comparison shows
I chose the epitope myself by inspecting the structure and surface hydrophobicity. Anthropic’s experiment tested whether Claude could choose where to design autonomously; mine did not. The comparison is therefore between two complete design procedures, not between identical autonomous agents given different compute budgets.
The same optimized-versus-unoptimized comparison on PD-L1 produced four binders out of ten, compared with one out of ten in the unoptimized arm. The effect was smaller than for TNFα, and two targets are not enough to establish general performance.
These are also binding results rather than functional results. I have not run a TNFR1 competition assay, so the data do not show that any design inhibits TNFα. The Anthropic campaign and mine used different ACROBiosystems human TNFα SKUs. Those differences limit the comparison between campaigns, although they do not affect the seven-of-ten versus one-of-ten comparison measured on the same plate in my experiment.
All twenty sequences, their per-engine scores, and the binding kinetics are in the data folder for this post under ODC-BY. Anthropic’s numbers can be reconstructed from its released dataset.
The practical difference in this campaign was simple. Instead of spending most of the budget searching for a backbone that already had the right sequence nearby, the pipeline kept the intended pose long enough to search further around each backbone. The comparison arm suggests that this changed what reached the plate: one binder out of ten without the optimization stage, and seven out of ten with it.
Both cost figures exclude wet-lab work. This trunk-conditioning method is patent pending.