The Score Most People Use to Rank Protein Binders Doesn’t Predict Binding
For the 2026 EGFR/cetuximab binder optimization competition, we developed a new in silico method to avoid co-folding confidence metrics such as ipTM and ipSAE. Why? Because when ranked by ipSAE, the tightest binder the wet lab validated (which our in silico method also ranked first) would not even have made our top-ten design shortlist.
The standard way to rank a designed protein binder is to co-fold it against its target with an AlphaFold-class model, read off confidence scores such as ipTM and ipSAE, and send the highest-ranked designs to the lab. For the 2026 EGFR/cetuximab binder optimization competition, we designed a new in silico method specifically to avoid ranking designs based on those scores. The result? All ten designs we sent to the lab bound their target, and three of them measured tighter than the previous record. The one we ranked first by our in silico method proved to be the tightest binder measured, binding 43.9-fold more tightly than cetuximab and 1.6-fold more tightly than the previous best binder.1
All ten designs we submitted for testing bound their target.
Each bar represents one design, measured against cetuximab, which binds at 9.94 nM. Higher bars indicate tighter binding. The four gray bars on the left are the references: cetuximab, the best binder from Adaptyv Rounds 1 and 2, and the previous best binder on this target.2 The dashed line extends the previous-best result across the chart, so a design rises above it only if it exceeds the existing record. The ten designs follow, ordered by their experimentally measured binding affinity, with the tightest binder first.
The designs span a 622-fold range, so the bar heights are compressed and the number printed on each bar shows the actual ratio. The label under each bar shows the rank that our in silico method assigned to that design before any measurement, which is why the order reads #1, #9, #6, #4, #3, and so on, rather than 1 through 10. Hovering over a bar reveals the design name and the results of both SPR replicates.3
The interface score (ipSAE) would have excluded our best design.
We co-folded every design against EGFR domain III with Boltz-2 and recorded its interface score, ipSAE. Below, the height of each bar is that design's ipSAE score.
The designs are in the same order as the chart above, tightest binder first. If ipSAE tracked binding, the tightest binders on the left would carry the highest scores, but they carry some of the lowest instead, and several weaker designs score higher. Across these ten designs, ipSAE shows correlations of −0.48 with measured binding affinity using Pearson’s correlation and −0.54 using Spearman’s correlation.4 The score points away from binding, not toward it.
Shortlisting designs based on that score would have picked the wrong designs. Ranked by ipSAE across the nineteen designs we co-folded, Interlit #1, the tightest binder measured in the competition, sat thirteenth, and a top-10 cutoff based on that score would have dropped it before it reached the lab.
We did not rank designs based on that score, a decision made before the lab experiments were conducted. Our pipeline weights every metric it touches by how much trust that metric has earned from comparisons with previous documented wet-lab data from the literature. By that measure, co-folding confidence has earned little trust for predicting affinity, so our pipeline used it as a viability gate, not in our ranking system.
The wet-lab measurements supported that decision: ipSAE ran opposite to measured affinity, so ranking by it would have dropped our tightest binder. This is the case for showing the reasoning behind a ranking rather than reporting a single score, because the same metric can be trustworthy as a gate but misleading as a ranking, and its value alone does not tell the full story.
Our ranking got the top right but the middle wrong.
Our #1-ranked design was the tightest binder measured. That is the result worth having.
That all came out of a single day of work. Admittedly, the rest of the ranking was noisier than the headline suggests. Our #2-ranked design showed only 0.3 times the binding affinity of cetuximab, meaning it binds about three times more weakly than the drug we started from. The designs ranked #5 and #7 had measured affinities of approximately 140 nM, making them roughly fourteen times weaker than cetuximab. Two of the three designs that bound more tightly than the previous best sat at ranks #6 and #9, near the bottom of our own list.
The #1 and #2 split deserves a closer look. What separated our first pick from our second was a tiebreaker, and that tiebreaker placed a design that binds three times worse than cetuximab in second place.
Conclusion
- This is not a universal law. The above-mentioned learnings and conclusions are based on only ten designs tested against a single target, so the sample size is small. However, our in silico method in principle should be generalizable to all protein binder design applications.
- Co-folding confidence is not useless. It is useful as a filter, not as a ranking metric, which is exactly how our pipeline used it and what our trust table predicted.
- Our ranking system is not perfect. We knew which metrics not to trust, which helped us rank the best design first. However, correctly predicting the full order is a much harder problem: predicting binding affinity remains one of the great unsolved challenges in the field, a holy grail that we are chasing.
- Affinities measured independently by Adaptyv Biosystems by SPR, two replicates per design, mid-July 2026. The cetuximab scFv we started from binds at 9.94 nM. Our best design measured 0.2263 nM with a 95% confidence interval of 0.1287 to 0.3978 nM. ↩
- Best measured binder on this target, from the public ProteinBase collections: 6.6 nM in Adaptyv round 1 (vast-bear-marble), 1.2 nM in round 2 (shy-shark-cypress), and 0.37 nM in the earlier Cradle EGFR competition (scarlet-seal-granite). The "previous best" is that 0.37 nM Cradle result, the tightest recorded before ours. ↩
- Bar heights carry a minimum so the weakest designs stay visible, which means height is indicative and the printed number is exact. The two designs below 1x bind more weakly than cetuximab, at 140.6812 nM and 139.9896 nM. ↩
- ipSAE versus measured fold-change over the ten submitted designs: Pearson r = −0.48, Spearman rho = −0.54. Scores are the interface aligned-error values recorded by our optimization pipeline. Interlit #1 ranks thirteenth by ipSAE among the nineteen candidates that reached a co-fold, so a top-10 cutoff based on that score excludes it. The ipTM columns carry no positive signal either. ↩