Anthropic just ran a real science experiment. It pointed Claude at 15 disease targets and told it to design proteins that stick to them, with no human guidance during the run. Independent labs then made the molecules and tested them. Claude’s designs worked on 14 of the 15 targets, hitting success rates of 22 to 35 percent against an industry norm of 10 to 15, and on one target it beat every human entrant in a public competition, 40 percent to 3.7.

What actually happened

Over a 48 hour session Claude produced 1,320 designs, and 354 of them turned out to be confirmed binders once Adaptyv Bio and Twist Bioscience synthesized and tested them in the wet lab. No prompt by prompt hand holding. The model ran the campaign, chose what to try, and iterated on its own. Anthropic is honest about the ceiling: a binder is not a drug, it is the first step of many. But this is a real, lab validated result, not a benchmark.

Why builders should care

This is the shift worth watching. The story of AI in 2026 stopped being “the model wrote good code” and became “the model ran an expert workflow end to end and beat the specialists.” Protein design today, but the pattern is general. Give an agent a clear goal, real tools, and a way to test its own output against reality, and it can close the loop without you. If you build agents, that is the lesson. The teams winning are not the ones with the cleverest prompt, they are the ones who wired their agent to a real world feedback signal so it can grade itself and improve. Find the loop your agent can close on its own.