Confidence and function are separate claims
Both experiments tested designs before knowing the answer, the only way to learn whether a model's confidence predicts real biology.
Confidence is not the same claim as function
Both studies tested designed proteins prospectively rather than retrospectively, the only way to learn whether a model's confidence predicts real biology.
AlphaDesign inhibitors
- Task
- Designed bacterial protein inhibitors, tested for in-vivo activity
- Sample size
- 88 designed constructs
- Hit rate
- 17/88 active, about 19%
What limited it: Expression and folding failures, not just binding geometry, separated plausible designs from active ones.
From rare hits to task-specific yields
| Experiment | Measured hit rate | What counted |
|---|---|---|
| Earlier EGFR Rosetta protocol | About 0.01% | Company-reported historical comparator, with a different campaign and selection process. |
| Adaptyv EGFR competition, 2024 | 5/201, 2.5% | Binder hit in the organiser's physical test. |
| Adaptyv second EGFR competition, 2025 | About 13% | Organiser-reported binder hits among 400 selected designs. |
| RFdiffusion paper, 2023 | 19% | Binding at least half the positive-control signal across five targets. |
| Bits to Binders, 2026 | 707/12,000, about 5.9% | Significant CD20-specific CAR-T proliferation in a pooled assay. |
This is a series of prospective experiments, not a controlled time series. Targets, selection, expression, and hit definitions differ; no single line should be extrapolated as a field-wide learning rate.
Four stages between a proposal and a product
Each stage can silently discard a design that looked correct at the stage before it.
Objective and representation
The desired geometry, interaction, dynamics, and context become a computable target.
- Measure
- Constraint coverage
- Failure boundary
- An incomplete objective can be satisfied perfectly and still miss what matters biologically.
Where the frontier moves
Objectives that encode function and context, not just static structure.
Generation and filtering
Models propose backbones and sequences, then rank structure and developability.
- Measure
- Designs/hour · diversity
- Failure boundary
- Model miscalibration can make implausible designs look confidently correct.
Where the frontier moves
Generation abundant enough that filtering, not proposing, becomes the bottleneck.
Build and assay
DNA synthesis, expression, purification, binding, and functional screens test reality.
- Measure
- Validated hits/design
- Failure boundary
- Lab throughput caps how many of the generated designs can actually be tested.
Where the frontier moves
Multiplexed assays that test thousands of sequences in parallel with function labels.
Product development
Stability, immunogenicity, delivery, manufacture, and in-vivo efficacy define usefulness.
- Measure
- Potency · yield · safety
- Failure boundary
- A design that works in a test tube can still fail translation into a manufacturable product.
Where the frontier moves
Designing for stability and manufacture before optimization even converges.
Sequence space is enormous; evidence is finite
Computation can search and compress priors, but each new function remains an empirical claim. Multiplexed experiments lower the cost of evidence without eliminating it. A commercial expression-and-binding service advertises a starting price of $99 per protein and a three-to-four-week turnaround; that is a vendor quote for one assay scope, not a universal physical floor. It excludes clinical safety, in-vivo efficacy, and often the full cost of synthesis and optimization.
Binding is easier to count than catalysis
Earlier computational Kemp eliminases had catalytic efficiencies of roughly 1–420 M⁻¹ s⁻¹, versus a cited median near 10⁵ M⁻¹ s⁻¹ for natural enzymes. A 2025 de novo metallohydrolase study reported a zero-shot design above 10⁴ M⁻¹ s⁻¹ among 96 directly tested candidates, within the range of native metallohydrolases on similar substrates. These are different reactions and assays, so the result shows progress in a harder task rather than a universal enzyme hit rate.
Translation has also begun: South Korea approved the SKYCovione vaccine in 2022 using a computationally designed protein nanoparticle, while Generate:Biomedicines reports its AI-engineered GB-0895 antibody is in Phase 3 asthma trials as of 2026. Approval of a nanoparticle vaccine and a candidate in trials are different evidence levels; neither makes an untested design a medicine.
The bottleneck moves into assays and objectives
Once generation is abundant, teams need scalable functional assays, negative-selection data, meaningful thresholds, and models that predict more than geometric plausibility.
Prospective benchmarks
Score hidden designs with standardized physical tests, not retrospective fits to known structures.
Multiplex assays
Measure thousands of sequences in parallel while preserving reliable function labels for each one.
Learn from negatives
Capture expression, toxicity, and specificity failures instead of only publishing the successes.
Design for product
Include stability, formulation, and manufacture as constraints before optimization converges, not after.
Who is building what
Structural diffusion models, multimodal biological transformers, sequence-to-function generative models, and automated high-throughput assays design proteins from scratch. Search the record, or filter by application domain.
Institute for Protein Design (Baker Lab)RFdiffusion & ProteinMPNNDenoising diffusion probabilistic models generating backbone structures directly from atomic coordinates, paired with ProteinMPNN for inverse sequence design
- Reported evidence
- Pioneered foundational open-source tools cited worldwide; designed high-affinity de novo binders targeting SARS-CoV-2, influenza, and cancer antigens.
- Announced next step
- Universal generative models designing functional binders, enzymes, and nanomachine assemblies in a single automated step.
- Unresolved risk
- Physical expressibility in E. coli or mammalian cell culture, aggregation in solution, and structural flexibility not captured by rigid models.
Google DeepMind / Isomorphic LabsAlphaFold 3Diffusion-based biomolecular architecture predicting 3D joint structures of proteins, DNA, RNA, post-translational modifications, and small molecule ligand complexes
- Reported evidence
- Demonstrated unprecedented accuracy on blind protein-ligand and protein-nucleic acid interaction benchmarks; partnered with Eli Lilly and Novartis.
- Announced next step
- De novo drug design moving from target structural prediction directly to high-affinity lead therapeutic molecules.
- Unresolved risk
- Predicting dynamic conformational state ensembles (induced-fit vs conformational selection) from single static structural predictions.
Generate:BiomedicinesThe Chroma ArchitectureGenerative biological diffusion models for antibodies, protein complexes, and peptide therapeutics
- Reported evidence
- Company-reported April 2026 pipeline: GB-0895 was in Phase 3 asthma trials; GB-4362 and GB-5267 had clinical-study plans. These are development stages, not approvals or proof that every program was designed de novo.
- Announced next step
- Demonstrate safety and efficacy for specific candidates in controlled clinical trials.
- Unresolved risk
- Clinical developability, including viscosity, formulation stability, and anti-drug-antibody immunogenicity.
EvolutionaryScaleESM3 Frontier Model98-billion parameter multimodal biological model simultaneously reasoning over protein sequence, 3D structure, and biological function tokens
- Reported evidence
- Demonstrated de novo generation of an entirely novel green fluorescent protein (esmGFP) with 58% sequence divergence from any known natural protein.
- Announced next step
- Generative biology models acting as foundational programming interfaces for synthetic enzyme and cellular circuit design.
- Unresolved risk
- Inference compute cost for multi-billion parameter models and bridging sequence tokens to multi-body biochemical kinetics.
Cradle BioGenerative Engineering PlatformModels trained on iterative wet-lab feedback to optimize protein stability, expression, and activity
- Reported evidence
- Company-reported partner work and activity improvements; no independently normalized cross-project hit-rate series is disclosed here.
- Announced next step
- Validate improvements against matched controls in partner processes.
- Unresolved risk
- Models may not generalize across enzyme families or small, sparse assay datasets.
Nabla BioHigh-Throughput Wet-Lab MLMultiplexed antibody screens coupled to predictive models
- Reported evidence
- Company-reported pharmaceutical partnerships and work on difficult membrane targets; public independent clinical outcomes and comparable hit rates are not established in this card.
- Announced next step
- Prospective antibody validation with developability and specificity profiling.
- Unresolved risk
- Assay artifacts and the gap between in-vitro binding and in-vivo pharmacokinetics.
Profluent BioOpenCRISPR-1 & Generative LLMsLarge language models trained on global protein diversity, generating fully synthetic CRISPR gene editors with distinct sequences from natural Cas enzymes
- Reported evidence
- Designed and synthesized OpenCRISPR-1, demonstrating comparable human genome editing efficiency to wild-type SpCas9 with lower off-target rates.
- Announced next step
- Complete generative design of bespoke gene-editing enzymes, molecular motors, and targeted delivery effectors.
- Unresolved risk
- Patent landscaping and intellectual property freedom-to-operate around synthetic variants of naturally occurring enzymes.
CASP / CAMEOBlind Prediction BenchmarksCritical Assessment of Structure Prediction (CASP) and Continuous Automated Model EvaluatiOn (CAMEO) conducting double-blind structural accuracy assessments
- Reported evidence
- The gold standard rigorous benchmark that crowned AlphaFold and continuously ranks global structural biology and protein design models.
- Announced next step
- Expanding rigorous blind evaluation to protein design tasks (de novo ligand binding, conformational dynamics, and enzyme catalytic efficiency).
- Unresolved risk
- Experimental crystallization bottlenecks verifying community-designed proteins in physical wet labs.
High in silico predicted binding scores and low RMSD do not guarantee biochemical expression. Solubility, thermal stability, aggregation propensity, and immunogenicity govern lab survival.
The optimistic view, with conditions
Protein design becomes an experimental compiler
Models will translate functional specifications into small, diverse libraries whose measured failures and successes continuously improve the system.
Test prospectively, always
Retrospective structure metrics do not substitute for measuring designs the model has never seen.
Scale the assay, not just the model
Multiplexed functional screens are what let generation volume translate into real evidence.
Design for the product from round one
Stability, immunogenicity, and manufacturability need to be constraints, not afterthoughts.
What a real protein-design pipeline actually needs
- Prospective validationDesigns scored on hidden targets, not fit to structures the model already knew.
- Scalable functional assaysMultiplexed testing that keeps pace with generation volume.
- Negative-result dataExpression, toxicity, and specificity failures captured, not discarded.
- Product-aware objectivesStability, formulation, and manufacturability as design constraints, not later fixes.
- Honest hit-rate reportingTask, assay, and threshold disclosed alongside every reported success rate.
The two curves must stay separate
The CASP archive records structure-prediction rounds from 1994 through CASP16 in 2024; its organisers identify CASP14 in 2020 as the sharp AlphaFold2 accuracy step. That is a structure-prediction curve, not an experimental binder or enzyme hit-rate curve. The table above instead compares prospective assays with different targets and thresholds; a pooled field-wide functional success rate cannot be inferred from them.
The published $99 per protein starting quote includes expression and binding screening under a vendor-defined scope, while the CAR-T pooled screen and enzyme catalytic assays have different costs and evidence outputs. No comparable historical public price series for a fixed protein-function assay was found. For programme cards, clinical phase, trial identifier and measured patient endpoint remain the relevant evidence; preclinical platform claims are company-reported until those data exist.
Sources, method, and boundaries
The examples are prospective experiments in different biological tasks. Their hit rates locate bottlenecks but are not normalized model comparisons.
- Prospective test
- Evaluating a model's designs against outcomes not yet known when the designs were generated.
- Structure confidence
- A model's self-reported certainty about a predicted fold, distinct from whether the protein actually functions.
- Developability
- A design's practical suitability for expression, purification, formulation, and manufacture.



.jpg)















