Can AI Design Proteins That Actually Work?

Generating plausible sequences is becoming cheap. Expression, folding, binding, function, specificity, manufacturability, and in-vivo behavior decide whether a design is real.

Last updated September 2026
Figure 1 · The evidence baseline

Functional hits are real, but assay specific

Sequence models and structure prediction prune an immense search space, but wet-lab results reveal expression failure, aggregation, toxicity, off-target binding, weak function, and context dependence that computational confidence can miss.

19%RFdiffusion designs meeting the paper's binder threshold across five targets.
707/12,000CD20-specific proliferating CAR-T designs in the pooled Bits to Binders experiment.
0.6–38.4%Range of team hit rates in that CAR-T assay; 28 teams participated.

Hit rates are task-, assay-, threshold-, and selection-specific. They should not be compared as general model rankings or clinical success probabilities.

The answer in one paragraph

Yes, in defined tasks. RFdiffusion produced binders at a 19% experimental hit rate under its stated threshold, and a pooled CAR-T study found 707 functional designs among 12,000 submissions. These are not universal success rates: different targets, assay thresholds, and downstream requirements make direct ranking misleading. The frontier is useful validated proteins per synthesized design and dollar spent on evidence, not sequences generated per hour.

  • AlphaDesign reported in-vivo activity for 17 of 88 designed bacterial protein inhibitors, about 19% in that task.
  • The CAR-T study found team hit rates ranging from 0.6% to 38.4%.
  • Global structure-confidence statistics did not predict that study's experimental outcomes.
  • Expression, affinity, specificity, catalysis, stability, and clinical benefit require different measurements.

All hit rates here are tied to the named assay and threshold. Company pipeline claims below are identified as company-reported.

Part I: Two prospective tests

Confidence and function are separate claims

Both experiments tested designs before knowing the answer, the only way to learn whether a model's confidence predicts real biology.

Figure 2 · Two prospective experiments

Confidence is not the same claim as function

Both studies tested designed proteins prospectively rather than retrospectively, the only way to learn whether a model's confidence predicts real biology.

Small, single-lab experiment

AlphaDesign inhibitors

Task
Designed bacterial protein inhibitors, tested for in-vivo activity
Sample size
88 designed constructs
Hit rate
17/88 active, about 19%

What limited it: Expression and folding failures, not just binding geometry, separated plausible designs from active ones.

Hit rates are task-, assay-, threshold-, and selection-specific. They should not be compared as general model rankings or clinical success probabilities.

From rare hits to task-specific yields

ExperimentMeasured hit rateWhat counted
Earlier EGFR Rosetta protocolAbout 0.01%Company-reported historical comparator, with a different campaign and selection process.
Adaptyv EGFR competition, 20245/201, 2.5%Binder hit in the organiser's physical test.
Adaptyv second EGFR competition, 2025About 13%Organiser-reported binder hits among 400 selected designs.
RFdiffusion paper, 202319%Binding at least half the positive-control signal across five targets.
Bits to Binders, 2026707/12,000, about 5.9%Significant CD20-specific CAR-T proliferation in a pooled assay.

This is a series of prospective experiments, not a controlled time series. Targets, selection, expression, and hit definitions differ; no single line should be extrapolated as a field-wide learning rate.

Part II: The physical stack

Four stages between a proposal and a product

Each stage can silently discard a design that looked correct at the stage before it.

01

Objective and representation

The desired geometry, interaction, dynamics, and context become a computable target.

Measure
Constraint coverage
Failure boundary
An incomplete objective can be satisfied perfectly and still miss what matters biologically.
Where the frontier moves

Objectives that encode function and context, not just static structure.

02

Generation and filtering

Models propose backbones and sequences, then rank structure and developability.

Measure
Designs/hour · diversity
Failure boundary
Model miscalibration can make implausible designs look confidently correct.
Where the frontier moves

Generation abundant enough that filtering, not proposing, becomes the bottleneck.

03

Build and assay

DNA synthesis, expression, purification, binding, and functional screens test reality.

Measure
Validated hits/design
Failure boundary
Lab throughput caps how many of the generated designs can actually be tested.
Where the frontier moves

Multiplexed assays that test thousands of sequences in parallel with function labels.

04

Product development

Stability, immunogenicity, delivery, manufacture, and in-vivo efficacy define usefulness.

Measure
Potency · yield · safety
Failure boundary
A design that works in a test tube can still fail translation into a manufacturable product.
Where the frontier moves

Designing for stability and manufacture before optimization even converges.

Part III: The measurement budget

Sequence space is enormous; evidence is finite

Computation can search and compress priors, but each new function remains an empirical claim. Multiplexed experiments lower the cost of evidence without eliminating it. A commercial expression-and-binding service advertises a starting price of $99 per protein and a three-to-four-week turnaround; that is a vendor quote for one assay scope, not a universal physical floor. It excludes clinical safety, in-vivo efficacy, and often the full cost of synthesis and optimization.

functional validated hits÷design-build-test cycles=protein design productivity

Binding is easier to count than catalysis

Earlier computational Kemp eliminases had catalytic efficiencies of roughly 1–420 M⁻¹ s⁻¹, versus a cited median near 10⁵ M⁻¹ s⁻¹ for natural enzymes. A 2025 de novo metallohydrolase study reported a zero-shot design above 10⁴ M⁻¹ s⁻¹ among 96 directly tested candidates, within the range of native metallohydrolases on similar substrates. These are different reactions and assays, so the result shows progress in a harder task rather than a universal enzyme hit rate.

Translation has also begun: South Korea approved the SKYCovione vaccine in 2022 using a computationally designed protein nanoparticle, while Generate:Biomedicines reports its AI-engineered GB-0895 antibody is in Phase 3 asthma trials as of 2026. Approval of a nanoparticle vaccine and a candidate in trials are different evidence levels; neither makes an untested design a medicine.

Part IV: The bottleneck shift

The bottleneck moves into assays and objectives

Once generation is abundant, teams need scalable functional assays, negative-selection data, meaningful thresholds, and models that predict more than geometric plausibility.

Prospective benchmarks

Score hidden designs with standardized physical tests, not retrospective fits to known structures.

Multiplex assays

Measure thousands of sequences in parallel while preserving reliable function labels for each one.

Learn from negatives

Capture expression, toxicity, and specificity failures instead of only publishing the successes.

Design for product

Include stability, formulation, and manufacture as constraints before optimization converges, not after.

Who is building what

Structural diffusion models, multimodal biological transformers, sequence-to-function generative models, and automated high-throughput assays design proteins from scratch. Search the record, or filter by application domain.

8 programmes
Institute for Protein Design (Baker Lab)RFdiffusion & ProteinMPNNDenoising diffusion probabilistic models generating backbone structures directly from atomic coordinates, paired with ProteinMPNN for inverse sequence design
Reported evidence
Pioneered foundational open-source tools cited worldwide; designed high-affinity de novo binders targeting SARS-CoV-2, influenza, and cancer antigens.
Announced next step
Universal generative models designing functional binders, enzymes, and nanomachine assemblies in a single automated step.
Unresolved risk
Physical expressibility in E. coli or mammalian cell culture, aggregation in solution, and structural flexibility not captured by rigid models.
Google DeepMind / Isomorphic LabsAlphaFold 3Diffusion-based biomolecular architecture predicting 3D joint structures of proteins, DNA, RNA, post-translational modifications, and small molecule ligand complexes
Reported evidence
Demonstrated unprecedented accuracy on blind protein-ligand and protein-nucleic acid interaction benchmarks; partnered with Eli Lilly and Novartis.
Announced next step
De novo drug design moving from target structural prediction directly to high-affinity lead therapeutic molecules.
Unresolved risk
Predicting dynamic conformational state ensembles (induced-fit vs conformational selection) from single static structural predictions.
Generate:BiomedicinesThe Chroma ArchitectureGenerative biological diffusion models for antibodies, protein complexes, and peptide therapeutics
Reported evidence
Company-reported April 2026 pipeline: GB-0895 was in Phase 3 asthma trials; GB-4362 and GB-5267 had clinical-study plans. These are development stages, not approvals or proof that every program was designed de novo.
Announced next step
Demonstrate safety and efficacy for specific candidates in controlled clinical trials.
Unresolved risk
Clinical developability, including viscosity, formulation stability, and anti-drug-antibody immunogenicity.
EvolutionaryScaleESM3 Frontier Model98-billion parameter multimodal biological model simultaneously reasoning over protein sequence, 3D structure, and biological function tokens
Reported evidence
Demonstrated de novo generation of an entirely novel green fluorescent protein (esmGFP) with 58% sequence divergence from any known natural protein.
Announced next step
Generative biology models acting as foundational programming interfaces for synthetic enzyme and cellular circuit design.
Unresolved risk
Inference compute cost for multi-billion parameter models and bridging sequence tokens to multi-body biochemical kinetics.
Cradle BioGenerative Engineering PlatformModels trained on iterative wet-lab feedback to optimize protein stability, expression, and activity
Reported evidence
Company-reported partner work and activity improvements; no independently normalized cross-project hit-rate series is disclosed here.
Announced next step
Validate improvements against matched controls in partner processes.
Unresolved risk
Models may not generalize across enzyme families or small, sparse assay datasets.
Nabla BioHigh-Throughput Wet-Lab MLMultiplexed antibody screens coupled to predictive models
Reported evidence
Company-reported pharmaceutical partnerships and work on difficult membrane targets; public independent clinical outcomes and comparable hit rates are not established in this card.
Announced next step
Prospective antibody validation with developability and specificity profiling.
Unresolved risk
Assay artifacts and the gap between in-vitro binding and in-vivo pharmacokinetics.
Profluent BioOpenCRISPR-1 & Generative LLMsLarge language models trained on global protein diversity, generating fully synthetic CRISPR gene editors with distinct sequences from natural Cas enzymes
Reported evidence
Designed and synthesized OpenCRISPR-1, demonstrating comparable human genome editing efficiency to wild-type SpCas9 with lower off-target rates.
Announced next step
Complete generative design of bespoke gene-editing enzymes, molecular motors, and targeted delivery effectors.
Unresolved risk
Patent landscaping and intellectual property freedom-to-operate around synthetic variants of naturally occurring enzymes.
CASP / CAMEOBlind Prediction BenchmarksCritical Assessment of Structure Prediction (CASP) and Continuous Automated Model EvaluatiOn (CAMEO) conducting double-blind structural accuracy assessments
Reported evidence
The gold standard rigorous benchmark that crowned AlphaFold and continuously ranks global structural biology and protein design models.
Announced next step
Expanding rigorous blind evaluation to protein design tasks (de novo ligand binding, conformational dynamics, and enzyme catalytic efficiency).
Unresolved risk
Experimental crystallization bottlenecks verifying community-designed proteins in physical wet labs.

High in silico predicted binding scores and low RMSD do not guarantee biochemical expression. Solubility, thermal stability, aggregation propensity, and immunogenicity govern lab survival.

The optimistic view, with conditions

Protein design becomes an experimental compiler

Models will translate functional specifications into small, diverse libraries whose measured failures and successes continuously improve the system.

Now

Test prospectively, always

Retrospective structure metrics do not substitute for measuring designs the model has never seen.

Near term

Scale the assay, not just the model

Multiplexed functional screens are what let generation volume translate into real evidence.

Structural

Design for the product from round one

Stability, immunogenicity, and manufacturability need to be constraints, not afterthoughts.

What a real protein-design pipeline actually needs

  1. Prospective validationDesigns scored on hidden targets, not fit to structures the model already knew.
  2. Scalable functional assaysMultiplexed testing that keeps pace with generation volume.
  3. Negative-result dataExpression, toxicity, and specificity failures captured, not discarded.
  4. Product-aware objectivesStability, formulation, and manufacturability as design constraints, not later fixes.
  5. Honest hit-rate reportingTask, assay, and threshold disclosed alongside every reported success rate.

The two curves must stay separate

The CASP archive records structure-prediction rounds from 1994 through CASP16 in 2024; its organisers identify CASP14 in 2020 as the sharp AlphaFold2 accuracy step. That is a structure-prediction curve, not an experimental binder or enzyme hit-rate curve. The table above instead compares prospective assays with different targets and thresholds; a pooled field-wide functional success rate cannot be inferred from them.

The published $99 per protein starting quote includes expression and binding screening under a vendor-defined scope, while the CAR-T pooled screen and enzyme catalytic assays have different costs and evidence outputs. No comparable historical public price series for a fixed protein-function assay was found. For programme cards, clinical phase, trial identifier and measured patient endpoint remain the relevant evidence; preclinical platform claims are company-reported until those data exist.

Sources, method, and boundaries

The examples are prospective experiments in different biological tasks. Their hit rates locate bottlenecks but are not normalized model comparisons.

Prospective test
Evaluating a model's designs against outcomes not yet known when the designs were generated.
Structure confidence
A model's self-reported certainty about a predicted fold, distinct from whether the protein actually functions.
Developability
A design's practical suitability for expression, purification, formulation, and manufacture.