A predicted fold is the start of the assay queue
Sequence models and structure prediction prune an immense search space, but wet-lab results reveal expression failure, aggregation, toxicity, off-target binding, weak function, and context dependence that computational confidence can miss.
Three numbers that locate the frontier
Hit rates are task-, assay-, threshold-, and selection-specific. They should not be compared as general model rankings or clinical success probabilities.
Validated function per design round is the curve
Design improves when prospective experiments yield more diverse, specific, stable functions and their failures update the next round. Compute throughput is valuable only when synthesis and assays keep pace.
Diversity matters because many near-identical hits can share the same hidden failure.
Assay realism matters because binding in vitro may not predict cellular or therapeutic function.
The headline metric sits on a system
Each layer can become the bottleneck even when the layer before it improves.
Objective and representation
The desired geometry, interaction, dynamics, and context become a computable target.
- Measure
- Constraint coverage
- Failure mode
- Incomplete objective
Generation and filtering
Models propose backbones and sequences, then rank structure and developability.
- Measure
- Designs/hour · diversity
- Failure mode
- Model miscalibration
Build and assay
DNA synthesis, expression, purification, binding, and functional screens test reality.
- Measure
- Validated hits/design
- Failure mode
- Lab throughput
Product development
Stability, immunogenicity, delivery, manufacture, and in-vivo efficacy define usefulness.
- Measure
- Potency · yield · safety
- Failure mode
- Translation
Sequence space is enormous; evidence is finite
Computation can search and compress priors, but each new function remains an empirical claim. Multiplexed experiments lower the cost of evidence without eliminating it.
The bottleneck moves into assays and objectives
Once generation is abundant, teams need scalable functional assays, negative-selection data, meaningful thresholds, and models that predict more than geometric plausibility.
Prospective benchmarks
Score hidden designs with standardized physical tests.
Multiplex assays
Measure thousands of sequences in parallel while preserving function labels.
Learn from negatives
Capture expression, toxicity, and specificity failures.
Design for product
Include stability, formulation, and manufacture before optimization converges.
An optimistic view, with conditions
Protein design becomes an experimental compiler
Models will translate functional specifications into small, diverse libraries whose measured failures and successes continuously improve the system.
Sources, method, and boundaries
The examples are prospective experiments in different biological tasks. Their hit rates locate bottlenecks but are not normalized model comparisons.



.jpg)















