A sequence is a measurement, and a construct is a product
Reading asks what is there and can be repeated until the answer is confident. Writing asks for something that did not exist and must be correct the first time it is used.
Obtain the material
Every genomic claim begins with cells someone consented to give. Sample quality, quantity, preservation, and provenance bound everything downstream.
- Measure
- Input mass, integrity, contamination, consent
- Failure boundary
- Degraded, mixed, or unrepresentative material, and populations absent from the record entirely
What the record shows at this step
- Reported evidence
- Formalin-fixed tissue, cell-free plasma, and ancient or forensic material each impose their own input limits regardless of platform.
- Where it is moving
- Less input and more tolerance of damaged material, which widens what can be sequenced rather than lowering the price of what already can.
- Principal risk
- Degraded, mixed, or unrepresentative material, and populations absent from the record entirely
Evidence statements describe published methods or commercially available platforms. “Where it is moving” is an editorial reading of direction, not an announced capability.
Build a library
DNA is fragmented, repaired, and given adapters so that a machine can grip it. This step converts biology into a format an instrument accepts.
- Measure
- Conversion efficiency, duplication rate, bias
- Failure boundary
- Preparation bias that silently under-samples the regions that matter most
What the record shows at this step
- Reported evidence
- Library preparation is now a larger share of the per-sample cost than the sequencing chemistry itself in small targeted assays.
- Where it is moving
- Automation and miniaturisation, which reduce a cost that did not fall with the chemistry.
- Principal risk
- Preparation bias that silently under-samples the regions that matter most
Evidence statements describe published methods or commercially available platforms. “Where it is moving” is an editorial reading of direction, not an announced capability.
Copy or observe directly
Most platforms amplify a molecule into a colony so its signal is loud enough to detect. Single-molecule platforms skip this and accept a noisier signal instead.
- Measure
- Amplification bias, duplicate rate, input requirement
- Failure boundary
- Polymerase preference distorts the apparent abundance of what was in the sample
What the record shows at this step
- Reported evidence
- Single-molecule approaches remove amplification bias and read native base modifications that copying erases.
- Where it is moving
- Amplification-free workflows that preserve epigenetic information alongside sequence.
- Principal risk
- Polymerase preference distorts the apparent abundance of what was in the sample
Evidence statements describe published methods or commercially available platforms. “Where it is moving” is an editorial reading of direction, not an announced capability.
Detect the signal
Fluorescence, ionic current, or pH change is recorded base by base. This is the step that became a parallel imaging and semiconductor problem, and it is the step whose cost collapsed.
- Measure
- Reads per run, per-base error, run time
- Failure boundary
- Systematic error modes that repeat identically across every read of a difficult region
What the record shows at this step
- Reported evidence
- Parallelism rose from one fragment per capillary to billions of fragments per flow cell, which is the whole of the cost curve in Figure 1.
- Where it is moving
- Denser flow cells and more pores, extending the same lever that produced the last eight orders of magnitude.
- Principal risk
- Systematic error modes that repeat identically across every read of a difficult region
Evidence statements describe published methods or commercially available platforms. “Where it is moving” is an editorial reading of direction, not an announced capability.
Call bases and assemble
Raw signal becomes sequence, sequence is aligned to a reference or assembled without one, and the two copies of each chromosome are separated where possible.
- Measure
- Mapping rate, phasing, assembly contiguity
- Failure boundary
- Reference bias, and repeats that short reads cannot span at any depth
What the record shows at this step
- Reported evidence
- Long reads and pangenome references resolve regions that were effectively invisible to a decade of short-read studies.
- Where it is moving
- Complete, phased, reference-free assemblies as the routine output rather than a specialist result.
- Principal risk
- Reference bias, and repeats that short reads cannot span at any depth
Evidence statements describe published methods or commercially available platforms. “Where it is moving” is an editorial reading of direction, not an announced capability.
Interpret the variation
A list of differences from a reference is not a finding. Each difference must be weighed for pathogenicity, penetrance, and relevance to the question actually asked.
- Measure
- Variants of uncertain significance, curation time, reclassification rate
- Failure boundary
- Confident interpretation built on databases that under-represent most of humanity
What the record shows at this step
- Reported evidence
- Interpretation and curation now dominate the per-case cost of clinical genome analysis in the model below.
- Where it is moving
- Large, diverse, well-phenotyped cohorts, which reduce uncertainty in a way no instrument can.
- Principal risk
- Confident interpretation built on databases that under-represent most of humanity
Evidence statements describe published methods or commercially available platforms. “Where it is moving” is an editorial reading of direction, not an announced capability.
Write it back
Synthesis turns a design into physical DNA. Chemistry is stepwise, so error compounds with length, and long constructs must be assembled and verified from short, imperfect pieces.
- Measure
- Coupling efficiency, full-length yield, verified cost per base pair
- Failure boundary
- Assembly and verification cost that grows faster than the raw synthesis cost falls
What the record shows at this step
- Reported evidence
- Phosphoramidite chemistry, commercialised in the 1980s, remains the dominant method four decades later.
- Where it is moving
- Enzymatic synthesis, which aims to replace the chemistry rather than parallelise it further.
- Principal risk
- Assembly and verification cost that grows faster than the raw synthesis cost falls
Evidence statements describe published methods or commercially available platforms. “Where it is moving” is an editorial reading of direction, not an announced capability.
Cheap reading made writing look expensive. It also made writing possible to check, which is why the two capabilities are locked together.
Reading is a statistics problem. Writing is a yield problem.
Both capabilities have a hard mathematical boundary, and the two boundaries have completely different shapes. That difference, not funding or attention, is most of why one curve collapsed and the other did not.
Detection here means at least 3 independent reads carrying the variant. A common germline variant is found reliably at ordinary depth; a variant present in a hundredth of the molecules needs orders of magnitude more reading, which is why liquid biopsy is expensive even when sequencing is cheap.
Why one curve and not the other
The divergence in Figure 1 is not a story about effort or investment. Four structural differences explain most of it.
Reading parallelised. Writing did not.
A flow cell reads billions of fragments in the same imaging step, so cost per base fell with feature density, exactly as it did for transistors. Synthesis adds one base at a time to each growing strand; running a million syntheses at once lowers cost per reaction but not the number of steps in any one of them.
Reading tolerates error. Writing does not.
A sequencing error is corrected by reading the same molecule again, so depth converts a noisy measurement into a confident one. A synthesis error is a defect in a physical product; it must be found and removed by sequencing the result, which makes reading a tax on writing.
Reading had one clear metric. Writing has several.
Cost per base is a fair summary of a read. A synthesised construct is judged on length, sequence fidelity, delivery format, turnaround, and whether it can be assembled into something larger, so competition never collapsed onto a single falling number.
Reading changed chemistry repeatedly. Writing changed once.
Chain termination, pyrosequencing, reversible terminators, and nanopores each reset the curve. Phosphoramidite chemistry has carried almost all of synthesis since the 1980s, and incremental improvement of one chemistry produces incremental price declines.
What the collapse actually bought
Sequencing became abundant faster than the interpretive and clinical infrastructure around it. That mismatch, rather than any remaining limit of the chemistry, is what now determines whether a genome changes a decision.
The discovery chain
Reading reset its chemistry four times in fifty years. Writing reset once, in 1981, and has been improved rather than replaced ever since.
- 1977
Chain-termination sequencing
Sanger's method made reading DNA a laboratory routine and defined the unit of cost that everything later would be measured against.
- 1981
Phosphoramidite synthesis
Caruthers established the chemistry for writing DNA. It is still the dominant method, which is most of the reason writing has no comparable cost curve.
- 1986
Automated fluorescent sequencing
Replacing radioactive labels with fluorescence made sequencing a machine-readable measurement rather than a manual interpretation.
- 1990-2003
The Human Genome Project
A reference sequence produced at a cost measured in billions established both the baseline and the demand for something far cheaper.
- 2005
Massively parallel sequencing
Reading many fragments simultaneously broke the one-read-per-capillary limit and started the collapse plotted in Figure 1.
- 2014-2019
Long reads become practical
Nanopore and single-molecule real-time platforms restored read length, recovering structure and phase that short reads had traded away.
- 2022-2025
The sub-$200 genome
Several platforms claimed genome costs below two hundred dollars, by which point sequencing was no longer the expensive part of a clinical answer.
Who is building what
Reading, writing, and interpretation are pursued by different organisations with different business models. Search the record, or filter by which part of the stack a programme works on.
IlluminaNovaSeq XHigh-density flow cells with reversible-terminator chemistry
- Reported evidence
- The platform behind most large-scale sequencing to date; the company announced genome costs around $200 at maximum instrument output.
- Announced next step
- Higher cluster density and further reduction in cost per gigabase.
- Unresolved risk
- Announced per-genome figures assume a fully loaded instrument and exclude preparation, analysis, and interpretation entirely.
Oxford NanoporeMinION, PromethIONIonic-current sensing through protein nanopores
- Reported evidence
- Reads spanning hundreds of kilobases, real-time output, native detection of base modifications, and a device small enough to use in the field.
- Announced next step
- Higher raw accuracy and pore density at lower cost per base.
- Unresolved risk
- Raw per-base error historically higher than sequencing by synthesis, requiring consensus depth on difficult sequence.
PacBioRevioSingle-molecule real-time sequencing with circular consensus
- Reported evidence
- Long reads at high consensus accuracy, now the basis of most complete genome assemblies and phased analyses.
- Announced next step
- Higher throughput to narrow the cost gap with short reads.
- Unresolved risk
- Higher cost per base and larger input requirements than short-read platforms.
MGI / Complete GenomicsDNBSEQDNA nanoball arrays with combinatorial probe anchor synthesis
- Reported evidence
- A competing high-throughput chemistry with published sub-$200 genome claims, and the main source of price pressure outside the United States.
- Announced next step
- Broader market access as patent and trade restrictions evolve.
- Unresolved risk
- Availability varies by jurisdiction for reasons unrelated to the technology.
Ultima GenomicsUG 100Open-substrate flow cell designed around cost per gigabase
- Reported evidence
- Announced a $100 genome by re-engineering the consumable rather than the chemistry.
- Announced next step
- Very high output applications such as single-cell and population sequencing.
- Unresolved risk
- A young platform with a smaller body of independent benchmarking than established instruments.
Twist BioscienceSilicon oligo synthesisMiniaturised phosphoramidite synthesis on a silicon chip
- Reported evidence
- Chip-based parallelism reduced oligo and gene prices substantially, reaching roughly a dime per base pair for synthetic genes.
- Announced next step
- Longer constructs and lower verified cost per base pair.
- Unresolved risk
- Parallelism lowers cost per reaction but does not shorten the stepwise chemistry that limits length.
Ansa BiotechnologiesEnzymatic synthesisTemplate-independent polymerase extension in aqueous conditions
- Reported evidence
- Reported synthesis of oligonucleotides substantially longer than the practical phosphoramidite ceiling.
- Announced next step
- A chemistry reset for writing comparable to what reading experienced four times.
- Unresolved risk
- Must match established chemistry on fidelity, cost, and scale before it displaces a forty-year-old process.
DNA ScriptSYNTAXBenchtop enzymatic DNA printing
- Reported evidence
- Moved synthesis from a centralised service to an instrument in the laboratory, changing turnaround rather than unit cost.
- Announced next step
- Same-day design-build cycles at the bench.
- Unresolved risk
- Benchtop economics favour speed over price; verification still requires sequencing.
Broad InstituteClinical and population genomicsLarge-scale sequencing with open variant resources
- Reported evidence
- Population reference datasets underpin most clinical variant frequency estimates in use today.
- Announced next step
- Broader ancestral representation in reference datasets.
- Unresolved risk
- Historical over-representation of European ancestry leaves uncertainty unevenly distributed across patients.
ClinGen and ClinVarShared variant evidenceCurated, publicly deposited variant classifications
- Reported evidence
- A shared archive that allows laboratories to see conflicting classifications and resolve them against common criteria.
- Announced next step
- Faster reclassification of variants of uncertain significance.
- Unresolved risk
- Curation is expert labour that does not parallelise and does not follow a hardware cost curve.
NIH All of UsDiverse cohort programmeLarge, deliberately diverse cohort with linked health records
- Reported evidence
- Built specifically to address the ancestry imbalance that limits interpretation for much of the population.
- Announced next step
- Enough diverse, phenotyped data to resolve uncertainty in bulk.
- Unresolved risk
- Cohort assembly, consent, and linkage run on a decade-scale clock no instrument can shorten.
Biosecurity screening frameworksSynthesis screeningCustomer and sequence screening obligations for synthesis providers
- Reported evidence
- Published guidance asks providers of synthetic DNA to screen orders and customers against sequences of concern.
- Announced next step
- Consistent international screening as synthesis capability spreads.
- Unresolved risk
- Screening cost and coverage scale with the number of providers, including benchtop instruments outside centralised services.
Platform specifications and pricing claims are reproduced from the record below. An announced cost per genome usually assumes a fully loaded instrument at maximum output and excludes preparation, analysis, and interpretation.
The unit that fell is not the unit that matters
Sequencing is sold per gigabase. Care is delivered per case. When the chemistry approaches zero, the price of an answer is set entirely by everything else.
What does one actionable answer cost?
Sequencing is priced per gigabase. Medicine is priced per answer. The gap between those two units is where the remaining cost lives.
Cost per case that changes a decision
- Sequencing chemistry
- $1080 · 18%
- Preparation and compute
- $160
- Interpretation
- $900
- Cases without an answer
- $3,974
3 samples per case; 3 gigabase target at 30× produces 90 Gb of data per sample. Excludes instrument capital, accreditation, sample collection, counselling, downstream confirmation, and the cost of the treatment the answer leads to.
Calculation & assumptions
Cost per answer = [ (target size × depth × price per gigabase + preparation) × samples per case + compute + interpretation ] ÷ diagnostic yield.
The chemistry term is the only one that has followed the cost curve in Figure 1. Move it to zero and the cancer panel scenario barely changes, because preparation, interpretation, and the cases that produce no answer carry the cost. That is what it means for a technology to stop being the bottleneck.
What pushes cost down next
Cohorts resolve uncertainty
Large, diverse, deeply phenotyped datasets reclassify uncertain variants in bulk, which is the only mechanism that lowers interpretation cost at scale.
Condition: recruitment beyond historically over-sampled populationsPreparation automates
Miniaturised and automated library preparation attacks the line that did not fall with the chemistry.
Condition: automation that does not require a high-volume laboratory to pay offReading absorbs more questions
One sufficiently deep genome can replace a sequence of separate targeted tests, moving cost from repeated assays to a single reusable measurement.
Condition: reanalysis pathways and storage that make the measurement durableEnzymatic synthesis changes the chemistry
Writing needs a reset of the kind reading had four times, not another increment of the same phosphoramidite process.
Condition: length and fidelity at least comparable to established chemistryThe remaining constraints are evidentiary, not physical
Nothing in physics prevents a cheaper base. What prevents a cheaper answer is uncertainty, and uncertainty is reduced by data collection, clinical evidence, and time.
Interpretation, not measurement
A genome is read in hours and argued over for years. Variants of uncertain significance accumulate faster than they are resolved, and resolution requires cohorts and phenotypes, not instruments.
Reference and cohort bias
Databases over-represent European ancestry, so the same variant carries different uncertainty depending on the patient. This is a data-collection problem that cheaper sequencing does not fix by itself.
The stepwise synthesis wall
Per-step yield compounds, so a chemistry that is 99.5% efficient still truncates half of all molecules by a few hundred bases. Longer constructs need assembly and verification, and both scale badly.
Verification is reading
Every synthesised construct must be sequenced to confirm it is what was ordered. Writing therefore inherits a cost from reading, which caps how cheap a verified base pair can become.
Clinical evidence and reimbursement
A test must show it changes management, not merely that it detects something. Evidence generation and payer negotiation run on a clock unrelated to the cost of chemistry.
Consent, privacy, and biosecurity
Sequence is identifying and permanent, and synthesis capability is dual-use. Screening obligations, data governance, and export controls are now part of the delivered cost of both capabilities.
An optimistic view, with conditions
Sequence becomes background infrastructure
The strongest plausible future is not a cheaper genome. It is a genome read once, stored well, reinterpreted repeatedly as knowledge improves, and consulted like any other part of a medical record, with writing following slowly behind, on a different curve, for a different set of reasons.
Interpretation industrialises
Curation pipelines, shared evidence bases, and structured reanalysis convert one-off expert judgement into a repeatable process with measurable reclassification rates.
The measurement becomes durable
A complete, phased genome taken once is reanalysed against better knowledge rather than re-sequenced, which shifts spending from chemistry to evidence.
Writing gets its own reset
If enzymatic or templated synthesis reaches useful length and fidelity, writing may finally follow a curve of its own, and the constraint moves again, to design and to biosafety.
View the analyst probability ranges
| Milestone | Date | Analyst probability |
|---|---|---|
| Routine clinical genome sequencing priced below $500 excluding interpretation | 2030 | 55-75% |
| Interpretation, not sequencing, is the majority of cost in most clinical genomic tests | 2028 | 70-85% |
| Diagnostic yield for rare disease exceeds 50% in a large unselected cohort | 2033 | 25-40% |
| Reference databases reach broad ancestral representation for common clinical use | 2035 | 20-35% |
| Enzymatic synthesis reaches routine commercial parity with phosphoramidite chemistry | 2032 | 30-50% |
| Verified synthetic DNA priced below one cent per base pair at gene length | 2035 | 20-35% |
Editorial judgements conditional on the record above, not published forecasts or company guidance.
There is more than one finish line
- Technically readBases are called at adequate depth and quality.
- Completely readRepeats, structure, and phase are resolved, not merely aligned.
- Confidently interpretedVariants are classified against evidence relevant to this patient.
- Clinically actionableThe finding changes a decision that someone can actually take.
- Economically routineDelivered cost and reimbursement support use outside specialist centres.
- WritableThe same sequence can be synthesised, verified, and assembled at comparable cost.
Sources, method, and boundaries
Figure 1 is an editorial reconstruction from published list prices, platform announcements, and period accounts, not a continuous price series; the reading and writing series measure different products and are plotted together deliberately. Figure 3 is a mathematical model rather than a measurement. The cost model in Figure 4 is illustrative structure with plausible magnitudes and is not a price list, a reimbursement schedule, or clinical advice. Diagnostic yields describe populations in published studies, not any individual patient.
- Raw base
- A base produced by an instrument, before alignment, consensus, or quality filtering.
- Verified base pair
- A base pair of synthesised double-stranded DNA confirmed by sequencing and delivered to a customer.
- Depth
- The average number of independent reads covering a position, which converts a noisy measurement into a confident one.
- Diagnostic yield
- The share of tested cases in which the test produces an answer that explains the presentation or changes management.