DNA

The Cost of Reading and Writing DNA

Reading a base became roughly a hundred million times cheaper. Writing one became about two hundred times cheaper. This is an account of why the two diverged, and what the gap costs.

Research through 20 September 2026
Figure 1

One curve collapsed. The other bent.

Both capabilities were laboratory crafts in 1990. Only one of them turned into a semiconductor problem.

Cost per DNA base read and written, 1990 to 2025On a logarithmic scale, the cost of reading a base falls by roughly eight orders of magnitude while the cost of writing a base falls by roughly two, opening a widening gap between the two capabilities.$10$1$10⁻⁴$10⁻⁶$10⁻⁸19901995200020052010201520202025Cost per base, logarithmicthe gap~10⁷×ReadingWriting
Reading · 2025

Sub-$200 genome claims from several platforms. Cost per raw base is now far below the cost of interpreting what the bases mean.

Figure 1: Approximate cost per base of reading and of writing DNA, on a logarithmic scale. The two series are deliberately not like-for-like: reading is priced per raw base produced by an instrument, while writing is priced per base pair of verified, delivered double-stranded DNA, which includes assembly and confirmation. That asymmetry is part of the argument rather than a defect in it, because verification is what makes a written base expensive. Points are editorial reconstructions from published list prices, platform announcements, and period accounts; they are not a market index, and premium or small-volume orders sit well above the plotted boundary.
~10⁸×Approximate fall in the cost of reading a base since 1990, faster over the same period than the fall in the cost of a transistor.
~10²×Approximate fall in the cost of writing a verified base pair over the same period, using essentially the same chemistry throughout.
Different problemsReading is a parallel measurement that tolerates error through repetition. Writing is a sequential synthesis in which error compounds and must be removed.

What the record shows

  • Sequencing stopped being the expensive part. In a targeted clinical assay, the chemistry can be a few percent of the delivered cost, while preparation, interpretation, and undiagnosed cases carry the rest.
  • The collapse came from parallelism, not from a single discovery. Reading many molecules in one imaging step put sequencing on the same kind of curve as feature density in semiconductors.
  • Writing did not follow, because stepwise chemistry compounds error with length. Long constructs must be assembled and verified from short pieces, and verification means sequencing them.
  • The binding constraint has moved from measurement to meaning. Interpretation depends on large, diverse, well-phenotyped cohorts, which accumulate on a clock no instrument controls.

This report separates the price of an instrument run, the cost per sample in a laboratory, the price a payer is charged, and the cost of producing a result that changes a decision. They are different quantities and should not be placed on one curve.

Figure 2: The reading trade space

Every platform buys length with cost

Read length, per-base accuracy, parallelism, and price trade against one another. No platform is best; each answers a different question.

The lowest cost per base ever achieved, and the basis of almost every large dataset

Signal
Reversible terminators imaged across a dense flow cell
Read length
100 to 300 bases
Per-base accuracy
High per base
Parallelism
Billions of reads per run

Binding constraint: Short fragments cannot span repeats, structural rearrangements, or phase two chromosomes apart

Cost per base is set almost entirely by parallelism, which is why reading followed a semiconductor curve. Length and accuracy are set by chemistry and physics, which is why they did not fall at the same rate.
Part I: First principles

A sequence is a measurement, and a construct is a product

Reading asks what is there and can be repeated until the answer is confident. Writing asks for something that did not exist and must be correct the first time it is used.

Clinical valueanswers found × actionability × timeliness÷uncertainty + cases without answers + delivered cost
01

Obtain the material

Every genomic claim begins with cells someone consented to give. Sample quality, quantity, preservation, and provenance bound everything downstream.

Measure
Input mass, integrity, contamination, consent
Failure boundary
Degraded, mixed, or unrepresentative material, and populations absent from the record entirely
What the record shows at this step
Reported evidence
Formalin-fixed tissue, cell-free plasma, and ancient or forensic material each impose their own input limits regardless of platform.
Where it is moving
Less input and more tolerance of damaged material, which widens what can be sequenced rather than lowering the price of what already can.
Principal risk
Degraded, mixed, or unrepresentative material, and populations absent from the record entirely

Evidence statements describe published methods or commercially available platforms. “Where it is moving” is an editorial reading of direction, not an announced capability.

02

Build a library

DNA is fragmented, repaired, and given adapters so that a machine can grip it. This step converts biology into a format an instrument accepts.

Measure
Conversion efficiency, duplication rate, bias
Failure boundary
Preparation bias that silently under-samples the regions that matter most
What the record shows at this step
Reported evidence
Library preparation is now a larger share of the per-sample cost than the sequencing chemistry itself in small targeted assays.
Where it is moving
Automation and miniaturisation, which reduce a cost that did not fall with the chemistry.
Principal risk
Preparation bias that silently under-samples the regions that matter most

Evidence statements describe published methods or commercially available platforms. “Where it is moving” is an editorial reading of direction, not an announced capability.

03

Copy or observe directly

Most platforms amplify a molecule into a colony so its signal is loud enough to detect. Single-molecule platforms skip this and accept a noisier signal instead.

Measure
Amplification bias, duplicate rate, input requirement
Failure boundary
Polymerase preference distorts the apparent abundance of what was in the sample
What the record shows at this step
Reported evidence
Single-molecule approaches remove amplification bias and read native base modifications that copying erases.
Where it is moving
Amplification-free workflows that preserve epigenetic information alongside sequence.
Principal risk
Polymerase preference distorts the apparent abundance of what was in the sample

Evidence statements describe published methods or commercially available platforms. “Where it is moving” is an editorial reading of direction, not an announced capability.

04

Detect the signal

Fluorescence, ionic current, or pH change is recorded base by base. This is the step that became a parallel imaging and semiconductor problem, and it is the step whose cost collapsed.

Measure
Reads per run, per-base error, run time
Failure boundary
Systematic error modes that repeat identically across every read of a difficult region
What the record shows at this step
Reported evidence
Parallelism rose from one fragment per capillary to billions of fragments per flow cell, which is the whole of the cost curve in Figure 1.
Where it is moving
Denser flow cells and more pores, extending the same lever that produced the last eight orders of magnitude.
Principal risk
Systematic error modes that repeat identically across every read of a difficult region

Evidence statements describe published methods or commercially available platforms. “Where it is moving” is an editorial reading of direction, not an announced capability.

05

Call bases and assemble

Raw signal becomes sequence, sequence is aligned to a reference or assembled without one, and the two copies of each chromosome are separated where possible.

Measure
Mapping rate, phasing, assembly contiguity
Failure boundary
Reference bias, and repeats that short reads cannot span at any depth
What the record shows at this step
Reported evidence
Long reads and pangenome references resolve regions that were effectively invisible to a decade of short-read studies.
Where it is moving
Complete, phased, reference-free assemblies as the routine output rather than a specialist result.
Principal risk
Reference bias, and repeats that short reads cannot span at any depth

Evidence statements describe published methods or commercially available platforms. “Where it is moving” is an editorial reading of direction, not an announced capability.

06

Interpret the variation

A list of differences from a reference is not a finding. Each difference must be weighed for pathogenicity, penetrance, and relevance to the question actually asked.

Measure
Variants of uncertain significance, curation time, reclassification rate
Failure boundary
Confident interpretation built on databases that under-represent most of humanity
What the record shows at this step
Reported evidence
Interpretation and curation now dominate the per-case cost of clinical genome analysis in the model below.
Where it is moving
Large, diverse, well-phenotyped cohorts, which reduce uncertainty in a way no instrument can.
Principal risk
Confident interpretation built on databases that under-represent most of humanity

Evidence statements describe published methods or commercially available platforms. “Where it is moving” is an editorial reading of direction, not an announced capability.

07

Write it back

Synthesis turns a design into physical DNA. Chemistry is stepwise, so error compounds with length, and long constructs must be assembled and verified from short, imperfect pieces.

Measure
Coupling efficiency, full-length yield, verified cost per base pair
Failure boundary
Assembly and verification cost that grows faster than the raw synthesis cost falls
What the record shows at this step
Reported evidence
Phosphoramidite chemistry, commercialised in the 1980s, remains the dominant method four decades later.
Where it is moving
Enzymatic synthesis, which aims to replace the chemistry rather than parallelise it further.
Principal risk
Assembly and verification cost that grows faster than the raw synthesis cost falls

Evidence statements describe published methods or commercially available platforms. “Where it is moving” is an editorial reading of direction, not an announced capability.

Cheap reading made writing look expensive. It also made writing possible to check, which is why the two capabilities are locked together.
Figure 3 · The two walls

Reading is a statistics problem. Writing is a yield problem.

Both capabilities have a hard mathematical boundary, and the two boundaries have completely different shapes. That difference, not funding or attention, is most of why one curve collapsed and the other did not.

Probability of detecting a variant against sequencing depthDetection probability rises steeply with depth and saturates, so more depth buys confidence with diminishing returns that depend on the allele fraction being sought.0%25%50%75%100%1×10×100×1000×Sequencing depth, logarithmic1× skim0.0%30× genome100%100× exome100%1000× tumour100%50% allele
Looking for:

Detection here means at least 3 independent reads carrying the variant. A common germline variant is found reliably at ordinary depth; a variant present in a hundredth of the molecules needs orders of magnitude more reading, which is why liquid biopsy is expensive even when sequencing is cheap.

Figure 3: Both panels are exact consequences of simple models (a binomial sampling process for reading, and a compounding per-step yield for writing) rather than measurements of any particular instrument. Real platforms depart from both: reads are not independent across difficult sequence, and synthesis errors include deletions as well as truncations. The shapes, not the precise values, are the point.

Why one curve and not the other

The divergence in Figure 1 is not a story about effort or investment. Four structural differences explain most of it.

Reading parallelised. Writing did not.

A flow cell reads billions of fragments in the same imaging step, so cost per base fell with feature density, exactly as it did for transistors. Synthesis adds one base at a time to each growing strand; running a million syntheses at once lowers cost per reaction but not the number of steps in any one of them.

Reading tolerates error. Writing does not.

A sequencing error is corrected by reading the same molecule again, so depth converts a noisy measurement into a confident one. A synthesis error is a defect in a physical product; it must be found and removed by sequencing the result, which makes reading a tax on writing.

Reading had one clear metric. Writing has several.

Cost per base is a fair summary of a read. A synthesised construct is judged on length, sequence fidelity, delivery format, turnaround, and whether it can be assembled into something larger, so competition never collapsed onto a single falling number.

Reading changed chemistry repeatedly. Writing changed once.

Chain termination, pyrosequencing, reversible terminators, and nanopores each reset the curve. Phosphoramidite chemistry has carried almost all of synthesis since the 1980s, and incremental improvement of one chemistry produces incremental price declines.

What the collapse actually bought

10⁹ readsA single modern flow cell produces on the order of a billion reads in one run, from one fragment per capillary forty years ago.
~200 ntThe practical ceiling for a single synthesised oligonucleotide, set by compounding per-step yield rather than by price.
HoursTime to generate a genome, against years to resolve what an uncertain variant in it means.

Sequencing became abundant faster than the interpretive and clinical infrastructure around it. That mismatch, rather than any remaining limit of the chemistry, is what now determines whether a genome changes a decision.

The discovery chain

Reading reset its chemistry four times in fifty years. Writing reset once, in 1981, and has been improved rather than replaced ever since.

  1. 1977

    Chain-termination sequencing

    Sanger's method made reading DNA a laboratory routine and defined the unit of cost that everything later would be measured against.

  2. 1981

    Phosphoramidite synthesis

    Caruthers established the chemistry for writing DNA. It is still the dominant method, which is most of the reason writing has no comparable cost curve.

  3. 1986

    Automated fluorescent sequencing

    Replacing radioactive labels with fluorescence made sequencing a machine-readable measurement rather than a manual interpretation.

  4. 1990-2003

    The Human Genome Project

    A reference sequence produced at a cost measured in billions established both the baseline and the demand for something far cheaper.

  5. 2005

    Massively parallel sequencing

    Reading many fragments simultaneously broke the one-read-per-capillary limit and started the collapse plotted in Figure 1.

  6. 2014-2019

    Long reads become practical

    Nanopore and single-molecule real-time platforms restored read length, recovering structure and phase that short reads had traded away.

  7. 2022-2025

    The sub-$200 genome

    Several platforms claimed genome costs below two hundred dollars, by which point sequencing was no longer the expensive part of a clinical answer.

Who is building what

Reading, writing, and interpretation are pursued by different organisations with different business models. Search the record, or filter by which part of the stack a programme works on.

12 programmes
IlluminaNovaSeq XHigh-density flow cells with reversible-terminator chemistry
Reported evidence
The platform behind most large-scale sequencing to date; the company announced genome costs around $200 at maximum instrument output.
Announced next step
Higher cluster density and further reduction in cost per gigabase.
Unresolved risk
Announced per-genome figures assume a fully loaded instrument and exclude preparation, analysis, and interpretation entirely.
Oxford NanoporeMinION, PromethIONIonic-current sensing through protein nanopores
Reported evidence
Reads spanning hundreds of kilobases, real-time output, native detection of base modifications, and a device small enough to use in the field.
Announced next step
Higher raw accuracy and pore density at lower cost per base.
Unresolved risk
Raw per-base error historically higher than sequencing by synthesis, requiring consensus depth on difficult sequence.
PacBioRevioSingle-molecule real-time sequencing with circular consensus
Reported evidence
Long reads at high consensus accuracy, now the basis of most complete genome assemblies and phased analyses.
Announced next step
Higher throughput to narrow the cost gap with short reads.
Unresolved risk
Higher cost per base and larger input requirements than short-read platforms.
MGI / Complete GenomicsDNBSEQDNA nanoball arrays with combinatorial probe anchor synthesis
Reported evidence
A competing high-throughput chemistry with published sub-$200 genome claims, and the main source of price pressure outside the United States.
Announced next step
Broader market access as patent and trade restrictions evolve.
Unresolved risk
Availability varies by jurisdiction for reasons unrelated to the technology.
Ultima GenomicsUG 100Open-substrate flow cell designed around cost per gigabase
Reported evidence
Announced a $100 genome by re-engineering the consumable rather than the chemistry.
Announced next step
Very high output applications such as single-cell and population sequencing.
Unresolved risk
A young platform with a smaller body of independent benchmarking than established instruments.
Twist BioscienceSilicon oligo synthesisMiniaturised phosphoramidite synthesis on a silicon chip
Reported evidence
Chip-based parallelism reduced oligo and gene prices substantially, reaching roughly a dime per base pair for synthetic genes.
Announced next step
Longer constructs and lower verified cost per base pair.
Unresolved risk
Parallelism lowers cost per reaction but does not shorten the stepwise chemistry that limits length.
Ansa BiotechnologiesEnzymatic synthesisTemplate-independent polymerase extension in aqueous conditions
Reported evidence
Reported synthesis of oligonucleotides substantially longer than the practical phosphoramidite ceiling.
Announced next step
A chemistry reset for writing comparable to what reading experienced four times.
Unresolved risk
Must match established chemistry on fidelity, cost, and scale before it displaces a forty-year-old process.
DNA ScriptSYNTAXBenchtop enzymatic DNA printing
Reported evidence
Moved synthesis from a centralised service to an instrument in the laboratory, changing turnaround rather than unit cost.
Announced next step
Same-day design-build cycles at the bench.
Unresolved risk
Benchtop economics favour speed over price; verification still requires sequencing.
Broad InstituteClinical and population genomicsLarge-scale sequencing with open variant resources
Reported evidence
Population reference datasets underpin most clinical variant frequency estimates in use today.
Announced next step
Broader ancestral representation in reference datasets.
Unresolved risk
Historical over-representation of European ancestry leaves uncertainty unevenly distributed across patients.
ClinGen and ClinVarShared variant evidenceCurated, publicly deposited variant classifications
Reported evidence
A shared archive that allows laboratories to see conflicting classifications and resolve them against common criteria.
Announced next step
Faster reclassification of variants of uncertain significance.
Unresolved risk
Curation is expert labour that does not parallelise and does not follow a hardware cost curve.
NIH All of UsDiverse cohort programmeLarge, deliberately diverse cohort with linked health records
Reported evidence
Built specifically to address the ancestry imbalance that limits interpretation for much of the population.
Announced next step
Enough diverse, phenotyped data to resolve uncertainty in bulk.
Unresolved risk
Cohort assembly, consent, and linkage run on a decade-scale clock no instrument can shorten.
Biosecurity screening frameworksSynthesis screeningCustomer and sequence screening obligations for synthesis providers
Reported evidence
Published guidance asks providers of synthetic DNA to screen orders and customers against sequences of concern.
Announced next step
Consistent international screening as synthesis capability spreads.
Unresolved risk
Screening cost and coverage scale with the number of providers, including benchtop instruments outside centralised services.

Platform specifications and pricing claims are reproduced from the record below. An announced cost per genome usually assumes a fully loaded instrument at maximum output and excludes preparation, analysis, and interpretation.

Part II: The economics of an answer

The unit that fell is not the unit that matters

Sequencing is sold per gigabase. Care is delivered per case. When the chemistry approaches zero, the price of an answer is set entirely by everything else.

Figure 4 · Interactive model

What does one actionable answer cost?

Sequencing is priced per gigabase. Medicine is priced per answer. The gap between those two units is where the remaining cost lives.

Rare-disease trio WGS$6,114/answer

Cost per case that changes a decision

Sequencing chemistry
$1080 · 18%
Preparation and compute
$160
Interpretation
$900
Cases without an answer
$3,974

3 samples per case; 3 gigabase target at 30× produces 90 Gb of data per sample. Excludes instrument capital, accreditation, sample collection, counselling, downstream confirmation, and the cost of the treatment the answer leads to.

Calculation & assumptions

Cost per answer = [ (target size × depth × price per gigabase + preparation) × samples per case + compute + interpretation ] ÷ diagnostic yield.

The chemistry term is the only one that has followed the cost curve in Figure 1. Move it to zero and the cancer panel scenario barely changes, because preparation, interpretation, and the cases that produce no answer carry the cost. That is what it means for a technology to stop being the bottleneck.

Figure 4: Illustrative structure with plausible magnitudes, not a price list or a reimbursement schedule. Diagnostic yields vary widely by indication, prior testing, ancestry, and the depth of phenotyping available to the analyst. A low-yield screening programme is not necessarily bad value: the cost per finding is high, but a single finding may avert a lifetime of disease.
InterpretationCuration by a qualified analyst does not parallelise, does not follow a semiconductor curve, and rises with the number of uncertain variants found.
YieldEvery case that produces no answer is charged to the cases that do. At 30% yield, two thirds of all work is carried by the third that succeeds.
PreparationLibrary preparation stayed roughly flat in price while the chemistry fell, so it now dominates small targeted assays.

What pushes cost down next

01

Cohorts resolve uncertainty

Large, diverse, deeply phenotyped datasets reclassify uncertain variants in bulk, which is the only mechanism that lowers interpretation cost at scale.

Condition: recruitment beyond historically over-sampled populations
02

Preparation automates

Miniaturised and automated library preparation attacks the line that did not fall with the chemistry.

Condition: automation that does not require a high-volume laboratory to pay off
03

Reading absorbs more questions

One sufficiently deep genome can replace a sequence of separate targeted tests, moving cost from repeated assays to a single reusable measurement.

Condition: reanalysis pathways and storage that make the measurement durable
04

Enzymatic synthesis changes the chemistry

Writing needs a reset of the kind reading had four times, not another increment of the same phosphoramidite process.

Condition: length and fidelity at least comparable to established chemistry
Part III: Where progress is stuck

The remaining constraints are evidentiary, not physical

Nothing in physics prevents a cheaper base. What prevents a cheaper answer is uncertainty, and uncertainty is reduced by data collection, clinical evidence, and time.

Interpretation, not measurement

A genome is read in hours and argued over for years. Variants of uncertain significance accumulate faster than they are resolved, and resolution requires cohorts and phenotypes, not instruments.

Reference and cohort bias

Databases over-represent European ancestry, so the same variant carries different uncertainty depending on the patient. This is a data-collection problem that cheaper sequencing does not fix by itself.

The stepwise synthesis wall

Per-step yield compounds, so a chemistry that is 99.5% efficient still truncates half of all molecules by a few hundred bases. Longer constructs need assembly and verification, and both scale badly.

Verification is reading

Every synthesised construct must be sequenced to confirm it is what was ordered. Writing therefore inherits a cost from reading, which caps how cheap a verified base pair can become.

Clinical evidence and reimbursement

A test must show it changes management, not merely that it detects something. Evidence generation and payer negotiation run on a clock unrelated to the cost of chemistry.

Consent, privacy, and biosecurity

Sequence is identifying and permanent, and synthesis capability is dual-use. Screening obligations, data governance, and export controls are now part of the delivered cost of both capabilities.

An optimistic view, with conditions

Sequence becomes background infrastructure

The strongest plausible future is not a cheaper genome. It is a genome read once, stored well, reinterpreted repeatedly as knowledge improves, and consulted like any other part of a medical record, with writing following slowly behind, on a different curve, for a different set of reasons.

Now to 2030

Interpretation industrialises

Curation pipelines, shared evidence bases, and structured reanalysis convert one-off expert judgement into a repeatable process with measurable reclassification rates.

2030s

The measurement becomes durable

A complete, phased genome taken once is reanalysed against better knowledge rather than re-sequenced, which shifts spending from chemistry to evidence.

Longer horizon

Writing gets its own reset

If enzymatic or templated synthesis reaches useful length and fidelity, writing may finally follow a curve of its own, and the constraint moves again, to design and to biosafety.

View the analyst probability ranges
MilestoneDateAnalyst probability
Routine clinical genome sequencing priced below $500 excluding interpretation203055-75%
Interpretation, not sequencing, is the majority of cost in most clinical genomic tests202870-85%
Diagnostic yield for rare disease exceeds 50% in a large unselected cohort203325-40%
Reference databases reach broad ancestral representation for common clinical use203520-35%
Enzymatic synthesis reaches routine commercial parity with phosphoramidite chemistry203230-50%
Verified synthetic DNA priced below one cent per base pair at gene length203520-35%

Editorial judgements conditional on the record above, not published forecasts or company guidance.

There is more than one finish line

  1. Technically readBases are called at adequate depth and quality.
  2. Completely readRepeats, structure, and phase are resolved, not merely aligned.
  3. Confidently interpretedVariants are classified against evidence relevant to this patient.
  4. Clinically actionableThe finding changes a decision that someone can actually take.
  5. Economically routineDelivered cost and reimbursement support use outside specialist centres.
  6. WritableThe same sequence can be synthesised, verified, and assembled at comparable cost.

Sources, method, and boundaries

Figure 1 is an editorial reconstruction from published list prices, platform announcements, and period accounts, not a continuous price series; the reading and writing series measure different products and are plotted together deliberately. Figure 3 is a mathematical model rather than a measurement. The cost model in Figure 4 is illustrative structure with plausible magnitudes and is not a price list, a reimbursement schedule, or clinical advice. Diagnostic yields describe populations in published studies, not any individual patient.

Raw base
A base produced by an instrument, before alignment, consensus, or quality filtering.
Verified base pair
A base pair of synthesised double-stranded DNA confirmed by sequencing and delivered to a customer.
Depth
The average number of independent reads covering a position, which converts a noisy measurement into a confident one.
Diagnostic yield
The share of tested cases in which the test produces an answer that explains the presentation or changes management.