Why Do Computers Use So Much Energy?

Landauer's limit prices irreversible erasure, but modern machines spend far more moving information through memory, packages, networks, and cooling. Bandwidth is the concrete energy problem.

Last updated September 2026
Figure 1 · The thermodynamic baseline

Arithmetic became abundant faster than proximity

HBM stacks memory beside compute on dense packages to create wide, short interfaces. Caches, tiling, operation fusion, quantization, sparsity, and near-memory designs attack the same problem from software and architecture.

kT ln 2Minimum heat for conventional irreversible bit erasure.
1.18 TB/sHBM throughput cited by NIST for a recent commercial generation.
$3.87bnExpected investment in an Indiana HBM packaging and R&D facility.

Landauer's principle is a thermodynamic limit, not measured chip energy. Vendor bandwidth is a peak component figure, not application throughput or energy efficiency.

The answer in one paragraph

The useful computing floor is not one transistor operation. It is the energy and time needed to deliver data to an operation and return a verified result. As arithmetic density rises, locality, HBM, packaging, interconnect, precision, software reuse, and utilization determine how much silicon becomes useful work, because in every real system, moving a bit costs far more energy than switching one.

  • At temperature T, conventional irreversible bit erasure has a minimum dissipation of kT ln 2.
  • Present systems remain many orders above that floor because they switch physical devices reliably and move information through long hierarchies in finite time.
  • NIST describes high-bandwidth memory as a core AI supply-chain constraint and cites a generation processing up to 1.18 terabytes per second.
  • Peak interface bandwidth omits capacity, latency, access patterns, energy per bit, package yield, network traffic, cooling, and application utilization.

Measured results, derived quantities, projections, targets, and editorial inference are identified by context. Announced capacity is never treated as operating performance.

Part I: The floor-to-reality ladder

Ten orders of magnitude between the floor and off-chip DRAM

Every layer a bit crosses on its way from storage to an arithmetic unit adds energy, and the largest jumps happen well before the network even gets involved.

Figure 2 · The floor-to-reality ladder

Every layer between compute and its data costs more than the Landauer floor

Every layer between compute and its data costs more than the Landauer floorEnergy per bit rises from about 2.85 zeptojoules at the Landauer floor to roughly 30 picojoules for off-chip DRAM access, a span of about ten orders of magnitude, plotted on a logarithmic scale across on-chip arithmetic, cache, HBM, DRAM, and network interconnect.1.0 zJ1000.0 zJ1 fJ1 pJ100 pJLandauer floor (300K)On-chip arithmeticOn-chip cache accessHBM package interfaceOff-chip DRAM accessNetwork interconnect123456Energy per bit, joules (log scale)
Landauer floor (300K) · 2.8 zJ

kT ln 2 at room temperature: the minimum thermodynamic dissipation for one irreversible bit erasure. No real device operates anywhere near this.

Figures combine NIST's thermodynamic floor discussion, Horowitz's widely cited 45nm energy breakdown (ISSCC 2014), and public HBM and data-center optical-interconnect figures. They describe representative orders of magnitude for a mix of technologies and process nodes, not one measured system.
Part II: The physical stack

Four layers, each with its own distance to close

The floor is physics; everything above it is architecture, packaging, and software choosing how far a bit has to travel.

01

Physical state

Transistors, memory cells, wires, and emerging devices encode, retain, and switch information under noise.

Measure
J/switch · error rate
Failure boundary
Leakage, reliability, and finite time all cost energy beyond the thermodynamic minimum.
Where the frontier moves

Devices that hold state reliably closer to the Landauer floor without sacrificing speed.

02

Package and memory

Registers, caches, HBM stacks, interposers, substrates, and host memory trade capacity against distance.

Measure
J/bit · GB/s · latency
Failure boundary
Thermals and package yield limit how close memory can physically sit to compute.
Where the frontier moves

Advanced packaging that keeps shortening the distance data has to travel.

03

System interconnect

Links, switches, collective communication, and storage move state beyond one package.

Measure
Bytes/result · network utilization
Failure boundary
Topology and synchronization overhead can dominate once data leaves the package.
Where the frontier moves

Co-packaged optics and other links that close the gap with on-package movement.

04

Algorithms and software

Precision, sparsity, locality, fusion, caching, routing, and scheduling decide which movement occurs.

Measure
Verified outcomes/joule
Failure boundary
Generality and poor reuse mean the same result gets recomputed or re-fetched needlessly.
Where the frontier moves

Software that treats data movement, not just arithmetic, as the resource being optimized.

Part III: The floor

Erasure has a floor; communication has a distance

Reversible operations can approach lower dissipation only under restrictive assumptions and time trade-offs. Real systems must maintain energy barriers against noise and charge physical wires; placing data closer and reusing it are therefore near-term gains far larger than approaching Landauer's limit.

energy per bit moved × bytes+switching and facility energy=energy per verified result
Part IV: The bottleneck shift

Packaging becomes architecture

When memory sits beside logic, bonding, interposers, thermals, test, repair, HBM supply, and software locality shape performance as directly as processor design. The bottleneck migrates from arithmetic units into the whole package and fabric.

Reuse before moving

Tile, fuse, cache, and restructure algorithms around locality instead of re-fetching the same data.

Represent with fewer bits

Use the lowest precision and compression consistent with verified output quality.

Shorten the link

Use HBM, advanced packaging, chiplets, optics, or near-memory compute where the workload justifies it.

Measure outcomes

Report whole-system joules and latency per accepted task, not isolated peak operations.

Who is building what

Co-packaged optics (CPO), silicon photonics, high-bandwidth memory (HBM), and universal chiplet standards dismantle the compute memory wall. Search the record, or filter by interconnect layer.

8 programmes
Celestial AIPhotonic FabricSilicon photonics platform delivering optical compute and memory interconnectivity, decoupling compute silicon from memory placement via optical links
Reported evidence
Secured major customer commitments and hyperscale investments; demonstrated multi-terabit low-latency die-to-die and die-to-memory interconnects.
Announced next step
Commercial deployment of disaggregated, optically pooled high-bandwidth memory (HBM) systems for multi-accelerator AI clusters.
Unresolved risk
Optical coupling losses at fiber-to-chip interfaces and manufacturing yield of dense silicon photonic waveguide arrays.
Ayar LabsTeraPHY & SuperNova LaserMonolithic electronic-photonic optical I/O chiplets utilizing micro-ring modulators driven by external multi-wavelength continuous-wave laser sources
Reported evidence
Demonstrated operational TeraPHY links integrated with Intel processors and NVIDIA GPUs; achieved <5 pJ/bit energy efficiency across multi-terabit links.
Announced next step
Integration of optical I/O directly into standard high-volume packaging substrates replacing electrical SerDes traces.
Unresolved risk
Thermal sensitivity of micro-ring modulators requiring active thermal tuning loops, and laser diode lifetime in data center environments.
LightmatterPassage Optical InterposerWafer-scale programmable silicon photonics interposer that routes optical signals between heterogeneous compute dies placed directly on top
Reported evidence
Demonstrated wafer-scale optical routing with orders-of-magnitude higher cross-sectional bandwidth density than organic packaging substrates.
Announced next step
Enabling rack-scale virtual compute chips consisting of hundreds of interconnected GPUs behaving as a single monolithic processor.
Unresolved risk
Mechanical stress and thermal expansion differences across full 300mm silicon interposers under high thermal dissipation loads.
BroadcomBailly 51.2 Tbps Co-Packaged OpticsCo-Packaged Optics (CPO) switch integrating 51.2 Tbps silicon switching ASICs with optical transceiver engines on a common substrate
Reported evidence
Manufactured and demonstrated working commercial 51.2 Tbps Bailly CPO systems, cutting optical interconnect power by over 30% versus pluggable transceivers.
Announced next step
Scaling CPO to 102.4 Tbps switch architectures, eliminating copper backplane trace losses in hyperscale AI network fabrics.
Unresolved risk
Serviceability and field repair: replacing a failed optical channel currently requires removing the entire multi-thousand-dollar switch package.
SK HynixHBM3E & HBM4 Memory Stacks12- and 16-high stacked DRAM dies interconnected via Through-Silicon Vias (TSVs) and Advanced Mass Reflow Molded Underfill (MR-MUF)
Reported evidence
Dominant supplier of HBM3E memory for NVIDIA H100 and B200 systems, delivering over 1.18 TB/s bandwidth per stack.
Announced next step
Transitioning to HBM4 featuring 2048-bit wide interfaces and logic base dies manufactured on advanced foundry nodes (TSMC).
Unresolved risk
Thermal dissipation from bottom DRAM layers trapped under top dies, and TSV micro-bump bonding yield across 16-layer stacks.
TSMC3DFabric & CoWoS (Co-Packaged Optics)Chip-on-Wafer-on-Substrate (CoWoS-R, CoWoS-L) and Compact Universal Photonic Engine (COUPE) integrating optics and compute in 3D packaging
Reported evidence
The manufacturing foundation for modern AI accelerators; expanding CoWoS capacity to multi-reticle dimensions exceeding 3.3× standard mask size.
Announced next step
Monolithic integration of electronic and photonic ICs (EPIC) with lowest parasitic capacitance and sub-picojoule transfer energy.
Unresolved risk
Substrate warpage during thermal cycling and severe global packaging substrate supply bottlenecks constraining AI accelerator shipments.
UCIe ConsortiumUniversal Chiplet Interconnect ExpressOpen industry standard defining physical, die-to-die adapter, and protocol layers (PCIe, CXL, streaming) for heterogeneously packaged chiplets
Reported evidence
Over 100 member companies including Intel, AMD, Arm, TSMC, and Qualcomm; validated interoperability across silicon vendor test chips.
Announced next step
Creating a thriving commercial merchant chiplet market where designers mix and match dies across different foundries and process nodes.
Unresolved risk
Testing and known-good-die (KGD) liability when chiplets from multiple competing vendors are integrated into a single expensive package.
IEEE Photonics / DARPA PIPESPhotonic Edge-to-Edge StandardsPublic research and standards initiatives benchmarking photonic link energy efficiency, latency, and optical interconnect architectures
Reported evidence
Funded foundational micro-ring, comb laser, and optical interposer developments that seeded leading commercial silicon photonics startups.
Announced next step
Establishing sub-1 pJ/bit interconnect metrics for exascale computing and disaggregated data center memory pools.
Unresolved risk
Academic prototypes demonstrating record efficiency in controlled settings struggling with manufacturing tolerances on commercial CMOS lines.

Optical I/O energy ratings (e.g. sub-5 pJ/bit) describe transceiver physical layer links; system-level energy must account for external continuous-wave laser thermal coolers and electrical serializer overhead.

The optimistic view, with conditions

The efficient computer is a balanced information machine

The largest practical gains will come from co-designing algorithms, numerics, memory hierarchy, packaging, networks, power, and cooling so that useful information travels less often and less far.

Package

Keep shortening the wire

HBM already beats off-chip DRAM per bit precisely by sitting closer, the same logic extends further with advanced packaging.

Network

Close the gap with co-packaged optics

Bringing optics onto the package can cut interconnect energy well below today's pluggable-module figures.

Software

Stop paying twice for the same bit

Locality-aware algorithms avoid re-fetching data that better reuse would have kept nearby.

What actually lowers energy per verified result

  1. Shorter distancesPackaging and memory placement that keep data physically close to where it's used.
  2. Fewer bits movedPrecision and compression matched to what the task actually needs.
  3. Higher reuseAlgorithms and caching that avoid re-fetching the same data repeatedly.
  4. Whole-system measurementJoules and latency reported per accepted task, not isolated peak operations.
  5. Co-designed layersAlgorithms, packaging, and networking optimized together, not independently.

Efficiency improved while the workload grew

The historical counterpoint is Koomey and colleagues’ computing-efficiency series: computations per kWh roughly doubled every 1.5 years across their long-run computer sample. That is a historical trend for a specified benchmark and hardware population, not a physical law or a forecast for today’s AI services. Dennard-style voltage and power-density scaling slowed after the mid-2000s, so adding transistors ceased to give the same automatic energy saving per operation; memory movement, interconnect and cooling gained weight.

System electricity now has a separate growth curve. IEA estimates 415 TWh of global data-centre electricity in 2024, about 1.5% of world demand, and models roughly 945 TWh in its 2030 base case. Those totals include conventional servers, AI accelerators, storage, networking and facility overhead. They cannot be divided by an assumed number of AI prompts to produce a measured joules-per-inference figure. A valid accelerator comparison would hold model quality, batch, latency, memory traffic and facility boundary fixed while reporting accepted outputs per kWh.

One floor cannot describe the whole computer

Landauer's kT ln 2 minimum at room temperature is about 2.9 × 10⁻²¹ joule per irreversible bit erasure. It is a physical lower bound for an ideal operation, not a budget for a transistor, DRAM access, HBM link, network transfer, power converter and cooling system. Reversible logic experiments investigate the gap but have not demonstrated a general-purpose computer that removes those full-system costs. A defensible current-node ladder would quote pJ/bit by memory technology, link length, data rate and test setup; 2014 Horowitz values remain pedagogical historical points rather than present hardware specifications.

The IEA projects global data-centre electricity near 950 TWh in 2030 from 485 TWh in 2025 in its current case. System demand can therefore rise while joules per arithmetic operation fall. Useful accepted work per facility MWh, at fixed model quality and latency, is the comparison that connects the chip, memory and demand curves.

Sources, method, and boundaries

The thermodynamic floor follows NIST's discussion of conventional and reversible computation. HBM figures come from NIST's project summary. The energy-per-bit ladder combines this floor with Horowitz's widely cited 45nm energy breakdown (ISSCC 2014) and public HBM and data-center optical-interconnect figures: representative orders of magnitude spanning different technologies and process nodes, not measurements of one system. The roofline relationship is a model, not a benchmark for a specific application.

Landauer's principle
The thermodynamic minimum energy required to irreversibly erase one bit of information at a given temperature.
Roofline model
A performance model showing that achievable compute throughput is capped by memory bandwidth for a given arithmetic intensity.
Co-packaged optics
Optical transceivers integrated directly onto a chip package rather than connected via a separate pluggable module.