Ten orders of magnitude between the floor and off-chip DRAM
Every layer a bit crosses on its way from storage to an arithmetic unit adds energy, and the largest jumps happen well before the network even gets involved.
Every layer between compute and its data costs more than the Landauer floor
kT ln 2 at room temperature: the minimum thermodynamic dissipation for one irreversible bit erasure. No real device operates anywhere near this.
Four layers, each with its own distance to close
The floor is physics; everything above it is architecture, packaging, and software choosing how far a bit has to travel.
Physical state
Transistors, memory cells, wires, and emerging devices encode, retain, and switch information under noise.
- Measure
- J/switch · error rate
- Failure boundary
- Leakage, reliability, and finite time all cost energy beyond the thermodynamic minimum.
Where the frontier moves
Devices that hold state reliably closer to the Landauer floor without sacrificing speed.
Package and memory
Registers, caches, HBM stacks, interposers, substrates, and host memory trade capacity against distance.
- Measure
- J/bit · GB/s · latency
- Failure boundary
- Thermals and package yield limit how close memory can physically sit to compute.
Where the frontier moves
Advanced packaging that keeps shortening the distance data has to travel.
System interconnect
Links, switches, collective communication, and storage move state beyond one package.
- Measure
- Bytes/result · network utilization
- Failure boundary
- Topology and synchronization overhead can dominate once data leaves the package.
Where the frontier moves
Co-packaged optics and other links that close the gap with on-package movement.
Algorithms and software
Precision, sparsity, locality, fusion, caching, routing, and scheduling decide which movement occurs.
- Measure
- Verified outcomes/joule
- Failure boundary
- Generality and poor reuse mean the same result gets recomputed or re-fetched needlessly.
Where the frontier moves
Software that treats data movement, not just arithmetic, as the resource being optimized.
Erasure has a floor; communication has a distance
Reversible operations can approach lower dissipation only under restrictive assumptions and time trade-offs. Real systems must maintain energy barriers against noise and charge physical wires; placing data closer and reusing it are therefore near-term gains far larger than approaching Landauer's limit.
Packaging becomes architecture
When memory sits beside logic, bonding, interposers, thermals, test, repair, HBM supply, and software locality shape performance as directly as processor design. The bottleneck migrates from arithmetic units into the whole package and fabric.
Reuse before moving
Tile, fuse, cache, and restructure algorithms around locality instead of re-fetching the same data.
Represent with fewer bits
Use the lowest precision and compression consistent with verified output quality.
Shorten the link
Use HBM, advanced packaging, chiplets, optics, or near-memory compute where the workload justifies it.
Measure outcomes
Report whole-system joules and latency per accepted task, not isolated peak operations.
Who is building what
Co-packaged optics (CPO), silicon photonics, high-bandwidth memory (HBM), and universal chiplet standards dismantle the compute memory wall. Search the record, or filter by interconnect layer.
Celestial AIPhotonic FabricSilicon photonics platform delivering optical compute and memory interconnectivity, decoupling compute silicon from memory placement via optical links
- Reported evidence
- Secured major customer commitments and hyperscale investments; demonstrated multi-terabit low-latency die-to-die and die-to-memory interconnects.
- Announced next step
- Commercial deployment of disaggregated, optically pooled high-bandwidth memory (HBM) systems for multi-accelerator AI clusters.
- Unresolved risk
- Optical coupling losses at fiber-to-chip interfaces and manufacturing yield of dense silicon photonic waveguide arrays.
Ayar LabsTeraPHY & SuperNova LaserMonolithic electronic-photonic optical I/O chiplets utilizing micro-ring modulators driven by external multi-wavelength continuous-wave laser sources
- Reported evidence
- Demonstrated operational TeraPHY links integrated with Intel processors and NVIDIA GPUs; achieved <5 pJ/bit energy efficiency across multi-terabit links.
- Announced next step
- Integration of optical I/O directly into standard high-volume packaging substrates replacing electrical SerDes traces.
- Unresolved risk
- Thermal sensitivity of micro-ring modulators requiring active thermal tuning loops, and laser diode lifetime in data center environments.
LightmatterPassage Optical InterposerWafer-scale programmable silicon photonics interposer that routes optical signals between heterogeneous compute dies placed directly on top
- Reported evidence
- Demonstrated wafer-scale optical routing with orders-of-magnitude higher cross-sectional bandwidth density than organic packaging substrates.
- Announced next step
- Enabling rack-scale virtual compute chips consisting of hundreds of interconnected GPUs behaving as a single monolithic processor.
- Unresolved risk
- Mechanical stress and thermal expansion differences across full 300mm silicon interposers under high thermal dissipation loads.
BroadcomBailly 51.2 Tbps Co-Packaged OpticsCo-Packaged Optics (CPO) switch integrating 51.2 Tbps silicon switching ASICs with optical transceiver engines on a common substrate
- Reported evidence
- Manufactured and demonstrated working commercial 51.2 Tbps Bailly CPO systems, cutting optical interconnect power by over 30% versus pluggable transceivers.
- Announced next step
- Scaling CPO to 102.4 Tbps switch architectures, eliminating copper backplane trace losses in hyperscale AI network fabrics.
- Unresolved risk
- Serviceability and field repair: replacing a failed optical channel currently requires removing the entire multi-thousand-dollar switch package.
SK HynixHBM3E & HBM4 Memory Stacks12- and 16-high stacked DRAM dies interconnected via Through-Silicon Vias (TSVs) and Advanced Mass Reflow Molded Underfill (MR-MUF)
- Reported evidence
- Dominant supplier of HBM3E memory for NVIDIA H100 and B200 systems, delivering over 1.18 TB/s bandwidth per stack.
- Announced next step
- Transitioning to HBM4 featuring 2048-bit wide interfaces and logic base dies manufactured on advanced foundry nodes (TSMC).
- Unresolved risk
- Thermal dissipation from bottom DRAM layers trapped under top dies, and TSV micro-bump bonding yield across 16-layer stacks.
TSMC3DFabric & CoWoS (Co-Packaged Optics)Chip-on-Wafer-on-Substrate (CoWoS-R, CoWoS-L) and Compact Universal Photonic Engine (COUPE) integrating optics and compute in 3D packaging
- Reported evidence
- The manufacturing foundation for modern AI accelerators; expanding CoWoS capacity to multi-reticle dimensions exceeding 3.3× standard mask size.
- Announced next step
- Monolithic integration of electronic and photonic ICs (EPIC) with lowest parasitic capacitance and sub-picojoule transfer energy.
- Unresolved risk
- Substrate warpage during thermal cycling and severe global packaging substrate supply bottlenecks constraining AI accelerator shipments.
UCIe ConsortiumUniversal Chiplet Interconnect ExpressOpen industry standard defining physical, die-to-die adapter, and protocol layers (PCIe, CXL, streaming) for heterogeneously packaged chiplets
- Reported evidence
- Over 100 member companies including Intel, AMD, Arm, TSMC, and Qualcomm; validated interoperability across silicon vendor test chips.
- Announced next step
- Creating a thriving commercial merchant chiplet market where designers mix and match dies across different foundries and process nodes.
- Unresolved risk
- Testing and known-good-die (KGD) liability when chiplets from multiple competing vendors are integrated into a single expensive package.
IEEE Photonics / DARPA PIPESPhotonic Edge-to-Edge StandardsPublic research and standards initiatives benchmarking photonic link energy efficiency, latency, and optical interconnect architectures
- Reported evidence
- Funded foundational micro-ring, comb laser, and optical interposer developments that seeded leading commercial silicon photonics startups.
- Announced next step
- Establishing sub-1 pJ/bit interconnect metrics for exascale computing and disaggregated data center memory pools.
- Unresolved risk
- Academic prototypes demonstrating record efficiency in controlled settings struggling with manufacturing tolerances on commercial CMOS lines.
Optical I/O energy ratings (e.g. sub-5 pJ/bit) describe transceiver physical layer links; system-level energy must account for external continuous-wave laser thermal coolers and electrical serializer overhead.
The optimistic view, with conditions
The efficient computer is a balanced information machine
The largest practical gains will come from co-designing algorithms, numerics, memory hierarchy, packaging, networks, power, and cooling so that useful information travels less often and less far.
Keep shortening the wire
HBM already beats off-chip DRAM per bit precisely by sitting closer, the same logic extends further with advanced packaging.
Close the gap with co-packaged optics
Bringing optics onto the package can cut interconnect energy well below today's pluggable-module figures.
Stop paying twice for the same bit
Locality-aware algorithms avoid re-fetching data that better reuse would have kept nearby.
What actually lowers energy per verified result
- Shorter distancesPackaging and memory placement that keep data physically close to where it's used.
- Fewer bits movedPrecision and compression matched to what the task actually needs.
- Higher reuseAlgorithms and caching that avoid re-fetching the same data repeatedly.
- Whole-system measurementJoules and latency reported per accepted task, not isolated peak operations.
- Co-designed layersAlgorithms, packaging, and networking optimized together, not independently.
Efficiency improved while the workload grew
The historical counterpoint is Koomey and colleagues’ computing-efficiency series: computations per kWh roughly doubled every 1.5 years across their long-run computer sample. That is a historical trend for a specified benchmark and hardware population, not a physical law or a forecast for today’s AI services. Dennard-style voltage and power-density scaling slowed after the mid-2000s, so adding transistors ceased to give the same automatic energy saving per operation; memory movement, interconnect and cooling gained weight.
System electricity now has a separate growth curve. IEA estimates 415 TWh of global data-centre electricity in 2024, about 1.5% of world demand, and models roughly 945 TWh in its 2030 base case. Those totals include conventional servers, AI accelerators, storage, networking and facility overhead. They cannot be divided by an assumed number of AI prompts to produce a measured joules-per-inference figure. A valid accelerator comparison would hold model quality, batch, latency, memory traffic and facility boundary fixed while reporting accepted outputs per kWh.
One floor cannot describe the whole computer
Landauer's kT ln 2 minimum at room temperature is about 2.9 × 10⁻²¹ joule per irreversible bit erasure. It is a physical lower bound for an ideal operation, not a budget for a transistor, DRAM access, HBM link, network transfer, power converter and cooling system. Reversible logic experiments investigate the gap but have not demonstrated a general-purpose computer that removes those full-system costs. A defensible current-node ladder would quote pJ/bit by memory technology, link length, data rate and test setup; 2014 Horowitz values remain pedagogical historical points rather than present hardware specifications.
The IEA projects global data-centre electricity near 950 TWh in 2030 from 485 TWh in 2025 in its current case. System demand can therefore rise while joules per arithmetic operation fall. Useful accepted work per facility MWh, at fixed model quality and latency, is the comparison that connects the chip, memory and demand curves.
Sources, method, and boundaries
The thermodynamic floor follows NIST's discussion of conventional and reversible computation. HBM figures come from NIST's project summary. The energy-per-bit ladder combines this floor with Horowitz's widely cited 45nm energy breakdown (ISSCC 2014) and public HBM and data-center optical-interconnect figures: representative orders of magnitude spanning different technologies and process nodes, not measurements of one system. The roofline relationship is a model, not a benchmark for a specific application.
- Landauer's principle
- The thermodynamic minimum energy required to irreversibly erase one bit of information at a given temperature.
- Roofline model
- A performance model showing that achievable compute throughput is capped by memory bandwidth for a given arithmetic intensity.
- Co-packaged optics
- Optical transceivers integrated directly onto a chip package rather than connected via a separate pluggable module.



















