Optical memory pooling still has a physical realisation problem
As AI systems evolve, balancing HBM cache, DDR5 capacity and optical I/O will make module granularity critical to building memory-pooling architectures that can be efficiently cooled, validated and scaled.
By Dr. Moh Kolbehdari, Sr. Director of IC/Packaging, Socionext US
Accelerators need high-bandwidth memory close to compute, but long-context and multi-user workloads can generate key-value caches that exceed practical local HBM capacity. Host DRAM offers greater capacity at lower cost, while storage provides even more capacity but introduces too much latency for frequently accessed inference data [1-3].
Optical memory pooling offers an attractive response: separate memory capacity from the accelerator, connect it through a low-latency optical fabric, and allow multiple servers or accelerators to share a larger memory resource.
Recent public architectures combine three complementary technologies [1-2]:
- DDR5 for economical, terabyte-scale capacity
- HBM3E for a high-bandwidth cache or fast memory tier
- Optical I/O for rack-scale reach and memory sharing
This is a compelling system concept. However, solving the rack-level connectivity problem does not eliminate the physical-realization challenges inside the memory module.
The central engineering question is: At what point does a disaggregated-memory architecture become too highly integrated within each package-plus-board module?
This question extends beyond module partitioning. The physical realization must also close DDR5 channel topology and achievable data rate, HBM cache behavior, package and PCB routing, power delivery, thermal distribution, optical stability, manufacturing variation, test, reliability, and serviceability. A more distributed implementation may improve some of these boundaries, for example by shortening DDR channels, reducing the number of PHYs and routes concentrated around one ASIC, distributing heat, and creating smaller units for validation and manufacturing learning. However, these benefits must be weighed against duplicated optical interfaces, control logic, packaging, power management, and system-level coordination. The objective is therefore not maximum distribution, but identifying the physical granularity that provides the best combined bandwidth, thermal, electrical, manufacturing, reliability, and cost result.
HBM for speed, DDR5 for capacity
The purpose of combining HBM and DDR5 is not to create two
equivalent memory pools.
DDR5 supplies most of the capacity. HBM acts as a high-bandwidth cache for active regions of the DDR address space, although the published architecture also allows HBM to be configured as additional pooled memory [1].
A recent publicly available Marvell technical paper describes a Photonic Fabric Memory Module containing two 36 GB HBM3E stacks, eight DDR5 interfaces supporting as much as 2 TB of DDR5 capacity, 7.2 Tbps of optical connectivity, a photonic integrated-circuit interposer, a fiber-array unit, and a memory-fabric ASIC. Sixteen such modules form a 32 TB shared-memory appliance [1].
The hierarchy is straightforward:
- HBM provides bandwidth
- DDR5 provides capacity
- Optics provides distance and sharing
The HBM cache is especially important because DDR5 alone may not sustain the peak optical-link bandwidth under every access pattern. The published evaluation provisions the HBM tier with bandwidth above the optical-link requirement so that the optical fabric can remain utilized when the DDR subsystem would otherwise become the limiting factor [1-2].
Caching, however, does not remove the DDR5 implementation burden. System performance still depends on cache locality, address interleaving, replacement policy, write behavior, and the proportion of transactions that ultimately reach the DIMMs.
The DDR5 topology creates a difficult choice
A public block diagram may show four DIMMs on either side of
the central package, but physical placement does not reveal electrical channel
topology [1].
There are two broad possibilities.
One DIMM per channel
If eight DIMMs use eight independent DDR5 interfaces, the
architecture avoids the severe loading associated with placing several DIMMs on
one channel.
However, the ASIC must support eight high-speed DDR interfaces. Even when one common controller complex manages them, the implementation still requires multiple channel data paths, PHYs, training engines, clocks, command/address resources, and package connections.
That creates substantial costs in:
- ASIC die-edge area
- Controller and PHY power
- Package pin count
- Package escape routing
- Board layer count
- Timing calibration
- Test coverage
- Power-delivery design
DDR5 modules also contain local power-management circuitry and additional active components. That reduces some motherboard rail-management burden, but it adds module-level power, monitoring, and thermal considerations [6].
The public technical description identifies eight DDR5 interfaces, which strongly suggests extensive parallel memory connectivity. It does not, however, disclose the supported DDR transfer rate, exact channel architecture, DIMM rank configuration, controller power, or full-population electrical margins [1].
Multiple DIMMs per channel
Using fewer controllers and PHYs would reduce ASIC area and
I/O concentration, but connecting two or especially four DIMMs to one channel
introduces a different problem.
Additional DIMMs increase capacitive loading, command/address loading, discontinuities, stubs, reflections, training complexity, and trace-length variation. These effects generally reduce the maximum practical operating rate.
The trade-off is visible in commercial server platforms. Intel documentation, for example, identifies lower maximum DDR5 rates for two-DIMM-per-channel configurations than for equivalent one-DIMM-per-channel configurations on several Xeon generations [4-5].
Four DIMMs per channel would be considerably more challenging at high DDR5 rates. It should not be assumed merely because four physical DIMMs appear on one side of a drawing.
The implementation therefore faces a fundamental trade-off:
- Fewer channels reduce ASIC complexity but increase channel loading.
- More channels preserve bandwidth but increase ASIC, package, power, and routing complexity.
The PCB remains part of the memory path
Optics may carry the transaction across racks, but the final
movement from the memory ASIC to DDR5 still occurs electrically across the
module PCB.
That path includes:
- ASIC package escape
- Vias and layer transitions
- Long board traces
- DIMM connectors
- Command/address and clock distribution
- Data and strobe matching
- Reference-plane transitions
- Return-path control
The propagation delay of the PCB trace is only one portion of total access latency. A cache-miss transaction may pass through:
Host CXL interface → optical NIC → optical fabric → photonic receiver → memory/fabric ASIC → HBM cache lookup → DDR controller and PHY → DIMM → return path.
The published work reports low-latency behavior based on hardware emulation and system simulation. The authors also state that end-to-end validation using physical appliance hardware and inference workloads remains pending [1].
That distinction matters because simulation can validate architecture and protocol behavior, while physical hardware must also close:
- DDR signal integrity
- Simultaneous-switching noise
- Package and PCB power integrity
- DIMM temperature
- Airflow
- Optical alignment
- Photonic temperature stability
- Manufacturing variation
- Long-duration traffic behavior
Thermal management is distributed, not eliminated
A module containing a central ASIC, two HBM stacks,
integrated photonics, eight DDR5 DIMMs, and multiple local power-management
circuits creates several thermal regions [1-6].
The central 2.5D package requires concentrated heat removal. The surrounding DIMMs need airflow or conduction paths that are not obstructed by the central heat sink. DIMM power-management devices and DRAM components can generate localized heating, while elevated
DRAM temperature may increase refresh activity or force performance throttling.
The optical region adds another constraint. Modulators, detectors, fiber coupling, and laser illumination must remain stable while adjacent digital and memory devices dissipate heat.
HBM reduces repeated DDR traffic when cache locality is favorable, but it does not remove DIMM thermal requirements. Under cache misses, streaming workloads, cache replacement, or write-through behavior, the backing DDR remains active.
The complete package-plus-board thermal field therefore matters more than any individual component temperature.
Has the module become a board-scale monolith?
The industry moved toward chiplets because putting every
function into one die eventually created unacceptable pressure on area, yield,
power, development cost, and design complexity.
The same principle can apply one level higher.
A system may be disaggregated across racks yet remain highly concentrated inside each memory module. Integrating eight DDR interfaces, two HBM stacks, high-bandwidth optical I/O, switching logic, cache management, and a large board-level memory subsystem into one realization object can reproduce many of the same challenges that chiplets were intended to relieve.
The issue is not whether the architecture is innovative. It clearly is.
The issue is whether the chosen module granularity is the most practical one.
Concentrated versus distributed implementation
Consider a conceptual alternative in which the capacity and
bandwidth of one concentrated module are partitioned across multiple smaller
modules connected in parallel through the optical fabric. This is not a
prescription for a fixed module count. The number of modules, DDR5 interfaces,
DIMMs, HBM or cache capacity, and optical bandwidth per module are design
variables.
Each smaller module could contain fewer DDR5 interfaces and DIMMs, a smaller memory/fabric ASIC, proportionally allocated high-bandwidth cache capacity, fewer package and board routes, and a lower optical-bandwidth allocation.
The distributed modules could collectively provide the same target capacity and aggregate bandwidth.
This finer-grained architecture could offer several advantages:
- Shorter DDR routes
- Fewer DDR PHYs per ASIC
- Smaller package and die size
- Distributed power and heat
- Simpler board stack-up
- Easier airflow
- Better fault isolation
- Incremental capacity expansion
- More manageable design verification
Shorter and less heavily loaded DDR channels may also improve electrical margin and create an opportunity to sustain higher practical DDR5 data rates. This is not automatic; the achievable rate must be demonstrated through channel analysis, training behavior, full-population testing, and hardware measurement.
A failure or thermal limit in one unit would affect only part of the memory pool. Manufacturing learning could also proceed on a smaller realization object rather than on the largest possible module.
But modularity has costs.
- A distributed implementation may duplicate
- ASIC control logic
- Optical interfaces
- Fiber-array attachments
- Package substrates
- Power-management circuitry
- Clocking
- Test infrastructure
- Cache-management resources
The software and fabric must also stripe memory efficiently across modules, preserve ordering and consistency, balance traffic, and prevent one slower module from creating tail latency.
The right answer may therefore be neither maximum integration nor maximum partitioning.
It may be the smallest module that provides acceptable economics while keeping electrical, thermal, optical, manufacturing, and validation complexity within a manageable boundary.
What physical validation must sh
Before any optical pooled-memory architecture can be judged at production scale, several
implementation factors must be addressed, including the sustained DDR5 rate
with every DIMM populated and whether each DIMM has a dedicated interface.
The associated controller and PHY area and power costs must also be considered, alongside how much traffic is served from HBM, how workload-dependent the cache hit rate is, and what bandwidth can be sustained during HBM misses.
Thermal management is another key challenge, particularly how the ASIC, HBM, photonics, DIMMs and DIMM PMICs can be cooled simultaneously, and how temperature variations affect DDR margins, refresh behaviour and optical stability. Finally, the architecture must account for module or optical-path failures and assess whether a distributed implementation could deliver better total cost, manufacturability, reliability and serviceability than a single concentrated module.
Optical memory pooling may become an important new AI-infrastructure tier. But optical reach alone does not prove physical scalability.
The deeper lesson is familiar:
- Disaggregation should not stop at the rack boundary.
- The package and module must also be partitioned at a scale that can be routed, powered, cooled, manufactured, tested, and repaired.
A system architecture becomes real only when its chosen physical granularity is practical.
Figure 1: Concentrated module versus distributed implementation The same aggregate memory capacity can be realized with different physical granularities. The optimum depends on ASIC cost, DDR signaling, cache behavior, thermal density, optical duplication, yield, and serviceability.













