1. Memory that belongs to no one machine
What this post covers, who it is for, and the one question it answers.
Most servers keep their memory to themselves; sharing data means messages over a network and a copy on each side. CXL breaks that habit. Several machines plug into the same memory and read and write it with the instructions they use for their own DRAM. No network stack on the access path, no serialization, no copies.
This post grew out of a talk we gave at ByteDance and two papers from our lab at UIUC. Both start from the question the hardware leaves open: the memory is shared, but can you actually share data on it? The answer is harder, and more fun, than it looks.
We assume no prior exposure to CXL or caches. If you know what RAM is, you have everything you need. In a hurry? Jump to Megalon; a short recap waits there.
- Papers
- 2
- OSDI '26 · SOSP '26
- Best-paper nominee
- 1
- Megalon at OSDI '26
- Peak speedup
- 15.1×
- Prism over a Tigon-style index
- Shared capacity
- TBs
- per device, in the model both papers target
2. The memory wall
Start with the problem that makes CXL worth inventing.
Cores multiplied. Memory per core shrank.
Talk, slide 5- CPU core count
- Memory capacity per core
- projected by the source
Show the numbers
| Year | Core count | Memory / core | Projected |
|---|---|---|---|
| 2012 | 100 | 100 | |
| 2014 | 130 | 96 | |
| 2016 | 160 | 88 | |
| 2018 | 250 | 76 | |
| 2020 | 400 | 62 | |
| 2022 | 600 | 50 | |
| 2024 | 1100 | 38 | yes |
| 2025 | 1600 | 32 | yes |
A line chart from 2012 to 2025. CPU core count rises steeply to roughly sixteen times its 2012 level, while memory capacity per core falls to roughly a third of its 2012 level. The last two points are the source's projections.
Server CPUs went from a dozen cores in 2012 to well over a hundred today. DRAM capacity per core has been falling for a decade. Two limits bind at once. First, DRAM density grows more slowly than core counts. Second, a CPU cannot simply add memory channels: each DDR channel takes roughly 200 signal pins, and pins are among the scarcest resources on a package.
Practitioners feel this as a wall. Databases, caches, and feature stores want more memory per machine every year, and the machine cannot offer it. Meanwhile a lot of memory sits stranded: a server has rented out all its cores while gigabytes of its DRAM stay idle. Microsoft measured up to 25% of Azure’s DRAM stranded this way, with DRAM up to 50% of server cost.
Unfortunately, you cannot fix a pin-count problem with software. You need a new wire.
3. Enter CXL
Memory over PCIe, spoken to with ordinary loads and stores.
Compute Express Link (CXL) is that wire. It is an open standard built on PCIe, the lanes that already carry your GPU and NVMe drives. A CXL memory device is a box of DRAM on PCIe, and the trick is how the CPU talks to it: not with I/O commands like a disk, but with ordinary loads and stores. Once mapped, it is just more memory to a program.
The economics come from the pins. A DDR5-6400 channel moves ~51 GB/s over ~200 signal pins. An x16 PCIe 5.0 link moves ~63 GB/s per direction over 64 signal pins, and PCIe 6.0 doubles that. Per pin, the serial link carries roughly 4× the bandwidth, so a CPU can afford far more of them. That is how CXL adds terabytes where the DDR bus cannot.
The price is distance. Local DRAM answers in roughly 110 ns; CXL memory takes two to three times that, about the cost of reaching the other socket of a two-socket server (a NUMA hop), which software tolerates every day. It is still far closer than a remote read over RDMA (microseconds) or an SSD (tens of microseconds).
Expansion products ship today from Samsung and Micron, and recent AMD and Intel CPUs speak CXL natively. Multi-host sharing, the subject of the rest of this post, is newer; we will be clear about what ships and what is only expected.
Two ways out of the socket
Sharma et al., CXL survey · PCI-SIGA diagram comparing a DDR channel, about 200 signal pins for about 51 gigabytes per second, against a CXL link over PCIe, 64 signal pins for about 63 gigabytes per second in each direction.
Where CXL lands in the hierarchy
Melody (ASPLOS '25); Next PlatformShow the numbers
| Tier | Latency | Note | Basis |
|---|---|---|---|
| L1 cache | ~1 ns | on the core itself | textbook order of magnitude |
| L2 cache | ~4 ns | private per core | textbook order of magnitude |
| L3 cache | ~15 ns | shared on the die | textbook order of magnitude |
| Local DRAM | ~111–117 ns | over the DDR bus | measured on three Melody platforms |
| CXL memory | ~170–400 ns | about one NUMA hop away | Melody measured 214–394 ns; industry quotes 170–250 ns |
| RDMA read | ~2–10 µs | network round trip | commonly reported range |
| NVMe SSD | ~10–100 µs | flash storage | commonly reported range |
A log-scale ladder of approximate access latencies from L1 cache at about one nanosecond, through local DRAM near 110 nanoseconds, CXL at roughly 170 to 400 nanoseconds, RDMA at microseconds, and NVMe SSDs at tens of microseconds.
5. Cache coherence, from zero
Why CPUs cache, why caching breaks sharing, and the protocol that fixes it inside one machine.
Most loads never reach DRAM. A trip to memory costs about 100 ns, so every CPU keeps recently used data in small, fast, private caches and serves most loads from there.
Caching creates an obvious hazard: copies. If CPU 1 has cached value A and CPU 2 changes A in memory, CPU 1’s copy is now wrong. Within one machine, hardware solves this with cache coherence: the CPUs watch each other’s memory traffic and follow a protocol that keeps every cached copy of a location consistent.
The classic protocol is MESI, after the four states a cached line can be in: Modified, Exclusive, Shared, Invalid. (It is also the Illinois protocol; it was invented where our lab sits.) You do not need the details, but the feel of it matters for everything that follows. Step through the widget once.
The takeaway is that coherence is not free. A write to shared data invalidates every other cache’s copy, and to do that the hardware tracks who might hold one. The bookkeeping grows with the number of caches and the amount of memory tracked. Hold that thought.
MESI, one write, no lies
Talk, slides 12–29Step 1 / 10 · Three CPUs, one memory
Memory holds A = 1. No CPU has it cached yet.
An interactive step-through of the MESI protocol: three CPUs read value A and hold Shared copies; CPU 2 invalidates the others and writes A equals 2, holding a Modified line; CPU 3's next read forces a writeback, leaving both with Shared copies of the new value.
6. The catch: partly coherent memory
The compromise the hardware is expected to make, and what silently breaks inside it.
Now scale that bookkeeping up to what CXL 3.x allows: several machines, dozens of caches each, sharing terabytes. To behave like memory inside one machine, the hardware would have to run coherence across hosts, over PCIe, for every cache line. The vendors are blunt about the arithmetic: AMD, Micron, and Samsung all say the machinery involved stops scaling somewhere between dozens and a few hundred megabytes.
So the hardware is expected to compromise, in what we call the partly coherent model, which both of our papers target. The memory splits in two. A small coherent region, the SCR, is a few hundred MB that hardware keeps coherent across hosts, like the widget above. A large non-coherent region, the LNR, is the remaining several TB, where the hardware does nothing about coherence across hosts. Each host’s own caches still work; they just never hear from the others.
What goes wrong in the LNR? Exactly what MESI exists to prevent, except now nobody prevents it. Host 1 reads object A and caches it. Host 2 overwrites it with A′. On one machine that write would invalidate host 1’s copy; across machines, in the LNR, no invalidation is ever sent. Host 1 reads again, its cache says “I have that,” and it returns the stale value. Silently. Your database just served data that is no longer fresh.
Why this post exists
Sharing memory is a hardware problem with a hardware solution. Sharing data correctly, on memory that is only partly coherent, is a software problem, and it is wide open. Both of our papers live in this gap.
The two regions, side by side
Megalon §2.1 · Tigon, citing AMDA few hundred MB on a device of several TB: roughly one part in ten thousand. The marker is widened so that it is visible at all. Everything the hardware promises about coherence across hosts happens inside it.
A long bar representing several terabytes of CXL memory. The hardware-coherent region, a few hundred megabytes, would be far thinner than a pixel; it is drawn as a widened marker at the left edge and shown magnified.
The stale read, step by step
Talk, slide 311 / 4 · Host 1 reads A from the LNR
The value drops into host 1’s CPU cache, as every read does.
An animation of two hosts sharing a value in the non-coherent region. Host 1 caches A; host 2 writes A-prime; host 1 reads again and its cache serves the stale A with a warning flag.
7. Sharing anyway, and how it collapses
If the hardware will not keep caches honest in the LNR, software has to. It works, until it churns.
The first idea everyone has is to keep shared data inside the SCR. But a few hundred megabytes shared among many hosts is nothing; the whole point was the terabytes. Rejected.
The serious idea keeps the data in the vast LNR and small metadata about it in the SCR, where hardware coherence makes the metadata trustworthy. Tigon (OSDI ’25), an early database for a multi-host CXL pod, works this way. Each shared object gets a coherence record, think of a version counter plus a lock, and a host checks the record before touching the object; if someone wrote it since the host last looked, the host flushes its stale cache lines and re-reads from CXL. Objects are coarse, rows or key-value pairs of a few KB rather than cache lines, so the metadata stays small, and a shared index in the SCR lets hosts find each object and its record. We call this scheme hardware-coherent metadata-based sharing, HCMeta for short. The first figure on the right steps through it.
HCMeta is correct, and for Tigon’s workload of occasional cross-partition transactions it works well. But push more data into sharing and a problem surfaces: the metadata grows with the number of shared objects, and the SCR does not. In the Megalon paper’s example with 40-byte keys, HCMeta spends 52 bytes of SCR per object, so a 100 MB SCR caps sharing at about 1.9M objects. A terabyte of small objects holds hundreds of millions.
At the cap, HCMeta unshares an old object to make room for a new one. Objects start rotating through the tiny window of shareability, and every rotation is expensive: the host that wants an unshared object must ask its owner and wait, about 55 µs per round trip in Tigon’s artifact. We call the rotation churn. In that artifact, growing the dataset from 2.4M to 24M objects at 20% cross-host transactions cuts throughput by 10×. A cliff, not a slope.
Summary. The partly coherent model makes shared metadata the key to correct sharing, then makes the only place it can live absurdly small. So our first paper asks: how can hosts share a huge number of objects when even the metadata is too big for the coherent region?
HCMeta, step by step
Megalon slide 7 · §2.21 / 4 · Host 1 reads A
Host 1 reads A from the LNR and notes its version from A’s coherence record in the SCR: 0.
An animation of HCMeta. Host 1 caches A at version 0; host 2 writes A-prime to the LNR and bumps A's version in the SCR to 1; host 1 checks the SCR, sees 0 does not match 1, invalidates its cached A, and re-reads A-prime from the LNR.
The collapse, measured
Megalon, Figure 1(c)- HCMeta, 100 MB SCR
- HCMeta, unlimited SCR
Show the numbers
| Dataset | HCMeta, 100 MB SCR | Unlimited SCR |
|---|---|---|
| 1.2M objects | 15.5 Mops/s | 15.5 Mops/s |
| 2.4M objects | 15.2 Mops/s | 15.3 Mops/s |
| 4.8M objects | 1.0 Mops/s | 15.1 Mops/s |
| 7.2M objects | 0.9 Mops/s | 15.0 Mops/s |
| 12M objects | 0.7 Mops/s | 14.6 Mops/s |
| 18M objects | 0.6 Mops/s | 13.8 Mops/s |
A line chart of throughput versus dataset size. With unlimited SCR, throughput holds near 15 million operations per second. With a realistic SCR, throughput collapses from about 15 to about 1 million operations per second once the dataset passes a few million objects.
Paper one · OSDI 2026 · Best-paper nominee
8. Megalon
Share the data. Split the metadata.
Skipped ahead? The setup in brief
CXL lets hosts share terabytes, but the model we target keeps only a small coherent region (the SCR, a few hundred MB) coherent across hosts; the large non-coherent region (the LNR) can silently serve stale cached data. The standard fix, HCMeta, keeps metadata for every shared object in the SCR, about 52 bytes each with 40-byte keys, so a 100 MB SCR caps sharing near 1.9M objects and throughput collapses past it. ↑ The partly coherent model · ↑ How HCMeta collapses
Megalon starts where HCMeta breaks, with one observation: the metadata is two very different things glued together. The index, which maps each object’s ID to its location, is big (every object’s key, tens of bytes each) but cold: it changes on inserts, deletes, and sharing-state changes, far less often than data is written. The coherence records are tiny (a lock bit and a counter, 4 bytes) but hot: touched on every write.
HCMeta stuffs both into the SCR. Megalon’s key idea is to split them and share each the way its nature demands. The hot, tiny records stay physically shared in the SCR, where hardware coherence is exactly the right tool; at 4 bytes, about 13× more of them fit than HCMeta’s 52-byte bundles. The big, cold index is logically shared: every host keeps a full replica in its local DRAM, of which it has hundreds of gigabytes, roughly 1000× the SCR. A 100 MB SCR that capped HCMeta near 2 million objects lets Megalon share about 25 million.
At first sight, replicating the index recreates the problem: N replicas must now be kept consistent, the very thing coherence was for. However, the index is the cold half. It changes rarely, so keeping replicas in sync is cheap, given a mechanism to do it.
That mechanism is a shared log in the CXL memory itself: an ordered sequence of entries in the LNR, with only its head and tail pointers in the SCR. To change the index, a host claims the next entry with a compare-and-swap (an atomic instruction) on the coherent tail, writes the entry, and flushes it. Every host applies entries in order to its replica, reading them past its cache, and checks the tail before any index read. One totally ordered log at memory speed: replicas cannot drift, and no message-based agreement is needed. The idea comes from Node Replication, which synchronized data structures across NUMA sockets the same way.
The log supports two further coherence techniques.
Dynamic coherence records. Objects that are only read need no record, since nothing changes under the readers. So Megalon allocates records only for objects that are read and written, and when the SCR fills it demotes cold ones and reassigns their records, announcing each change through the log. Read-shared objects are no longer capped by the SCR, and churn becomes an appended entry rather than a round trip: about 8× cheaper.
Dual-path coherence. With records coming and going, an object can gain or lose its record while you are reading it, so the record alone cannot prove your cache is fresh. Hosts therefore also check, at the end of a read, whether the log recorded an allocation event for the object; if so, they flush and retry.
Other uses of the log. Since every index change is ordered by the log, Megalon can also keep a host’s private objects in its local DRAM, and cache read copies of hot shared objects there, where access is faster than CXL (paper §3.5).
Against HCMeta, Megalon delivers 15× on read-only workloads with large datasets and 10× at 5% writes once metadata outgrows the SCR. The cost is host DRAM for the index replicas, 7.6% more memory in the 24M-object read-only run.
Split sharing: each half where it belongs
Megalon §3.2, Figure 2The index moved to where memory is plentiful; only what must be coherent stays where coherence lives.
A diagram of Megalon's split: index replicas in each host's local DRAM point to data objects in the LNR and to small coherence records in the SCR. Occupancy bars compare 52 bytes per object for HCMeta against 4 for Megalon.
A shared log, in the memory itself
Megalon §3.3, §4Only the head and tail pointers need hardware coherence. The entries sit in the LNR, written with a flush and read past the cache, and order does the rest.
A shared log laid out in the non-coherent region with its head and tail pointers in the coherent region. Hosts append entries and replay them into local index replicas.
Flat where HCMeta collapses
Megalon Figure 6(a) · Talk, slide 51- HCMeta
- Megalon
Show the numbers
| Shared objects | HCMeta | Megalon |
|---|---|---|
| 4.8M | 16.6 Mops/s | 17.7 Mops/s |
| 7.2M | 5.7 Mops/s | 17.2 Mops/s |
| 12M | 2.5 Mops/s | 17.4 Mops/s |
| 24M | 1.7 Mops/s | 17.4 Mops/s |
Grouped bars of throughput for 4.8, 7.2, 12 and 24 million shared objects under 5 percent writes. Megalon stays near 17 million operations per second at every size; HCMeta falls from 16.6 to 1.7, a ten times gap at 24 million objects.
9. The index problem
Sharing objects is not enough. A store also has to find them.
“Give me every order between Monday and Wednesday.” Queries like that need a range index, the sorted tree at the heart of every database, living in CXL where all hosts can search it. An index turns out to be the worst-case tenant for partly coherent memory, for a reason that takes one paragraph to see.
The obvious move is to port a state-of-the-art in-memory index, the adaptive radix tree (ART) or a B+-tree, with the software-coherence recipe: nodes in the LNR, one version number per node in the SCR, bumped on every change and checked by readers. These indexes already keep per-node versions for concurrency control, so the port is natural. We built both: SC-ART and SC-BTree, SC for software-coherent.
Unfortunately, both fail, in opposite directions. Small nodes, too many versions. ART’s nodes are small and numerous: 100M keys make ~40M inner nodes, and at 8 bytes per version only about 40% of them fit in a 128 MB SCR. A version that spills into the LNR needs its own freshness check, a cache flush on every access, whether or not anything changed; on read-heavy workloads SC-ART runs about 10× slower than with unlimited SCR. Big nodes, false invalidations. Fatten the nodes to 128 keys and the versions fit, but one version now covers 128 keys: update any one and every host that cached any of the other 127 must flush and re-read the whole node. Under 50% writes, SC-BTree collapses to 1.8 Mops/s.
The trade-off is fundamental, and the paper names its cause: the updatable surface area, the part of a structure that can be modified in place. In ART and B-trees any node from root to leaf can be rewritten, so the surface is the entire index and every node needs a version. Track it finely and versions overflow; track it coarsely and false invalidations eat you.
So the real question, the one our second paper poses, is not how to port an index. It is: what is the right way to build an index for partly coherent memory?
Larger nodes trade version space for false invalidations
Prism §3, Figures 1–2modeled: versions fit, bystanders pay
every version fits, but one write already invalidates 22 unmodified keys; shrink the nodes to spare them and the versions overflow again
An interactive slider over index node size. At small node sizes, tens of millions of version numbers overflow the coherent region. At large node sizes, a single update falsely invalidates every key in a 128-key node. The two ends are measured; the middle is a model.
Paper two · SOSP 2026
10. Prism
The right index was on disk all along.
Prism finds the answer in the least likely place: a data structure family built to reduce disk writes. Log-structured merge trees (LSMs), the engine inside RocksDB, LevelDB, Bigtable, and Cassandra, were designed around a disk’s hatred of random writes. All writes land in one small in-memory buffer, the memtable. When it fills, it is frozen and written out as an immutable sorted file, an SSTable; background compaction merges SSTables down through levels by replacing old files with new ones. Below the memtable, nothing is modified in place.
Read that again with the last section’s lens. The updatable surface area of an LSM is the memtable, a few dozen megabytes. Perhaps surprisingly, a structure built for a different medium has exactly the shape partly coherent memory demands. The memtable is small, so it lives entirely inside the SCR, coherent across hosts with no software checks. SSTables are immutable, so they need no per-node tracking and sit in the terabyte LNR. The one care is recycling: each SSTable carries a strictly increasing ID in a small manifest in the SCR, and a host that meets a new ID flushes that region once. Coarse, safe, and rare.
A straight port, SC-LSM, already beats SC-BTree by 2.2× at 50% writes, because false invalidations simply stop existing. But it inherits the LSM’s classic ailment: at high write rates the memtable flushes faster than compaction can drain, level-0 SSTables pile up, and a read must probe every one of them, 49 per read at 50% writes.
Here Prism notices that the old cure was disk-specific. On disk you cannot update upper LSM levels in place, because random I/O is what LSMs exist to avoid. In memory, in-place updates are cheap. So Prism inserts a middle tier, the Bounded Updatable Layer (BUL), an in-place-updatable tree that absorbs repeated writes, so hot keys overwrite themselves instead of spawning SSTable after SSTable. Doesn’t that reopen the updatable-surface problem? It would, so Prism bounds it: the BUL is sized so that all of its per-node versions fit in the SCR. In the paper’s ablation at 5% writes, the BUL alone cuts SSTables probed per read from 15 to 2.
One more inversion completes the design. The memtable stops being a write buffer and becomes a cache for proven-hot keys. New keys go straight to the BUL; only a key written again is promoted. One-hit wonders never waste the scarcest bytes in the system.
| Tier | Lives in | Coherence | Cost |
|---|---|---|---|
| Memtable | SCR (64 MB) | Hardware coherence | No software checks |
| BUL | Data in LNR; version numbers in SCR (62 MB) | Per-node version check | Fine-grained, bounded |
| SSTables | LNR (the terabytes) | Immutable + rising IDs; one flush on first sight | Coarse but safe |
A 2 MB manifest in the SCR lists the live SSTables: 64 + 62 + 2 = 128 MB.
Prism's results
On emulated CXL with 128 MB of SCR: up to 6.3× over SC-ART on read-only workloads, the LSM’s traditional worst case; 7.4× and 9.4× over SC-BTree on the write-heavy YCSB-A and YCSB-F; 3.6–5.7× over Chime-CXL, an RDMA index ported to CXL; and 8.6–15.1× over Tigon-SWcc, a Tigon-style index. With 128 MB of SCR, Prism matches the throughput SC-ART needs unlimited SCR to reach on YCSB-A. Against Megalon’s own index it wins with writes and loses on read-only, where Megalon’s replica sits in local DRAM.
Three tiers, one coherence mechanism each
Prism §5, Figure 5SSTables are immutable for their whole lifetime, so they need no coherence tracking at all; strictly increasing IDs handle the one case that remains, recycling.
Prism's architecture: a memtable inside the coherent region, a bounded updatable layer with data in the non-coherent region and version numbers in the coherent region, and immutable SSTable levels below.
A key-value store under YCSB
Prism Figure 9 · Talk, slide 90- SC-ART
- SC-BTree
- Prism
Show the numbers
| Workload | SC-ART | SC-BTree | Prism |
|---|---|---|---|
| YCSB-A · 50% writes | 2.2 Mops/s | 1.6 Mops/s | 11.8 Mops/s |
| YCSB-B · 5% writes | 3.1 Mops/s | 13.0 Mops/s | 18.5 Mops/s |
| YCSB-C · read-only | 3.3 Mops/s | 19.5 Mops/s | 21.0 Mops/s |
| YCSB-F · 50% RMW | 1.9 Mops/s | 1.5 Mops/s | 14.1 Mops/s |
Grouped bars across YCSB workloads A, B, C and F. Prism reaches 11.8, 18.5, 21.0 and 14.1 million operations per second; SC-BTree collapses to 1.6 and 1.5 on the two write-heavy workloads; SC-ART stays below 3.5 throughout.
The honest chart: range scans
Prism Figure 8 · Talk, slide 110Show the numbers
| Write ratio | SC-BTree | Prism | Prism ÷ SC-BTree | Winner |
|---|---|---|---|---|
| 5% | 13.8 Mops/s | 9.8 Mops/s | 0.7× | SC-BTree |
| 10% | 9.0 Mops/s | 10.1 Mops/s | 1.1× | Prism |
| 20% | 4.0 Mops/s | 10.9 Mops/s | 2.7× | Prism |
| 30% | 2.7 Mops/s | 10.5 Mops/s | 3.9× | Prism |
| 40% | 2.0 Mops/s | 10.2 Mops/s | 5.1× | Prism |
| 50% | 1.6 Mops/s | 9.5 Mops/s | 6.0× | Prism |
A line of Prism's range-scan throughput relative to SC-BTree as write ratio grows: about 0.7 times at five percent writes, crossing 1.0 at ten percent, reaching six times at fifty percent.
How we got here
Selected milestones: CXL data systems are the third act of a decade-long story.
Act one · Tiering and pooling (2017–2023)
Memory could be slower but bigger; the game was deciding which pages live where, one host at a time. Pond added pooling: segments of one device handed to different hosts.
- Thermostat, ASPLOS ’17
First fully application-transparent page management for two-tier main memory.
- HeMem, SOSP ’21
Tiered memory management built to scale for big-data applications.
- TMO, ASPLOS ’22
Transparent memory offloading, in production across Meta's fleet.
- Pond, ASPLOS ’23
CXL memory pooling for Azure, within tight latency targets.
- TPP, ASPLOS ’23
CXL page placement for the Linux kernel.
Act two · First sharing (2025)
Tigon faces the partly coherent model first. Its HCMeta-style design is the baseline our work stands on.
- Tigon, OSDI ’25
An early database for a multi-host CXL pod, and the first system built around the partly coherent model.
Act three · General sharing (2026 →)
Megalon makes many objects shareable despite the tiny SCR; Prism gives partly coherent memory a range index. The rest of the stack is next.
- Megalon, OSDI ’26
General data sharing beyond the coherent region's capacity.
our lab
- Prism, SOSP ’26
A range index built for partly coherent memory.
our lab
Takeaways
Five things worth keeping if you keep nothing else.
CXL exists because of pins. Core counts outran the 200-pin DDR channel. CXL adds terabytes over a 64-pin PCIe link at roughly the latency of a NUMA hop, a new rung in the memory hierarchy.
Shared addresses are not shared coherence. In the model the hardware is heading for, a few hundred MB stay coherent across hosts; the terabytes beyond can silently serve stale caches. Sharing memory is solved. Sharing data is not.
The coherent region is a scarce metadata budget. Every correct sharing scheme keeps per-object metadata there, and naive schemes churn and collapse when it fills. Megalon splits the metadata: replicate the big cold index in host DRAM, keep only 4-byte hot records coherent, and order everything through a shared log in CXL, for up to 10× under writes and 15× on reads.
Updatable surface area decides which structures survive. If most of a structure can be modified in place, partly coherent memory punishes it: version overflow on one side, false invalidations on the other.
Old designs, new medium. Disk-born LSMs, with a tiny mutable buffer and an immutable everything else, fit partly coherent CXL better than the in-memory indexes we ported: up to 9.4× over them and 15.1× over a Tigon-style index.
Sources & further reading
Everything this post leans on: our papers first, then every external reference.
Sources of truth
Megalon: Efficient Data Sharing for Partly Coherent CXL Memory. OSDI 2026 (best-paper nominee).
Jiyu Hu, Seokjoo Cho, Landon Johnson, Kiran Hombal, Shreesha G. Bhat, Marcos K. Aguilera, Ramnatthan Alagappan, Aishwarya Ganesan
paper (PDF) · slides (PDF) · USENIX page · artifact on GitHub
Disk-Based LSMs: An Unexpectedly Good Index for Partly Coherent CXL Memory. SOSP 2026.
Kiran Hombal, Jiyu Hu, Marcos K. Aguilera, Ramnatthan Alagappan, Aishwarya Ganesan
Intro to CXL: Data Sharing on Partly Coherent CXL Memory. Talk at ByteDance, 2026.
Kiran Hombal and Jiyu Hu. The deck is not public; slide numbers on this page refer to it.
The numbers in this post come from those three sources. Where a chart’s values were read off a paper figure rather than taken from a table, its caption says so. The background facts below were checked against each linked page; the one whose URL no longer resolves says so in its entry.
CXL and hardware
- CXL Consortium: About CXL, and the 2.0/3.0/3.2 specification releases
- Sharma, Blankenship, Berger. An Introduction to the Compute Express Link (CXL) Interconnect. ACM Computing Surveys, 2024
- Sun et al. Demystifying CXL Memory with Genuine CXL-Ready Systems and Devices. MICRO 2023
- Liu et al. Systematic CXL Memory Characterization and Performance Analysis at Scale (Melody). ASPLOS 2025
- Just How Bad Is CXL Memory Latency? The Next Platform, 2022
- Jain et al. (AMD). Memory Sharing with CXL: Hardware and Software Design Approaches. arXiv, 2024
- Micron's Perspective on Impact of CXL on DRAM Bit Growth Rate. Whitepaper, 2023 (the chart behind the memory-wall figure, as reproduced in our talk; Micron's original URL no longer resolves)
- Micron CZ120 CXL memory expansion modules: launch announcement, 2023
- Samsung CMM-D (CXL Memory Module, DRAM): product page
- CXL Memory Disaggregation and Tiering: Lessons Learned from Storage. SNIA SDC 2023
Tiering and sharing systems
- Agarwal & Wenisch. Thermostat: Application-transparent Page Management for Two-tiered Main Memory. ASPLOS 2017
- Raybuck et al. HeMem: Scalable Tiered Memory Management for Big Data Applications and Real NVM. SOSP 2021
- Weiner et al. TMO: Transparent Memory Offloading in Datacenters. ASPLOS 2022
- Li et al. Pond: CXL-Based Memory Pooling Systems for Cloud Platforms. ASPLOS 2023
- Maruf et al. TPP: Transparent Page Placement for CXL-Enabled Tiered-Memory. ASPLOS 2023
- Huang et al. Tigon: A Distributed Database for a CXL Pod. OSDI 2025
Data structures and protocols
- Papamarcos & Patel. A low-overhead coherence solution for multiprocessors with private cache memories (MESI's origin). ISCA 1984
- Calciu et al. Black-box Concurrent Data Structures for NUMA Architectures (Node Replication). ASPLOS 2017
- Leis, Kemper, Neumann. The Adaptive Radix Tree: ARTful Indexing for Main-Memory Databases. ICDE 2013
- Leis et al. The ART of Practical Synchronization (Optimistic Lock Coupling). DaMoN 2016
- O'Neil et al. The Log-Structured Merge-Tree (LSM-Tree). Acta Informatica, 1996
- RocksDB (Meta) and LevelDB (Google): LSMs in production