Skip to content

One memory,
many machines

CXL lets whole servers share the same terabytes of RAM. The hardware arrived with one catch, and the catch is where things get interesting.

Four hosts, one shared CXL memoryFour outlined host machines sit at the corners. Beams from each converge on a single memory slab labeled CXL memory, terabytes, shared by all four at once.Host 1Host 2Host 3Host 4CXL MEMORY · TBsone device · four address spaces · zero copies

Blog by Kiran Hombal and Jiyu Hu

Based on our papers Megalon: Efficient Data Sharing for Partly Coherent CXL Memory (OSDI 2026, best-paper nominee) by Jiyu Hu, Seokjoo Cho, Landon Johnson, Kiran Hombal, Shreesha G. Bhat, Marcos K. Aguilera, Ramnatthan Alagappan and Aishwarya Ganesan and Disk-Based LSMs: An Unexpectedly Good Index for Partly Coherent CXL Memory (SOSP 2026) by Kiran Hombal, Jiyu Hu, Marcos K. Aguilera, Ramnatthan Alagappan and Aishwarya Ganesan, and on a talk we gave at ByteDance.

DASSL Lab, UIUC and NVIDIA

1. Memory that belongs to no one machine

What this post covers, who it is for, and the one question it answers.

Most servers keep their memory to themselves; sharing data means messages over a network and a copy on each side. CXL breaks that habit. Several machines plug into the same memory and read and write it with the instructions they use for their own DRAM. No network stack on the access path, no serialization, no copies.

This post grew out of a talk we gave at ByteDance and two papers from our lab at UIUC. Both start from the question the hardware leaves open: the memory is shared, but can you actually share data on it? The answer is harder, and more fun, than it looks.

We assume no prior exposure to CXL or caches. If you know what RAM is, you have everything you need. In a hurry? Jump to Megalon; a short recap waits there.

Papers
2
OSDI '26 · SOSP '26
Best-paper nominee
1
Megalon at OSDI '26
Peak speedup
15.1×
Prism over a Tigon-style index
Shared capacity
TBs
per device, in the model both papers target

2. The memory wall

Start with the problem that makes CXL worth inventing.

Cores multiplied. Memory per core shrank.

Talk, slide 5
CPU core count versus memory capacity per core, 2012 to 2025A qualitative redraw. Both series indexed to 2012 equals 100 on a log axis. Core count rises to roughly sixteen times its 2012 level while memory per core falls to roughly one third; the 2024 and 2025 points are the source's projections, drawn dashed.core countmemory / core
  • CPU core count
  • Memory capacity per core
  • projected by the source
Core count against memory capacity per core, indexed to 2012 = 100, log scale. A qualitative redraw of the Micron chart in our talk; the source marks 2024 and 2025 as projections, drawn dashed. The shapes are the argument.
Show the numbers
Cores multiplied. Memory per core shrank.
YearCore countMemory / coreProjected
2012100100
201413096
201616088
201825076
202040062
202260050
2024110038yes
2025160032yes

A line chart from 2012 to 2025. CPU core count rises steeply to roughly sixteen times its 2012 level, while memory capacity per core falls to roughly a third of its 2012 level. The last two points are the source's projections.

Server CPUs went from a dozen cores in 2012 to well over a hundred today. DRAM capacity per core has been falling for a decade. Two limits bind at once. First, DRAM density grows more slowly than core counts. Second, a CPU cannot simply add memory channels: each DDR channel takes roughly 200 signal pins, and pins are among the scarcest resources on a package.

Practitioners feel this as a wall. Databases, caches, and feature stores want more memory per machine every year, and the machine cannot offer it. Meanwhile a lot of memory sits stranded: a server has rented out all its cores while gigabytes of its DRAM stay idle. Microsoft measured up to 25% of Azure’s DRAM stranded this way, with DRAM up to 50% of server cost.

Unfortunately, you cannot fix a pin-count problem with software. You need a new wire.

3. Enter CXL

Memory over PCIe, spoken to with ordinary loads and stores.

Compute Express Link (CXL) is that wire. It is an open standard built on PCIe, the lanes that already carry your GPU and NVMe drives. A CXL memory device is a box of DRAM on PCIe, and the trick is how the CPU talks to it: not with I/O commands like a disk, but with ordinary loads and stores. Once mapped, it is just more memory to a program.

The economics come from the pins. A DDR5-6400 channel moves ~51 GB/s over ~200 signal pins. An x16 PCIe 5.0 link moves ~63 GB/s per direction over 64 signal pins, and PCIe 6.0 doubles that. Per pin, the serial link carries roughly 4× the bandwidth, so a CPU can afford far more of them. That is how CXL adds terabytes where the DDR bus cannot.

The price is distance. Local DRAM answers in roughly 110 ns; CXL memory takes two to three times that, about the cost of reaching the other socket of a two-socket server (a NUMA hop), which software tolerates every day. It is still far closer than a remote read over RDMA (microseconds) or an SSD (tens of microseconds).

Expansion products ship today from Samsung and Micron, and recent AMD and Intel CPUs speak CXL natively. Multi-host sharing, the subject of the rest of this post, is newer; we will be clear about what ships and what is only expected.

Two ways out of the socket

Sharma et al., CXL survey · PCI-SIG
DDR versus CXL pin budgetsA CPU connects to local DRAM through a wide ribbon of about 200 signal pins carrying about 51 gigabytes per second, and to CXL memory through a narrow ribbon of 64 signal pins carrying about 63 gigabytes per second in each direction at PCIe 5.0 rates.CPULOCAL DRAMhundreds of GBDDR5-6400 · ~200 signal pins · ~51 GB/sCXLCTRLCXL MEMORYterabytesx16 PCIe 5.0 · 64 signal pins · ~63 GB/s per direction
Bandwidth per signal pin is the whole story: close to 1 GB/s per pin for the PCIe link against about 0.25 for DDR5. The survey we cite quotes 256 GB/s in its introduction; its own bandwidth section gives the per-direction figure used here.

A diagram comparing a DDR channel, about 200 signal pins for about 51 gigabytes per second, against a CXL link over PCIe, 64 signal pins for about 63 gigabytes per second in each direction.

Where CXL lands in the hierarchy

Melody (ASPLOS '25); Next Platform
Approximate access latency by tierA log-scale dot plot from L1 cache near one nanosecond to NVMe SSDs near one hundred microseconds. CXL memory sits at roughly 170 to 400 nanoseconds, between local DRAM and an RDMA round trip. Each dot is a representative point inside a range; the table gives the ranges and where each one comes from.L1 cache~1 ns on the core itselfL2 cache~4 ns private per coreL3 cache~15 ns shared on the dieLocal DRAM~111–117 ns over the DDR busCXL memoryabout one NUMA hop away ~170–400 nsRDMA readnetwork round trip ~2–10 µsNVMe SSDflash storage ~10–100 µs
Approximate latency, log scale. Each dot is a point inside a range; the table says which rows are measured and which are textbook orders of magnitude. CXL sits in the long gap between DRAM and the network.
Show the numbers
Where CXL lands in the hierarchy
TierLatencyNoteBasis
L1 cache~1 nson the core itselftextbook order of magnitude
L2 cache~4 nsprivate per coretextbook order of magnitude
L3 cache~15 nsshared on the dietextbook order of magnitude
Local DRAM~111–117 nsover the DDR busmeasured on three Melody platforms
CXL memory~170–400 nsabout one NUMA hop awayMelody measured 214–394 ns; industry quotes 170–250 ns
RDMA read~2–10 µsnetwork round tripcommonly reported range
NVMe SSD~10–100 µsflash storagecommonly reported range

A log-scale ladder of approximate access latencies from L1 cache at about one nanosecond, through local DRAM near 110 nanoseconds, CXL at roughly 170 to 400 nanoseconds, RDMA at microseconds, and NVMe SSDs at tens of microseconds.

4. One memory, many machines

The three CXL capabilities that matter here, and the one that changes the rules.

CXL 1.1 (2019) does expansion: one host, more memory. CXL 2.0 (2020) adds switching and pooling: a rack-level pool carved into segments, each still owned by one host at a time. CXL 3.0 (2022) and its successors (3.2 in 2024, 4.0 in late 2025) add the radical part: multi-host shared memory. Several machines map the same region into their address spaces, at the same time, and all of them issue loads and stores against it.

Stop and appreciate how strange that is. Two servers with separate power supplies, operating systems, and failure domains dereference pointers into the same DRAM. A structure written by host 1 is already there when host 2 looks. Databases, key-value stores, and file systems could keep one copy of their data where every host can reach it.

That is shared CXL memory, the “one memory, many machines” of the title. If it sounds too good to be true, it is, exactly once. To see the catch we need a short detour into the deepest habit CPUs have: caching.

Three generations, three relationships to memory

CXL Consortium spec releases
CXL 3.x: SharingFabrics and true multi-host sharing: several machines map the same region, at the same time.Host 1Host 2Host 3Host 4the same bytes, in four address spaces at onceSharing: Fabrics and true multi-host sharing: several machines mapthe same region, at the same time.
Toggle the generations. Expansion gives one host more memory; pooling gives each host its own segment; sharing lets every host map the same region at once. The overlap is the new thing.

An interactive diagram showing CXL 1.1 with one host and one expander, CXL 2.0 with a switch assigning pool segments to individual hosts, and CXL 3.0 with four hosts mapping one shared region simultaneously.

5. Cache coherence, from zero

Why CPUs cache, why caching breaks sharing, and the protocol that fixes it inside one machine.

Most loads never reach DRAM. A trip to memory costs about 100 ns, so every CPU keeps recently used data in small, fast, private caches and serves most loads from there.

Caching creates an obvious hazard: copies. If CPU 1 has cached value A and CPU 2 changes A in memory, CPU 1’s copy is now wrong. Within one machine, hardware solves this with cache coherence: the CPUs watch each other’s memory traffic and follow a protocol that keeps every cached copy of a location consistent.

The classic protocol is MESI, after the four states a cached line can be in: Modified, Exclusive, Shared, Invalid. (It is also the Illinois protocol; it was invented where our lab sits.) You do not need the details, but the feel of it matters for everything that follows. Step through the widget once.

The takeaway is that coherence is not free. A write to shared data invalidates every other cache’s copy, and to do that the hardware tracks who might hold one. The bookkeeping grows with the number of caches and the amount of memory tracked. Hold that thought.

MESI, one write, no lies

Talk, slides 12–29
Three CPUs, one memoryMemory holds A = 1. No CPU has it cached yet.CPU 1cache emptyCPU 2cache emptyCPU 3cache emptyBUSMAIN MEMORYA = 1

Step 1 / 10 · Three CPUs, one memory

Memory holds A = 1. No CPU has it cached yet.

Three CPUs share value A through their private caches. Play the sequence: three reads spread copies, one write destroys them, and the writeback repairs memory. States: Modified · Exclusive · Shared · Invalid.

An interactive step-through of the MESI protocol: three CPUs read value A and hold Shared copies; CPU 2 invalidates the others and writes A equals 2, holding a Modified line; CPU 3's next read forces a writeback, leaving both with Shared copies of the new value.

6. The catch: partly coherent memory

The compromise the hardware is expected to make, and what silently breaks inside it.

Now scale that bookkeeping up to what CXL 3.x allows: several machines, dozens of caches each, sharing terabytes. To behave like memory inside one machine, the hardware would have to run coherence across hosts, over PCIe, for every cache line. The vendors are blunt about the arithmetic: AMD, Micron, and Samsung all say the machinery involved stops scaling somewhere between dozens and a few hundred megabytes.

So the hardware is expected to compromise, in what we call the partly coherent model, which both of our papers target. The memory splits in two. A small coherent region, the SCR, is a few hundred MB that hardware keeps coherent across hosts, like the widget above. A large non-coherent region, the LNR, is the remaining several TB, where the hardware does nothing about coherence across hosts. Each host’s own caches still work; they just never hear from the others.

What goes wrong in the LNR? Exactly what MESI exists to prevent, except now nobody prevents it. Host 1 reads object A and caches it. Host 2 overwrites it with A′. On one machine that write would invalidate host 1’s copy; across machines, in the LNR, no invalidation is ever sent. Host 1 reads again, its cache says “I have that,” and it returns the stale value. Silently. Your database just served data that is no longer fresh.

Why this post exists

Sharing memory is a hardware problem with a hardware solution. Sharing data correctly, on memory that is only partly coherent, is a software problem, and it is wide open. Both of our papers live in this gap.

The two regions, side by side

Megalon §2.1 · Tigon, citing AMD
SCR versus LNR, not to scaleA bar representing an illustrative four terabytes of CXL memory. The hardware-coherent region, a few hundred megabytes, would be far thinner than a pixel; it is drawn as a widened marker at the left edge and magnified in a callout above.SCR · 100s MBhardware coherentLNR · several TB · no coherence across hosts04 TB

A few hundred MB on a device of several TB: roughly one part in ten thousand. The marker is widened so that it is visible at all. Everything the hardware promises about coherence across hosts happens inside it.

A device of several TB with a coherent region of a few hundred MB: roughly one part in ten thousand. The marker is widened to be visible at all; at true scale you could not see it, which is the point.

A long bar representing several terabytes of CXL memory. The hardware-coherent region, a few hundred megabytes, would be far thinner than a pixel; it is drawn as a widened marker at the left edge and shown magnified.

The stale read, step by step

Talk, slide 31
Host 1 reads A from the LNRThe value drops into host 1’s CPU cache, as every read does.Host 1cacheAHost 2cacheread ALNR · no coherenceA

1 / 4 · Host 1 reads A from the LNR

The value drops into host 1’s CPU cache, as every read does.

Host 2’s write reaches the memory, but no invalidation reaches host 1’s cache, so host 1’s next read is answered by its own cache, wrongly.

An animation of two hosts sharing a value in the non-coherent region. Host 1 caches A; host 2 writes A-prime; host 1 reads again and its cache serves the stale A with a warning flag.

7. Sharing anyway, and how it collapses

If the hardware will not keep caches honest in the LNR, software has to. It works, until it churns.

The first idea everyone has is to keep shared data inside the SCR. But a few hundred megabytes shared among many hosts is nothing; the whole point was the terabytes. Rejected.

The serious idea keeps the data in the vast LNR and small metadata about it in the SCR, where hardware coherence makes the metadata trustworthy. Tigon (OSDI ’25), an early database for a multi-host CXL pod, works this way. Each shared object gets a coherence record, think of a version counter plus a lock, and a host checks the record before touching the object; if someone wrote it since the host last looked, the host flushes its stale cache lines and re-reads from CXL. Objects are coarse, rows or key-value pairs of a few KB rather than cache lines, so the metadata stays small, and a shared index in the SCR lets hosts find each object and its record. We call this scheme hardware-coherent metadata-based sharing, HCMeta for short. The first figure on the right steps through it.

HCMeta is correct, and for Tigon’s workload of occasional cross-partition transactions it works well. But push more data into sharing and a problem surfaces: the metadata grows with the number of shared objects, and the SCR does not. In the Megalon paper’s example with 40-byte keys, HCMeta spends 52 bytes of SCR per object, so a 100 MB SCR caps sharing at about 1.9M objects. A terabyte of small objects holds hundreds of millions.

At the cap, HCMeta unshares an old object to make room for a new one. Objects start rotating through the tiny window of shareability, and every rotation is expensive: the host that wants an unshared object must ask its owner and wait, about 55 µs per round trip in Tigon’s artifact. We call the rotation churn. In that artifact, growing the dataset from 2.4M to 24M objects at 20% cross-host transactions cuts throughput by 10×. A cliff, not a slope.

Summary. The partly coherent model makes shared metadata the key to correct sharing, then makes the only place it can live absurdly small. So our first paper asks: how can hosts share a huge number of objects when even the metadata is too big for the coherent region?

HCMeta, step by step

Megalon slide 7 · §2.2
Host 1 reads AHost 1 reads A from the LNR and notes its version from A’s coherence record in the SCR: 0.Host 1cacheA: 0Host 2cacheread Aversion 0SCR · hardware coherentcoherence record: 0LNR · no coherenceA

1 / 4 · Host 1 reads A

Host 1 reads A from the LNR and notes its version from A’s coherence record in the SCR: 0.

A version record per object in the SCR, which hardware keeps coherent; a reader that finds a newer version drops its copy and re-reads. Simplified: the lock, the fences, and the mid-read retry are left out.

An animation of HCMeta. Host 1 caches A at version 0; host 2 writes A-prime to the LNR and bumps A's version in the SCR to 1; host 1 checks the SCR, sees 0 does not match 1, invalidates its cached A, and re-reads A-prime from the LNR.

The collapse, measured

Megalon, Figure 1(c)
HCMeta throughput collapses once the dataset outgrows the coherent regionWith unlimited SCR throughput holds near 15 Mops per second across every dataset size. With a realistic 100 MB SCR it falls from 15.2 to 1.0 between 2.4 and 4.8 million objects and keeps sinking. Values were read off the paper's plot.the SCR fills hereunlimited SCR (unrealistic)
  • HCMeta, 100 MB SCR
  • HCMeta, unlimited SCR
A key-value store using HCMeta, read-only, 100 MB SCR, values read off the paper’s plot. The unlimited-SCR variant stays flat; the real one falls from 15.2 to 1.0 Mops/s (million operations per second) between 2.4M and 4.8M objects and keeps sinking.
Show the numbers
The collapse, measured
DatasetHCMeta, 100 MB SCRUnlimited SCR
1.2M objects15.5 Mops/s15.5 Mops/s
2.4M objects15.2 Mops/s15.3 Mops/s
4.8M objects1.0 Mops/s15.1 Mops/s
7.2M objects0.9 Mops/s15.0 Mops/s
12M objects0.7 Mops/s14.6 Mops/s
18M objects0.6 Mops/s13.8 Mops/s

A line chart of throughput versus dataset size. With unlimited SCR, throughput holds near 15 million operations per second. With a realistic SCR, throughput collapses from about 15 to about 1 million operations per second once the dataset passes a few million objects.

Paper one · OSDI 2026 · Best-paper nominee

8. Megalon

Share the data. Split the metadata.

Skipped ahead? The setup in brief

CXL lets hosts share terabytes, but the model we target keeps only a small coherent region (the SCR, a few hundred MB) coherent across hosts; the large non-coherent region (the LNR) can silently serve stale cached data. The standard fix, HCMeta, keeps metadata for every shared object in the SCR, about 52 bytes each with 40-byte keys, so a 100 MB SCR caps sharing near 1.9M objects and throughput collapses past it. ↑ The partly coherent model · ↑ How HCMeta collapses

Megalon starts where HCMeta breaks, with one observation: the metadata is two very different things glued together. The index, which maps each object’s ID to its location, is big (every object’s key, tens of bytes each) but cold: it changes on inserts, deletes, and sharing-state changes, far less often than data is written. The coherence records are tiny (a lock bit and a counter, 4 bytes) but hot: touched on every write.

HCMeta stuffs both into the SCR. Megalon’s key idea is to split them and share each the way its nature demands. The hot, tiny records stay physically shared in the SCR, where hardware coherence is exactly the right tool; at 4 bytes, about 13× more of them fit than HCMeta’s 52-byte bundles. The big, cold index is logically shared: every host keeps a full replica in its local DRAM, of which it has hundreds of gigabytes, roughly 1000× the SCR. A 100 MB SCR that capped HCMeta near 2 million objects lets Megalon share about 25 million.

At first sight, replicating the index recreates the problem: N replicas must now be kept consistent, the very thing coherence was for. However, the index is the cold half. It changes rarely, so keeping replicas in sync is cheap, given a mechanism to do it.

That mechanism is a shared log in the CXL memory itself: an ordered sequence of entries in the LNR, with only its head and tail pointers in the SCR. To change the index, a host claims the next entry with a compare-and-swap (an atomic instruction) on the coherent tail, writes the entry, and flushes it. Every host applies entries in order to its replica, reading them past its cache, and checks the tail before any index read. One totally ordered log at memory speed: replicas cannot drift, and no message-based agreement is needed. The idea comes from Node Replication, which synchronized data structures across NUMA sockets the same way.

The log supports two further coherence techniques.

Dynamic coherence records. Objects that are only read need no record, since nothing changes under the readers. So Megalon allocates records only for objects that are read and written, and when the SCR fills it demotes cold ones and reassigns their records, announcing each change through the log. Read-shared objects are no longer capped by the SCR, and churn becomes an appended entry rather than a round trip: about 8× cheaper.

Dual-path coherence. With records coming and going, an object can gain or lose its record while you are reading it, so the record alone cannot prove your cache is fresh. Hosts therefore also check, at the end of a read, whether the log recorded an allocation event for the object; if so, they flush and retry.

Other uses of the log. Since every index change is ordered by the log, Megalon can also keep a host’s private objects in its local DRAM, and cache read copies of hot shared objects there, where access is faster than CXL (paper §3.5).

Against HCMeta, Megalon delivers 15× on read-only workloads with large datasets and 10× at 5% writes once metadata outgrows the SCR. The cost is host DRAM for the index replicas, 7.6% more memory in the 24M-object read-only run.

Split sharing: each half where it belongs

Megalon §3.2, Figure 2
Split metadata sharingHosts hold replicated indexes in local DRAM. The CXL memory holds coherence records in the small coherent region and data objects in the large non-coherent region. Occupancy bars compare 52 bytes per object for HCMeta (with 40-byte keys) against 4 bytes for Megalon.Host 1 · local DRAMINDEX REPLICAHost 2 · local DRAMINDEX REPLICAHost N · local DRAMINDEX REPLICACXL MEMORYSCRcoherence records · 4 Bdata objects · 1–4 KB eachLNRSCR SPENT PER SHARED OBJECTHCMeta52 BMegalon4 B · ~13× more objects

The index moved to where memory is plentiful; only what must be coherent stays where coherence lives.

The big, cold index is replicated into each host’s DRAM; only the tiny, hot records occupy the SCR. For the paper’s 40-byte-key example, 52 bytes of SCR per object becomes 4.

A diagram of Megalon's split: index replicas in each host's local DRAM point to data objects in the LNR and to small coherence records in the SCR. Occupancy bars compare 52 bytes per object for HCMeta against 4 for Megalon.

A shared log, in the memory itself

Megalon §3.3, §4
The CXL shared logA log of ordered entries lives in the non-coherent region; only its head and tail pointers live in the coherent region. Hosts append via compare-and-swap on the tail and replay entries into their index replicas.Host 1Host 2append · compare-and-swap on the tailreplay → replicaLNR1234567shared log · ordered · entries flushed on write, read past the cacheSCRhead = 1tail = 7

Only the head and tail pointers need hardware coherence. The entries sit in the LNR, written with a flush and read past the cache, and order does the rest.

Hosts claim entries with a compare-and-swap on the tail and replay them into their replicas. Ordering without messages; the only coordination is atomic operations on two pointers in the SCR.

A shared log laid out in the non-coherent region with its head and tail pointers in the coherent region. Hosts append entries and replay them into local index replicas.

Flat where HCMeta collapses

Megalon Figure 6(a) · Talk, slide 51
Megalon versus HCMeta as the number of shared objects growsGrouped bars, 5 percent writes and a 200 MB coherent region: Megalon holds near 17 Mops per second from 4.8 to 24 million objects while HCMeta falls from 16.6 to 1.7, about 10 times lower at 24 million. Values read off the paper's Figure 6(a).4.8M7.2M12M1.717.424Mshared objects10×
  • HCMeta
  • Megalon
5% writes, Zipfian, 200 MB SCR, values read off the paper’s Figure 6(a). From 4.8M to 24M shared objects HCMeta churns and collapses while Megalon holds near 17 Mops/s: 10× at 24M. Even with a 32 MB SCR (not shown) Megalon keeps an 8.4× lead: smaller records, cheaper churns.
Show the numbers
Flat where HCMeta collapses
Shared objectsHCMetaMegalon
4.8M16.6 Mops/s17.7 Mops/s
7.2M5.7 Mops/s17.2 Mops/s
12M2.5 Mops/s17.4 Mops/s
24M1.7 Mops/s17.4 Mops/s

Grouped bars of throughput for 4.8, 7.2, 12 and 24 million shared objects under 5 percent writes. Megalon stays near 17 million operations per second at every size; HCMeta falls from 16.6 to 1.7, a ten times gap at 24 million objects.

9. The index problem

Sharing objects is not enough. A store also has to find them.

“Give me every order between Monday and Wednesday.” Queries like that need a range index, the sorted tree at the heart of every database, living in CXL where all hosts can search it. An index turns out to be the worst-case tenant for partly coherent memory, for a reason that takes one paragraph to see.

The obvious move is to port a state-of-the-art in-memory index, the adaptive radix tree (ART) or a B+-tree, with the software-coherence recipe: nodes in the LNR, one version number per node in the SCR, bumped on every change and checked by readers. These indexes already keep per-node versions for concurrency control, so the port is natural. We built both: SC-ART and SC-BTree, SC for software-coherent.

Unfortunately, both fail, in opposite directions. Small nodes, too many versions. ART’s nodes are small and numerous: 100M keys make ~40M inner nodes, and at 8 bytes per version only about 40% of them fit in a 128 MB SCR. A version that spills into the LNR needs its own freshness check, a cache flush on every access, whether or not anything changed; on read-heavy workloads SC-ART runs about 10× slower than with unlimited SCR. Big nodes, false invalidations. Fatten the nodes to 128 keys and the versions fit, but one version now covers 128 keys: update any one and every host that cached any of the other 127 must flush and re-read the whole node. Under 50% writes, SC-BTree collapses to 1.8 Mops/s.

The trade-off is fundamental, and the paper names its cause: the updatable surface area, the part of a structure that can be modified in place. In ART and B-trees any node from root to leaf can be rewritten, so the surface is the entire index and every node needs a version. Track it finely and versions overflow; track it coarsely and false invalidations eat you.

So the real question, the one our second paper poses, is not how to port an index. It is: what is the right way to build an index for partly coherent memory?

Larger nodes trade version space for false invalidations

Prism §3, Figures 1–2
Larger nodes trade version space for false invalidationsAt small node sizes the version numbers overflow the coherent region; at large node sizes they fit but a single update falsely invalidates every other key in the node. The two ends are measured; the middle is a model.VERSION NUMBERS vs THE SCR (128 MB)spilled into the LNR7.0M versions needed · 16M fit · all fit · modeledONE KEY WRITTEN → KEYS FLUSHEDfan-out 23: the write touches 1 key, and invalidates 22 that nobody changed

modeled: versions fit, bystanders pay

every version fits, but one write already invalidates 22 unmodified keys; shrink the nodes to spare them and the versions overflow again

The ends are the paper’s measurements: SC-ART at ART’s smallest node and SC-BTree at 128 keys per node. In between is a model that scales the paper’s 40M ART nodes inversely with node size, labeled as such. There is no throughput axis, because the paper measures only the ends.

An interactive slider over index node size. At small node sizes, tens of millions of version numbers overflow the coherent region. At large node sizes, a single update falsely invalidates every key in a 128-key node. The two ends are measured; the middle is a model.

Paper two · SOSP 2026

10. Prism

The right index was on disk all along.

Prism finds the answer in the least likely place: a data structure family built to reduce disk writes. Log-structured merge trees (LSMs), the engine inside RocksDB, LevelDB, Bigtable, and Cassandra, were designed around a disk’s hatred of random writes. All writes land in one small in-memory buffer, the memtable. When it fills, it is frozen and written out as an immutable sorted file, an SSTable; background compaction merges SSTables down through levels by replacing old files with new ones. Below the memtable, nothing is modified in place.

Read that again with the last section’s lens. The updatable surface area of an LSM is the memtable, a few dozen megabytes. Perhaps surprisingly, a structure built for a different medium has exactly the shape partly coherent memory demands. The memtable is small, so it lives entirely inside the SCR, coherent across hosts with no software checks. SSTables are immutable, so they need no per-node tracking and sit in the terabyte LNR. The one care is recycling: each SSTable carries a strictly increasing ID in a small manifest in the SCR, and a host that meets a new ID flushes that region once. Coarse, safe, and rare.

A straight port, SC-LSM, already beats SC-BTree by 2.2× at 50% writes, because false invalidations simply stop existing. But it inherits the LSM’s classic ailment: at high write rates the memtable flushes faster than compaction can drain, level-0 SSTables pile up, and a read must probe every one of them, 49 per read at 50% writes.

Here Prism notices that the old cure was disk-specific. On disk you cannot update upper LSM levels in place, because random I/O is what LSMs exist to avoid. In memory, in-place updates are cheap. So Prism inserts a middle tier, the Bounded Updatable Layer (BUL), an in-place-updatable tree that absorbs repeated writes, so hot keys overwrite themselves instead of spawning SSTable after SSTable. Doesn’t that reopen the updatable-surface problem? It would, so Prism bounds it: the BUL is sized so that all of its per-node versions fit in the SCR. In the paper’s ablation at 5% writes, the BUL alone cuts SSTables probed per read from 15 to 2.

One more inversion completes the design. The memtable stops being a write buffer and becomes a cache for proven-hot keys. New keys go straight to the BUL; only a key written again is promoted. One-hit wonders never waste the scarcest bytes in the system.

Prism’s three tiers and their coherence mechanisms
TierLives inCoherenceCost
MemtableSCR (64 MB)Hardware coherenceNo software checks
BULData in LNR; version numbers in SCR (62 MB)Per-node version checkFine-grained, bounded
SSTablesLNR (the terabytes)Immutable + rising IDs; one flush on first sightCoarse but safe

A 2 MB manifest in the SCR lists the live SSTables: 64 + 62 + 2 = 128 MB.

Prism's results

On emulated CXL with 128 MB of SCR: up to 6.3× over SC-ART on read-only workloads, the LSM’s traditional worst case; 7.4× and 9.4× over SC-BTree on the write-heavy YCSB-A and YCSB-F; 3.6–5.7× over Chime-CXL, an RDMA index ported to CXL; and 8.6–15.1× over Tigon-SWcc, a Tigon-style index. With 128 MB of SCR, Prism matches the throughput SC-ART needs unlimited SCR to reach on YCSB-A. Against Megalon’s own index it wins with writes and loses on read-only, where Megalon’s replica sits in local DRAM.

Three tiers, one coherence mechanism each

Prism §5, Figure 5
Prism: memtable, BUL, SSTablesThe coherent region holds the manifest, the memtable and the BUL’s version numbers. The non-coherent region holds the BUL’s data and immutable SSTable levels.SCR · 128 MB · hardware coherentMANIFEST2 MBMEMTABLE · 64 MBa cache for proven-hot keys · all in-place writesBUL VERSION #s62 MB · one per node── the coherence boundary ──BUL · in-place updatable · data here, versions abovebounded so every version number fits in the SCRL0flush ↓L1Lncompact ↓

SSTables are immutable for their whole lifetime, so they need no coherence tracking at all; strictly increasing IDs handle the one case that remains, recycling.

The memtable, a hot-key cache, lives in the SCR; the BUL keeps its data in the LNR with every version anchored in the SCR; SSTables below are immutable. The updatable surface is exactly as large as the SCR can protect.

Prism's architecture: a memtable inside the coherent region, a bounded updatable layer with data in the non-coherent region and version numbers in the coherent region, and immutable SSTable levels below.

A key-value store under YCSB

Prism Figure 9 · Talk, slide 90
Prism versus ported in-memory indexes under YCSBGrouped bars read off the paper's Figure 9: Prism reaches 11.8 Mops per second on YCSB-A, 18.5 on B, 21.0 on C and 14.1 on F; SC-BTree falls to 1.6 and 1.5 on the two write-heavy workloads; SC-ART stays at or below 3.3 throughout.YCSB-A50% writesYCSB-B5% writesYCSB-Cread-only1.91.514.1YCSB-F50% RMW9.4×
  • SC-ART
  • SC-BTree
  • Prism
Values read off the paper’s Figure 9. The heavier the write pressure, the wider Prism’s lead: false invalidations ruin SC-BTree on A and F, and version spill drags SC-ART everywhere. Even on read-only C, the LSM’s traditional weakness, Prism edges ahead. F is 50% read-modify-write (RMW).
Show the numbers
A key-value store under YCSB
WorkloadSC-ARTSC-BTreePrism
YCSB-A · 50% writes2.2 Mops/s1.6 Mops/s11.8 Mops/s
YCSB-B · 5% writes3.1 Mops/s13.0 Mops/s18.5 Mops/s
YCSB-C · read-only3.3 Mops/s19.5 Mops/s21.0 Mops/s
YCSB-F · 50% RMW1.9 Mops/s1.5 Mops/s14.1 Mops/s

Grouped bars across YCSB workloads A, B, C and F. Prism reaches 11.8, 18.5, 21.0 and 14.1 million operations per second; SC-BTree collapses to 1.6 and 1.5 on the two write-heavy workloads; SC-ART stays below 3.5 throughout.

The honest chart: range scans

Prism Figure 8 · Talk, slide 110
Prism's range-scan throughput relative to SC-BTreeA line rising from 0.7 times at five percent writes, crossing parity at 10 percent, and reaching 6.0 times at fifty percent writes.parity: SC-BTree wins below this line0.7× at 5%the flip6.0×
A real trade-off, with both throughputs read off the paper’s Figure 8. At 5% writes SC-BTree wins range scans, because sorted neighbors in one node are a genuine advantage. The lead flips to Prism at 10% writes and reaches 6.0× at 50%.
Show the numbers
The honest chart: range scans
Write ratioSC-BTreePrismPrism ÷ SC-BTreeWinner
5%13.8 Mops/s9.8 Mops/s0.7×SC-BTree
10%9.0 Mops/s10.1 Mops/s1.1×Prism
20%4.0 Mops/s10.9 Mops/s2.7×Prism
30%2.7 Mops/s10.5 Mops/s3.9×Prism
40%2.0 Mops/s10.2 Mops/s5.1×Prism
50%1.6 Mops/s9.5 Mops/s6.0×Prism

A line of Prism's range-scan throughput relative to SC-BTree as write ratio grows: about 0.7 times at five percent writes, crossing 1.0 at ten percent, reaching six times at fifty percent.

How we got here

Selected milestones: CXL data systems are the third act of a decade-long story.

Act one · Tiering and pooling (2017–2023)

Memory could be slower but bigger; the game was deciding which pages live where, one host at a time. Pond added pooling: segments of one device handed to different hosts.

  1. Thermostat, ASPLOS ’17

    First fully application-transparent page management for two-tier main memory.

  2. HeMem, SOSP ’21

    Tiered memory management built to scale for big-data applications.

  3. TMO, ASPLOS ’22

    Transparent memory offloading, in production across Meta's fleet.

  4. Pond, ASPLOS ’23

    CXL memory pooling for Azure, within tight latency targets.

  5. TPP, ASPLOS ’23

    CXL page placement for the Linux kernel.

Act two · First sharing (2025)

Tigon faces the partly coherent model first. Its HCMeta-style design is the baseline our work stands on.

  1. Tigon, OSDI ’25

    An early database for a multi-host CXL pod, and the first system built around the partly coherent model.

Act three · General sharing (2026 →)

Megalon makes many objects shareable despite the tiny SCR; Prism gives partly coherent memory a range index. The rest of the stack is next.

  1. Megalon, OSDI ’26

    General data sharing beyond the coherent region's capacity.

    our lab

  2. Prism, SOSP ’26

    A range index built for partly coherent memory.

    our lab

Takeaways

Five things worth keeping if you keep nothing else.

  1. CXL exists because of pins. Core counts outran the 200-pin DDR channel. CXL adds terabytes over a 64-pin PCIe link at roughly the latency of a NUMA hop, a new rung in the memory hierarchy.

  2. Shared addresses are not shared coherence. In the model the hardware is heading for, a few hundred MB stay coherent across hosts; the terabytes beyond can silently serve stale caches. Sharing memory is solved. Sharing data is not.

  3. The coherent region is a scarce metadata budget. Every correct sharing scheme keeps per-object metadata there, and naive schemes churn and collapse when it fills. Megalon splits the metadata: replicate the big cold index in host DRAM, keep only 4-byte hot records coherent, and order everything through a shared log in CXL, for up to 10× under writes and 15× on reads.

  4. Updatable surface area decides which structures survive. If most of a structure can be modified in place, partly coherent memory punishes it: version overflow on one side, false invalidations on the other.

  5. Old designs, new medium. Disk-born LSMs, with a tiny mutable buffer and an immutable everything else, fit partly coherent CXL better than the in-memory indexes we ported: up to 9.4× over them and 15.1× over a Tigon-style index.

Sources & further reading

Everything this post leans on: our papers first, then every external reference.

Sources of truth

  • Megalon: Efficient Data Sharing for Partly Coherent CXL Memory. OSDI 2026 (best-paper nominee).

    Jiyu Hu, Seokjoo Cho, Landon Johnson, Kiran Hombal, Shreesha G. Bhat, Marcos K. Aguilera, Ramnatthan Alagappan, Aishwarya Ganesan

    paper (PDF) · slides (PDF) · USENIX page · artifact on GitHub

  • Disk-Based LSMs: An Unexpectedly Good Index for Partly Coherent CXL Memory. SOSP 2026.

    Kiran Hombal, Jiyu Hu, Marcos K. Aguilera, Ramnatthan Alagappan, Aishwarya Ganesan

    paper (PDF) · ACM DOI

  • Intro to CXL: Data Sharing on Partly Coherent CXL Memory. Talk at ByteDance, 2026.

    Kiran Hombal and Jiyu Hu. The deck is not public; slide numbers on this page refer to it.

The numbers in this post come from those three sources. Where a chart’s values were read off a paper figure rather than taken from a table, its caption says so. The background facts below were checked against each linked page; the one whose URL no longer resolves says so in its entry.