AI inference server expansion: should you add host memory or VRAM first?
Add VRAM first when the workload stalls with accelerator memory full, and add host DDR4 first when the accelerator has free memory but the host is paging. Two different shortages, not two versions of the same upgrade: HBM sets how much fits on an accelerator and how fast it can be read, while host RDIMM decides whether the machine driving that accelerator can feed it. The failure modes look different in a log and are fixed with different parts.
Start with three questions
Get a reading from the machine while it is slow. Three questions resolve almost every case.
- Where is memory actually full when it stops? If the accelerator exhausts memory and the process aborts, the limit is accelerator capacity. If the accelerator has headroom but the host is paging, it is host capacity.
- Does it stop because it cannot hold more, or cannot move enough? A run that dies at a fixed context length or batch size is hitting a capacity ceiling. A run that completes but produces tokens slowly, with accelerator memory utilisation high and compute low, is hitting a bandwidth ceiling.
- If memory were free and unlimited, would throughput rise? If the accelerator cores are already saturated, more memory changes nothing. Run at the concurrency you need to serve, and check whether compute is the limit.
Bandwidth-bound loads versus capacity-bound loads
A bandwidth-bound load reads more bytes per unit of work than it can afford to wait for. Each generated token requires the weights and the key-value cache to be read from accelerator memory, so the ceiling is bytes moved per second, not bytes stored. Long-context attention behaves the same way — its working set barely changes size but is read repeatedly. Extra capacity does not help.
A capacity-bound load stops because something does not fit: prefilling a very long prompt, holding many concurrent sequences, running a model whose weights plus key-value cache exceed the accelerator, or a host staging a dataset and keeping model files in the page cache.
When both are needed — a load that fits at low concurrency and stops fitting as concurrency rises, on a host that is also short of memory — that is two independent shortages. Size the accelerator side first, because it is the harder of the two to change, then the host.
The bandwidth gap is why that ordering exists. An HBM stack uses a 1024-bit interface, so bandwidth per stack is the data rate times 1024 divided by 8; a DDR4-2933 module moves data over a 64-bit path, which PC4-23400 encodes as roughly 23.5 GB/s.
| Device | Interface | Data rate | Bandwidth per device |
|---|---|---|---|
| DDR4-2933 ECC RDIMM (host) | 64-bit | 2933 MT/s | ≈ 23.5 GB/s |
| HBM2 8GB stack | 1024-bit | 2000–2400 MT/s | 256–307 GB/s |
| HBM2e 8GB / 16GB stack | 1024-bit | 3200–3600 MT/s | 410–461 GB/s |
| HBM3 16GB / 24GB stack | 1024-bit | 5600–6400 MT/s | 717–819 GB/s |
| HBM3e 24GB / 36GB stack | 1024-bit | 6000–9200 MT/s | 768–1178 GB/s |
How to read HBM generation, stack height and density
An HBM part is specified on three axes at once. The generation — HBM2, HBM2e, HBM3, HBM3e — fixes the interface and the bandwidth class. The stack height, written 8-Hi or 12-Hi, is the number of DRAM dies stacked in one device. Density per device follows from those two: within a generation a taller stack carries more dies, and therefore more capacity.
Our stock shows the rule at HBM3e: an 8-Hi device is 24GB, a 12-Hi device of the same generation is 36GB, and both are held at 9200 MT/s from three vendors — so either density comes from the same generation and speed class. At HBM3 the same rule gives 16GB at 8-Hi and 24GB at 12-Hi.
Generation and stack height are independent, which is what catches people out. Moving from HBM2e to HBM3e raises the bandwidth class from 3200–3600 MT/s to 6000–9200 MT/s without changing the die count; moving from 8-Hi to 12-Hi raises capacity without changing the generation.
| Generation | Stack heights in stock | Capacity per device | Data rate | Bandwidth per stack |
|---|---|---|---|---|
| HBM2 | 8-Hi | 8GB | 2000 / 2400 MT/s | 256 / 307 GB/s |
| HBM2e | 4-Hi, 8-Hi | 8GB, 16GB | 3200 / 3600 MT/s | 410 / 461 GB/s |
| HBM3 | 8-Hi, 12-Hi | 16GB, 24GB | 5600 / 6400 MT/s | 717 / 819 GB/s |
| HBM3e | 8-Hi, 12-Hi | 24GB, 36GB | 6000 / 8000 / 9200 MT/s | 768 / 1024 / 1178 GB/s |
One consequence changes what "add VRAM" can mean. The stack count on an accelerator is fixed by the package and interposer design, so a finished card's total accelerator memory is set when it is built. In the parts channel HBM is a repair and refit device: replacing a failed stack restores the designed capacity and bandwidth without raising the ceiling above it.
How much host DDR you add
Size the host side against what the host process actually holds, not against accelerator capacity. Measure the resident set of the serving or training process at the concurrency you need, add the page cache you want for model files and datasets, and leave the rest to the operating system.
Two mechanical points decide whether the capacity you buy is usable. First, populate every memory channel: a DDR4-2933 module moves about 23.5 GB/s over one 64-bit path, so half-populating the channels halves aggregate host bandwidth regardless of total capacity. Second, match rank and device width to the platform's population rules — printed together, as in 2Rx4, where the leading figure is the rank count and the trailing figure the device width. Higher rank count reaches higher density but loads the bus harder than lower-rank parts.
Our Mac Pro 2019 kits are assembled at 6 or 12 modules so the customer receives a matched batch rather than modules from mixed sources: 96GB and 192GB at six, 192GB and 384GB at twelve.
What this looks like against our stock
Stock is held in Shenzhen and can be inspected before shipment. HBM is sold by lot, with quantity per lot depending on the stack height of the devices in it. Registered DDR4 lines carry a published list price, with the 32GB line quoting at a minimum of 12 pieces; HBM and kits are quoted on request.
| SKU | Item | Side | Price |
|---|---|---|---|
| RDIMM-16G-DDR4-2933-1Rx4 | 16GB DDR4-2933 PC4-23400 ECC RDIMM, 1Rx4 | Host capacity | US$140.00 |
| RDIMM-32G-DDR4-2933-2Rx4 | 32GB DDR4-2933 PC4-23400 ECC RDIMM, 2Rx4 | Host capacity | US$310.00 · MOQ 12 pcs |
| KIT-MACPRO19-6X16G | Mac Pro 2019 kit, 6 × 16GB = 96GB | Host capacity | Request quote |
| KIT-MACPRO19-12X32G | Mac Pro 2019 kit, 12 × 32GB = 384GB | Host capacity | Request quote |
| VRAM-HBM2-SAM-01 | HBM2 8GB 8-Hi (Samsung), 2000 MT/s | Accelerator memory | Request quote |
| VRAM-HBM2e-SAM-05 | HBM2e 16GB 8-Hi (Samsung), 3200 MT/s | Accelerator memory | Request quote |
| VRAM-HBM3-HYN-16 | HBM3 24GB 12-Hi (SK hynix), 5600 MT/s | Accelerator memory | Request quote |
| VRAM-HBM3e-HYN-20 | HBM3e 24GB 8-Hi (SK hynix), 8000 MT/s | Accelerator memory | Request quote |
| VRAM-HBM3e-HYN-22 | HBM3e 36GB 12-Hi (SK hynix), 9200 MT/s | Accelerator memory | Request quote |
| VRAM-HBM3e-MIC-25 | HBM3e 36GB 12-Hi (Micron), 9200 MT/s | Accelerator memory | Request quote |
That is a slice of 25 HBM lines across four generations and two stack heights. Part numbers not listed can often be sourced against a confirmed enquiry.
When you should not add either one
- Compute is already the limit. If the accelerator cores are saturated while memory sits partly idle, more bandwidth and more capacity both change nothing.
- The bottleneck is the interconnect. Host RDIMM does not widen the host-to-device link, and accelerator memory does not fix a transfer that never reaches memory.
- The platform does not accept that generation. A board built for HBM2 or HBM2e devices cannot be refitted with HBM3 or HBM3e — interface, signalling and controller differ. The same applies to DDR5 modules in a DDR4 platform.
- All slots are already populated. The answer is a different density per module, which changes rank and device width, not more modules.
- The board is healthy. If nothing has failed, an HBM purchase does not add capacity to a working card.
- The grade is untested and you have no test capability. Pulled-untested stock is priced for buyers who validate themselves and can absorb a failure rate.
Common questions
How do I tell whether my workload is bandwidth-bound or capacity-bound?
Watch what fails and what is idle. If the process aborts because accelerator memory is exhausted, or stops fitting above a fixed context length or batch size, the limit is capacity. If it completes but tokens arrive slowly while accelerator memory utilisation is high and compute is low, the limit is bandwidth. A paging host with idle accelerators points to the host memory.
Can I add HBM to an accelerator that already has it, to raise its capacity?
No. The stack count is fixed by the package and interposer design, so a finished card's total accelerator memory is determined when it is built. Replacing a failed stack restores the capacity and bandwidth the board was designed for. Raising the ceiling means more accelerators, or a card built with taller stacks or a newer generation.
Will adding host DDR4 memory make token generation faster?
Not on its own. Token generation reads weights and key-value cache from accelerator memory, so its speed is governed by accelerator bandwidth. Host memory helps when the host is the constraint — a paging process, model files re-read from disk, a dataset that does not fit — and those show up as stalls, not slow tokens.
Does a taller stack always mean more bandwidth?
No. Stack height sets capacity within a generation; the generation and its data rate set the bandwidth class. At HBM3e an 8-Hi device is 24GB and a 12-Hi device is 36GB, and both in our stock run at 9200 MT/s, so bandwidth per stack is the same.
Which condition grade should a production inference node use?
New, refurbished and pulled-and-tested lines are all functionally verified before packing; the difference is who carries the risk. Pulled-untested stock is sold as-is at a discount and suits buyers who run their own validation. Grade is recorded per line and appears on the quotation, and every shipment carries lot photographs and test reports.
Still not sure which part fits your platform? Ask Ms Aya — or send the part number, quantity and destination port and we will quote the lot in writing.