A Mac You Can Put 512GB Into — What Changes, Seen From a 16GB Wall

This page contains advertising (affiliate links). See our Privacy Policy for details.

3 September 2026

I had been stuck at a 16GB wall, adding an external GPU to a mini PC. Models that fit in 16GB run fast; cross that line and they stop running at all. Measuring where that line sits has been my work for months.

Then, on 25 August 2026, Apple announced a machine that takes 512GB of memory: the Mac Studio.

From where I sit, worrying about 16GB, 512GB looks like another world. But the arithmetic should be the same, only with more zeros. What actually happens over there?

This article is compiled from published information. I own neither machine discussed here (checked 3 September 2026).

Sponsored

What I looked at, and how I checked it

How local LLMs behave on a machine with a very large pool of memory, and what someone stuck at 16GB can learn from it. The comparisons are the other machines built for the same job: NVIDIA’s DGX Spark, and the discrete RTX 5090 and 4090.

I cross-checked Apple’s announcement and specification pages, public information from NVIDIA and around it, records of people running very large models on the previous generation (M3 Ultra, 512GB), and Japanese retail prices. Not one number here was measured by me — with one exception, the EVO-X2 figures near the end, which are mine.

Sponsored

The numbers, side by side

MachineMemoryBandwidthPrice
Mac mini (M6)32GB170 GB/s$899
Mac mini (M5 Pro)64GB307 GB/s$1,699
Mac Studio (M5 Max, 32-core GPU)36GB460 GB/s$2,499
Mac Studio (M5 Max, 40-core GPU)128GB614 GB/s
Mac Studio (M5 Ultra)96GB–512GB1.2 TB/s$5,499 and up
DGX Spark128GB273 GB/s$4,699
GMKtec EVO-X2 (Ryzen AI MAX+ 395)128GB256 GB/sabout ¥520,000
RTX 509032GB1,792 GB/slist $1,999
RTX 409024GB1,008 GB/s
Reference: M3 Ultra (previous gen)512GB819 GB/s

Bandwidth roughly doubles at each grade: 170, 307, 460/614, 1,200. It is a separate axis from capacity, and confusing the two leads to the wrong purchase.

Sponsored

The same name, two different bandwidths

There are two M5 Max chips. The 32-core GPU version gives 460 GB/s; the 40-core GPU version gives 614 GB/s. Same “M5 Max Mac Studio" on the box, and a 25% difference in how fast it writes.

The base configuration is the 32-core part with 36GB fixed. Asking for 128GB moves you automatically to the 40-core side, so “buy the cheap M5 Max and load it with 128GB" is not a thing you can do.

The M5 Ultra behaves differently: both configurations give the same 1.2 TB/s. It is two dies joined together, so the number of memory channels does not change. On the M5 Max, the lower version has channels cut.

So if you choose an Ultra, writing speed is the same either way. What the upper configuration adds is CPU and GPU cores — reading speed, and work that is compute-heavy.

Sponsored

On bandwidth alone, it does not reach a GPU

MachineBandwidth
RTX 50901,792 GB/s
M5 Ultra1,200 GB/s
RTX 40901,008 GB/s
M5 Max614 GB/s
DGX Spark273 GB/s

The M5 Ultra lands between the 4090 and the 5090. It loses to the 5090, but not by an order of magnitude.

That is not where the decision is made

The 5090 has 1,792 GB/s across 32GB. The M5 Ultra has 1,200 GB/s across 512GB — sixteen times the capacity.

For a model that fits in 32GB, the 5090 wins outright: half again the bandwidth and more compute. No contest.

One byte past 32GB, the 5090 stops. What spills goes out to system memory over PCIe Gen5 x16 at 64 GB/s — one twenty-eighth of its own bandwidth. The speed collapses. The Mac has no such cliff: the same bandwidth continues all the way to 512GB.

ModelRTX 5090 (32GB)M5 Ultra (512GB)
8B at 4-bit (about 5GB)wins easilyslower
32B at 4-bit (about 20GB)winsslower
Qwen3-235B at 4-bit (about 130GB plus context)does not runruns
DeepSeek R1 671B at 4-bit (404GB)does not runruns

The question is not fast or slow. It is runs or does not run.

Sponsored

What people actually got running

On the previous generation — M3 Ultra with 512GB at 819 GB/s — there are records of very large models running.

ModelQuantisationMemory usedSpeed
DeepSeek R1 671B4-bit404GB (VRAM allocation raised to 448GB by hand)17–18 tok/s
DeepSeek V34-bit20 tok/s
Qwen3-235B4-bit MLX272GB24 tok/s

Power draw stayed under 200W — a different order from a multi-GPU rig.

Note the Qwen3-235B line: 272GB in use. The weights are 130–240GB, but adding the area that holds the context pushes it past 256GB. That single fact decides the configuration question later on.

Reading speed and writing speed are separate subjects

DGX SparkMac Studio
Compute1 PFLOP (FP4)no dedicated FP4 path
Bandwidth273 GB/s1.2 TB/s
Reading (prefill)3–4x fasterslower
Writing (decode)slowerabout 3x faster

Reading is decided by compute, writing by bandwidth. So the machine with compute wins on long prompts, and the machine with bandwidth wins on generation.

Sponsored

The same shape showed up on my own bench

I measure a mini PC with external GPUs attached. Swapping a 16GB card for a 12GB one, reading speed fell on every model and writing speed rose on every model — the card changed, so the ratio of compute to bandwidth changed.

The same ruler works on a machine costing a few hundred thousand yen and on one costing a few million.

256GB or 512GB?

ConfigurationCPUGPUMemoryBandwidth
M5 Ultra base30-core64-core96GB1.2 TB/s
M5 Ultra upper36-core80-core256GB / 512GB1.2 TB/s

Between 256GB and 512GB there is no difference except capacity. Same bandwidth, same CPU, same GPU. There is a step before that, though: 256GB and above requires the upper configuration, which from 96GB raises the CPU by 1.2x and the GPU by 1.25x.

The 512GB price has not been announced (shipping late October). Two things are known: the M5 Ultra base with 96GB and a 1TB SSD is $5,499, and going from 96GB to 256GB adds $4,000. On the previous generation, 512GB cost 1.8 times what 256GB cost. At the same ratio the total for a 512GB machine lands somewhere around ¥2.4–3.0 million, against about ¥1.88 million for 256GB. That estimate is mine, not Apple’s.

There is a precedent worth remembering: the M3 Ultra’s 512GB option was withdrawn during the global DRAM shortage, and the 256GB upgrade price rose at the same time. That the M5 Ultra’s 256GB upgrade starts high suggests the shortage is already priced in.

Which leaves one question: do you intend to run a 404GB model? If not, there is no reason to wait — the extra money buys capacity and nothing else. And last time, waiting meant the option disappeared.

Sponsored

Who this machine suits

It suits someone running a large model alone, conversationally, on data they do not want to send anywhere, who cares about silence and power draw.

It does not suit someone who wants a model that fits in 32GB to run fast (a GPU wins outright), someone whose work is mostly long prompts (DGX Spark reads 3–4x faster), someone training or fine-tuning, someone generating images or video, or someone serving several users at once.

It is not a machine you buy for speed. 17–18 tok/s on a 671B model does not reach the 35–50 tok/s of a cloud API. It comes down to whether a large model running on your own desk is worth something to you.

Sponsored

Next to the 128GB machine I actually own

MachineMemoryBandwidthPrice (Japan)
GMKtec EVO-X2128GB256 GB/sabout ¥520,000
DGX Spark128GB273 GB/s¥980,000–1,260,000
Mac Studio (M5 Max, 40-core GPU)128GB614 GB/sprice not confirmed

The EVO-X2 and the DGX Spark sit in the same bandwidth band — 256 and 273 GB/s — and in the same category of 128GB unified-memory machines.

These are my own measurements on the EVO-X2’s integrated GPU, with llama.cpp over Vulkan on 25 August 2026.

ModelReading pp512Writing tg128
llama3.2 1B7,238 tok/s155.62 tok/s
phi4-mini2,408 tok/s79.01 tok/s
qwen3 8B1,214 tok/s40.95 tok/s
qwen3 14B744 tok/s24.28 tok/s
qwen3-coder 30B (MoE)1,219 tok/s91.70 tok/s
llama3.3 70B (dense)109 tok/s5.28 tok/s

Look at the 30B and the 70B. The 30B is the smaller model, and it writes at 91.70 against the 70B’s 5.28 — a factor of seventeen. The 30B is a mixture-of-experts model that reads only part of itself per token; the 70B is dense and reads all of it every time. Same bandwidth, completely different outcome, because of how the model is built.

Set that beside the records above: 256 GB/s here against 819 GB/s there is 3.2x the bandwidth, and the Mac runs a model ten times larger more than three times faster. Mixture-of-experts is doing that work. Large capacity pays off when the model is MoE. Make a dense model bigger and the capacity does not bring speed with it — 5.28 tok/s on a 70B says so.

Sponsored

What the DGX Spark premium actually buys

Same capacity, near-identical bandwidth, and a gap of ¥460,000–740,000 — enough to buy a second EVO-X2. The difference is not bandwidth. It is two things the EVO-X2 does not have: compute (1 PFLOP at FP4, which shows up in reading speed) and CUDA.

In practice the second one matters more. Inference alone runs on AMD through llama.cpp over Vulkan — that is how I measure. But training and fine-tuning are effectively CUDA-only; serving engines like vLLM and SGLang are CUDA-centred; some quantisation formats assume CUDA; new models arrive on the CUDA side first; and image and video generation is far faster there. The premium is an entry fee for CUDA.

Sponsored

Choosing inside the 128GB band

What you want to doWhat to pickWhy
Run larger models and use themAMD AI machines (EVO-X2 and similar)128GB for about ¥520,000 — the cheapest in this band
You need CUDA (training, vLLM, image generation)DGX Sparkone of the few ways to get 128GB with CUDA
You need more capacity stillMac Studiogoes to 512GB — but there is no CUDA
Sponsored

What I did not verify

  • I own neither the Mac Studio nor the DGX Spark. Everything except the EVO-X2 table is from published sources
  • The 512GB price is unannounced; the figures here are my estimate
  • No measurement of an actual M5 Ultra has surfaced. The speeds quoted are from the previous generation, the M3 Ultra
  • DGX Spark pricing is moving: ¥700,000–800,000 at the end of March, ¥984,980 at one retailer on 21 August, ¥1,259,625 on Amazon when I checked. NVIDIA raised the list price from $3,999 to $4,699 in February, citing memory supply
  • Japanese pricing for the Mac mini (M5 Pro) and the current RTX 4090 price were not confirmed
  • MoE models should run lighter than their total size suggests, but no M5 Ultra measurement was found to confirm it
Sponsored

In summary: the ruler is the same at every price

What surprised me about looking into the 512GB world is that the ruler did not change. Reading speed is decided by compute, writing speed by bandwidth, and everything stops the moment you exceed capacity — the same shapes I keep hitting at 16GB on a mini PC.

My own EVO-X2, at 128GB, stands in the same category as the DGX Spark: 256 against 273 GB/s, ¥520,000 against ¥980,000–1,260,000. I thought I was reading about machines in another price class, and the ground turned out to be continuous.

One more thing came out of it. Moving the wall back is not enough on its own. On my machine a dense 70B writes at 5.28 tok/s and a 30B MoE writes at 91.70 — seventeen times apart on the same hardware. Buying capacity is also buying the assumption that you will run mixture-of-experts models. That part does not appear on the price list.

Sponsored