DGX Spark Gets a 64GB Version and the 128GB Rises to $6,950: What Fits in Half the Memory, Checked Against EVO-X2 Measurements

本ページは広告(アフィリエイトプログラム)を含みます。詳しくはプライバシーポリシーをご覧ください。

9 October 2026

NVIDIA’s small AI computer, the DGX Spark, is getting a 64GB version with half the memory. It was announced on 2 October 2026 at $4,999, going on sale on 23 October. On the same day the direct price of the 128GB version rose to $6,950. In August I compared three 128GB-class machines here: the EVO-X2 mini PC, the Mac Studio and the DGX Spark. The Japanese price of the DGX Spark at that time was 730,000-760,000 yen. What can you run with half the memory? I check it against the announcement and my own measurements.

This reflects information as of 3 October 2026. I do not own a DGX Spark; this article is based on published information. For “will the model fit" and the speed reference numbers, I used values I measured on the EVO-X2.

What was announced: a 64GB version and a price rise for the 128GB version

ItemDGX Spark 64GB (new)DGX Spark 128GB (existing)
Announced2 October 2026Launched October 2025
Price (US)From $4,999$6,950 (revised 2 October; $3,999 at launch, then $4,699 from February 2026)
On sale23 OctoberOn sale now
Sold byAcer, ASUS, Dell, GIGABYTE, HP, MSI (no NVIDIA-direct Founders Edition)NVIDIA direct plus partners
ChipGB10 Grace Blackwell (same)
Memory64GB shared128GB shared
Memory bandwidth273GB/s (nominal, same)
NetworkingConnectX-7 (two units can be linked directly)
Model size that runs on one unit (NVIDIA)Up to 100 billion parametersUp to 200 billion parameters

Sources are AI Watch (2 October), reports from overseas outlets, and quotes from NVIDIA’s announcement. Reports cite tight memory supply and rising memory prices as the reason for the price increase, but I could not find a sentence in which NVIDIA itself names a single reason.

How prices have moved: 1.7x in a year

WhenUS price of the 128GB versionStreet price in Japan (my check)
October 2025 (launch)$3,999–
February 2026$4,699–
August 2026$4,699About 730,000-760,000 yen (this blog’s August comparison)
2 October 2026$6,950About 1,045,000-1,200,000 yen (listings on Kakaku.com, ELSA and TSUKUMO)

In Japan the street price rose by roughly 280,000-470,000 yen in the two months from August to October. No Japanese price for the 64GB version had been announced by the vendors as of this article. Converting the US price of $4,999 straightforwardly gives around 750,000 yen. Since the Japanese price of the 128GB version has tracked about 1.0-1.2 times the US price, the 64GB version could land in the 800,000-900,000 yen range. That is a guess, not a confirmed figure.

What fits in 64GB: how I checked, against my own measurements

Whether a model file fits in shared memory is the same question for the DGX Spark and the EVO-X2. I take the models I measured on the EVO-X2 (128GB) and test them against a 64GB line by file size. Speeds are values on the EVO-X2’s integrated GPU (llama.cpp b11192, October 2026), not DGX Spark speeds. Because memory bandwidth is close (256GB/s versus 273GB/s), I think they are a usable guide for writing speed, but that is an inference from bandwidth alone. GPU design and software optimization could make a difference.

A few short definitions. tok/s is tokens written per second; a token is roughly a word fragment. GiB is a unit of file size (about 1.07 billion bytes). MoE (mixture of experts) holds many specialists and uses only some per question, so it runs fast for its file size. Shared memory means the CPU and GPU use the same memory, and this capacity takes the place of a graphics card’s VRAM as the place the model sits.

Model (quantization per row)File [GiB]64GB line (judged by file size only)Writing speed on EVO-X2 [tok/s]
Qwen3.8 27B (4-bit)15.3Inside the line13.0
Gemma 4 31B (4-bit)17.1Inside the line11.6
Qwen3 Coder 30B (4-bit, MoE)17.3Inside the line91.1
Nemotron 3 Nano 30B (4-bit, MoE)22.6Inside the line69.0
Llama 3.3 70B (4-bit)39.6Inside the line5.3
gpt-oss 120B (MXFP4, about 4-bit)59.0Inside the line, but little headroom55.2
GLM-4.5-Air (4-bit)67.8Outside the line25.6 (b10605)
Qwen3 235B (2-bit)79.8Outside the line21.8 (b10605)
MiniMax M2.7 (4-bit)101.0Outside the line28.9

The “line" in the table is drawn by file size alone, and I have not confirmed on a DGX Spark whether the models actually run. Memory is also needed for the OS and for the conversation’s memory (the KV cache). The 64GB line sits just above gpt-oss 120B (59GB), and a 120B-class model will run, if at all, with little room to spare. That matches NVIDIA saying “up to 100 billion parameters". Up to 27-70B there is room to spare, and the feel on my machine is 70-90 tok/s for MoE models and 5-13 tok/s for dense models. In plain terms, MoE models write faster than most people read, and dense models are at or somewhat below reading speed.

You need 128GB if you want to run models above 60GB. GLM-4.5-Air (68GB), the 2-bit Qwen3 235B (80GB) and MiniMax M2.7 (101GB) are in this band. For this group, the 64GB version needs two units linked together.

Linking two units gets you 128GB, but 64GB x 2 costs more than one 128GB

NVIDIA is pushing the use of two 64GB units linked directly by a ConnectX-7 QSFP cable. Linked, they total 128GB, run models up to 200 billion parameters, and are described as up to 1.7 times the performance of one unit. A Sync Cluster Assistant that sets this up automatically is also coming. Sync Model Launcher, which distributes a model to one or two units and starts it, is due at the end of October, with Qwen3.8 27B given as an example.

But adding up the prices, two 64GB units cost $9,998, about $3,000 more than one 128GB unit at $6,950. It looks right to read this not as a cheap way to reach 128GB but as a way to start with one unit and add another later. The 1.7x figure is also for a single model spread across both units, and I have not verified it.

How does it differ from the EVO-X2 (128GB mini PC)?

The EVO-X2, the machine I measure on, is also a 128GB shared-memory machine. Here are the differences.

ItemDGX Spark 64GBDGX Spark 128GBGMKtec EVO-X2 128GB
Memory64GB128GB128GB
Memory bandwidth (nominal)273GB/s273GB/s256GB/s
GPU toolchainCUDA (NVIDIA), ARM CPUROCm / Vulkan (AMD), x86 CPU
Price (as of 3 October)$4,999 (Japan: TBD)$6,950 / about 1,045,000-1,200,000 yen in Japan583,000 yen on the official store (128GB + 2TB)
My measurementsNoneNoneYes (over fifty models)

Since the bandwidths are close, I expect the writing speed for the same model at the same 4-bit level not to differ greatly, but this is an inference from bandwidth alone and I have not measured a DGX Spark. The difference is the toolchain. If you use software that needs CUDA (NVIDIA’s GPU computing platform, which a lot of AI software assumes) for things like training, fine-tuning and some image generation, the DGX Spark is the choice; if you only run inference with llama.cpp or Ollama, the EVO-X2 stands on the same ground. The CPUs differ too: the DGX Spark is ARM and the EVO-X2 is x86, like a regular PC. Some Windows software does not run on ARM, so check that before buying. Compared at 128GB each, the DGX Spark costs about 1.8-2.1 times the EVO-X2 (Japanese prices).

What I Haven’t Confirmed in This Article

  • Speed on a real DGX Spark (neither 64GB nor 128GB is measured). The speeds in the table are EVO-X2 values
  • The Japanese price and release date of the 64GB version (no vendor announcements as of 3 October)
  • Performance when two units are linked (NVIDIA’s explanation only)
  • The reason for the price rise (reports cite memory prices, but I found no official explanation from NVIDIA)
  • What Sync Model Launcher does (due at the end of October, not yet public)

Summary: who is fine with 64GB, who needs 128GB

The 64GB version is the entry point for people who want to run 4-bit 27-70B models and up to gpt-oss 120B on one unit. If you want models above 60GB (GLM-4.5-Air, 235B, MiniMax), you need the 128GB version or two 64GB units. The price of the 128GB version rose 1.7 times in a year and passed 1,000,000 yen in Japan. The 128GB EVO-X2 is about 580,000 yen on the official store. Whether you need CUDA is the dividing line.

As a next step, I suggest first checking the file size of the model you want to run and seeing which side of the 64GB line it falls on. Model sizes and speeds are in the by-use summary and the 128GB machine comparison.

What runs by memory size (VRAM)

Hardware covered in this article

GMKtec EVO-X2 (Ryzen AI Max+ 395 / 128GB / 2TB)A8限定クーポン「A82608」で5,000円OFF(2026/10/31まで)/公式オンラインストアのクーポン「KANSYA2026」で2,000円OFF(2026/10/18まで)

¥583,000 Amazon・2026-10-04調べ

公式サイトのみ ¥5,000引きクーポン配布中 クーポンコード A82608(2026-10-31まで)