Can Moving 3D Gaussian Splatting Be Streamed by Sending Only What Changes? Formats and Real Uses of 4DGS Delivery

October 11, 2026

This page contains advertising (affiliate links). See our Privacy Policy for details.

11 October 2026

I usually write about local generative AI here, but this time it is about the moving version of 3D Gaussian Splatting, a technique that builds 3D scenes from photos. Footage is appearing in which you can watch a person singing or dancing from any angle you like. Could such footage be streamed live?

On a stage the background barely moves; only the performers and the lighting do. In that case, sending the first frame in full and then only what changes should make the data small. In this article I looked at that approach of “one reference frame plus only the changes" through research papers, file-format standardization and examples in actual use.

This reflects information as of 28 September 2026.

Sponsored

What are 3DGS and 4DGS? Definitions

3DGS: 3D built from blurry dots

3D Gaussian Splatting (3DGS) is a way to build a 3D scene from photos. Based on photos taken from many angles, it assembles the scene from a collection of small colored blurs (Gaussians). Each splat holds position, orientation, size, transparency and color. The finished scene stays close to photographic and can be viewed from any position.

4DGS: 3DGS stretched along time

3DGS represents a still scene. Adding a time axis so that it can represent a moving scene is called 4D Gaussian Splatting (4DGS): width, height and depth plus time makes four dimensions. To make one, you surround the scene with dozens of cameras and shoot at the same moment.

Frames and differences

A frame is one picture of a video. How many frames are shown per second is FPS (frames per second); at 30 FPS it is 30 frames a second. A difference is data containing only what changed from the previous frame. In this approach, only the first frame (the reference frame) is made in full, and from the second frame on only differences are made.

Codecs

A codec is a set of rules for shrinking video to send it and restoring it on the receiving end. H.264 is the best-known one, and phones and PCs have dedicated circuits that restore it quickly. Ordinary video does not actually send every frame in full either; it shrinks by using differences from the previous frame.

The flow of sending a reference frame plus differences: frame 1 is everything in full (heavy), frame 2 is only the differences of splats that moved, frame 3 is only the differences of splats that moved, and so on, adding splats for objects that newly appear.

Sponsored

What I looked into and how I checked

For sending moving 3DGS as “reference frame + differences", I looked at three things. I did not run any of this on this blog’s machines; this is research based on papers, official announcements and press reports.

  • In the research papers, what form of data is made, and how many MB per frame? How much smaller is it than making every frame in full?
  • For still 3DGS and moving 4DGS, how far have file-format standards been decided? How much smaller does each format make a file?
  • Where is it used in actual works and services?

For the papers I took up four from 2024 (3DGStream, QUEEN, HiCoM, V3). I only list numbers I confirmed in the body or abstract of the papers. The numbers in the papers are under different conditions, such as the scene, the number of cameras and the GPU, so lining them up side by side cannot decide winners and losers.

Sponsored

What is sent as the difference? The data forms in four papers

All four make a reference 3DGS and express following frames as differences from it. What differs is the form in which they hold the frame-to-frame change.

3DGStream (CVPR 2024): hold how splats move in a small network

From the second frame on, 3DGStream does not rebuild the splats; it learns only “how much to move and rotate" the existing ones. This movement is memorized by a small neural network called NTC (Neural Transformation Cache). Splats specific to the frame are added only when something that was not in the previous frame appears.

The data sent is the combination of the 3DGS of the first frame, an NTC per frame, and the added splats.

QUEEN (NeurIPS 2024, NVIDIA): thin out per-splat differences

QUEEN learns directly how much each splat’s values (position, color and so on) changed between frames. For position differences, splats that did not move are set to zero and thinned out. Differences other than position are rounded to few levels (quantized) to make them small.

The paper’s abstract says it uses a signal that tells the still parts of the scene from the moving parts. The idea is to recognize parts that stay still, like the background, and thin out their differences.

HiCoM (NeurIPS 2024): hold it as coherent motion

HiCoM expresses the motion between adjacent frames with a small number of values, using the property that nearby splats move together. The paper title’s “Hierarchical Coherent Motion" refers to this mechanism.

V3 (SIGGRAPH Asia 2024): hold it as ordinary video

V3 reworks moving 3DGS as “2D video". It lays splat values (position, color and so on) out in the cells of an image, one image per kind of value. Making this image for each frame gives a video per kind of value.

Because that video is shrunk with H.264 and sent, the video circuits already in phones and PCs can restore it. The work of shrinking using frame-to-frame differences is left to an existing codec.

Sponsored

How many MB per frame? Sizes and speeds from the papers

I lined up the values written in the papers. Training time is the time to make one frame’s difference. Display speed (FPS) is how fast the finished data can be drawn on screen.

MethodSize per frame [MB]Training time per frame [s]Display speed [FPS]Conditions given in the paper
(Baseline) make a full 3DGS every frame47.1about 498 (8.3 minutes)390Comparison table in the 3DGStream paper. N3DV, RTX 3090
3DGStream7.6about 12215N3DV (21 cameras), RTX 3090
QUEEN0.7under 5about 350Several scenes with large motion. GPU not stated in the abstract
HiCoM(about 85% smaller than the latest method)under 2 on average(not stated in the abstract)When several frames are trained in parallel
V30.51-0.53about 49-53 (0.82-0.89 minutes)435 (RTX 3090) / 96 (iPad, M2) / 27 (iPhone, A15)Displayed at 1920×1080
From the body and abstract of each paper. The first row and 3DGStream are the same paper under the same conditions. The others do not share conditions, so no ranking can be drawn from side-by-side comparison.

There are two pairs that can be compared under the same conditions. In the 3DGStream paper, on N3DV, a multi-view video set commonly used for evaluation, making a full 3DGS every frame costs 47.1MB per frame, against 7.6MB for 3DGStream, about one sixth. In the V3 paper, 3DGStream and V3 are lined up under the same conditions on different data. V3 is 0.51-0.53MB, about one fourteenth to one fifteenth of 3DGStream (7.6MB).

Training time runs in the opposite order. In the V3 paper’s table, 3DGStream takes about 7-11 seconds per frame (0.12-0.19 minutes) and V3 about 49-53 seconds. This differs from 3DGStream’s own paper (about 12 seconds) because the data differs. The smaller the shrinking, the longer it takes to make.

3DGStream and V3 measure display speed on an RTX 3090. 3DGStream reports making one frame’s difference within about ten-odd seconds and drawing more than 200 frames a second on the RTX 3090, a top-end consumer GPU. V3 is written to have drawn 27 frames a second even on an iPhone (A15 Bionic, in the iPhone 13 and others), reaching the point where it can be viewed on a phone.

How much would it be on a network line?

Multiplying the size per frame by the number of frames per second gives a rough amount to put on the line. I calculated it for 30 FPS (treating 1MB, megabyte, as 8Mb, megabits). As a yardstick I also list the values YouTube recommends for ordinary live video.

MethodPer frame [MB]At 30 FPS [Mbps]
Full 3DGS every frame47.1about 11,304
3DGStream7.6about 1,824
QUEEN0.7about 168
V30.51-0.53about 122-127
(Reference) DNE x Gracia “Open" stream–17-75 (published value)
(Reference) YouTube live 1080p, 30fps–10 (recommended for H.265 and AV1; for H.264 it is 14)
(Reference) YouTube live 4K, 30fps–30 (recommended for H.265 and AV1; for H.264 it is 42)
Rows calculated from the papers are my own calculation. Mbps is megabits sent per second.

Even V3, the smallest of the papers, is about four times YouTube’s recommended value for 4K live. The DNE and Gracia stream introduced later is published at 17-75 Mbps, smaller than sending the papers’ values as they are, and about 0.6 to 2.5 times the 4K live recommendation. How the product shrinks it could not be determined from public information.

Sponsored

Are the file formats settled? Standardization for still and moving

File formats for still 3DGS have firmed up over about the past year. On the other hand, in what I checked I found no public standard that decides how to store moving 4DGS. The four papers each hold data in their own form, which cannot be read by one another.

FormatMade byContentsSize compared with .plyMoving scenes
.ply–Splat values laid out as they are. 59 numbers per splatBaseline (about 240 bytes per splat)Some examples use it as a numbered sequence, one file per frame
.splatantimatter15 (individual developer)Drops the view-dependent color information and packs each splat into 32 bytesAbout 1/7 to 1/8Not stated
.spzNianticRounds values and shrinks with the ZSTD compression method. SPZ 4 released 5 May 2026About 1/10Not stated in the official description
.sogPlayCanvasStores values laid out in WebP imagesAbout 1/15 to 1/20Not stated in the official description
KHR_gaussian_splattingKhronos GroupAn extension to glTF, the common 3D data format. Release candidate announced 3 February 2026, now ratified–Not stated in the specification
Gaussian Splat CodingMPEGStandardization of a compression method is under study–For now still scenes. Moving ones are a long-term study item
From each body’s and developer’s official pages, as of 28 September 2026. The size of one .ply splat and the .splat ratio are my own calculation (one number = 4 bytes).

In numbers, a scene of one million splats is about 240MB as .ply, about 24MB as .spz and about 12-16MB as .sog (my calculation). As an actual example, SuperSplat’s developer writes that a scene of 4.4 million splats was 990MB as .ply. That is about 225 bytes per splat, roughly in line with the calculation.

Applying this to 4DGS shows the limit of shrinking with still-image formats alone. Even cutting the 3DGStream paper’s 47.1MB to one twentieth with .sog leaves about 2.4MB per frame, still larger than the difference methods QUEEN (0.7MB) and V3 (0.51-0.53MB) (my calculation).

V3 re-laying values into images and video, and SOG storing them in images, are very similar ideas. Put the splat values in the cells of an image and you can use image and video compression polished over many years as it is.

What MPEG is doing

MPEG is the international group that has set video standards such as H.264. Its official page says that for 3DGS compression it will “focus on interoperability for static scenes for now and study dynamic scenes over a longer horizon."

Meanwhile, Radiance Fields, a news site covering 3DGS (4 August 2026), reports that a call for test material for moving 3DGS has been decided, at the 155th meeting in Geneva. The deadline is 15 October 2026, and the call for proposals on the compression technology itself is reported to be in preparation. A standard format would come after that.

Sponsored

Where is moving 3DGS used? Works and services

I could confirm a few cases where moving 3DGS was used. But all of them are “record, then stream" or “demonstration of capture"; I found no example streamed live.

ExampleWhat it isLive?Source
A$AP Rocky “Helicopter" music videoEvercoast captured people with 56 RGB-D cameras. About 30 minutes of footage exported as a numbered .ply sequence and composited with CGNo (used in production)Press (Radiance Fields)
DNE x Gracia “Open"A 4-minute performance by singer Amy May. Plays without an app in a browser, on a phone or on Meta Quest 3. 17-75 MbpsNo (pre-recorded)Press (CG Channel, 27 April 2026)
4DV.ai x OBSBOTAt NAB Show 2026, showed a capture rig of about 60 OBSBOT Tail 2 cameras. Visitors viewed it in VR headsetsNo (demo of capture and reconstruction)OBSBOT official / press (CineD)
SplatLabsSays it can stream live events and sports at 60fps with under 50 ms latencyClaims only. No customers listedOwn site
What I could confirm as of 28 September 2026

A$AP Rocky “Helicopter"

It is reported that nearly all the people in the music video were captured as volumetric 3D with Evercoast’s system, using 56 RGB-D cameras (cameras that record color and depth together), synchronized by two Dell workstations. The export was a numbered .ply sequence, a full file for every frame. It is an example of using 4DGS as footage for video production, and likely a different thing from streaming with differences.

DNE x Gracia “Open"

A 4-minute musical performance released by the capture studio DNE and Gracia. CG Channel introduced it as “the first streamable 4DGS musical performance that plays in a browser". Playback uses WebGPU (a way for the browser to use the GPU), on a PC, phone or Meta Quest 3.

“Streamed in real time" here means playing back a pre-recorded performance as it is delivered. The performance itself was filmed in advance. Gracia also distributes a free 3DGS viewer, which I covered in a related article.

4DV.ai x OBSBOT

At the 2026 broadcast equipment show NAB Show, 4DV.ai and OBSBOT exhibited a portable capture setup. OBSBOT’s announcement says it uses about 60 cameras and is designed to scale to over 200 in production. Stated uses include stadiums, concert stages and film sets.

In CineD’s article, 4DV.ai’s Jiaming Sun describes the weakness of making 4DGS one frame at a time: it “does not share information across time and ignores the fact that most of the scene barely changes from one moment to the next." That points in the same direction as this article’s idea of sending only what changed.

SplatLabs

A service that says it can stream live events, sports and conferences in 3D in the browser. Its site says 60fps, under 50 ms latency and 3 ms rendering per frame. At the same time it says “looking for first partners" and “early Alpha", and lists no customers. The way the numbers were measured is not published, and I found no third-party verification.

Sponsored

Is the difference approach an advantage for live events with a still background?

From here on is this blog’s own reasoning based on what I looked into.

If the stage background is built in full in frame 1, it hardly needs to be sent from frame 2 on. The four papers also take the form of “make the first frame once, heavily, then make only differences", so they should suit scenes where the background does not move. QUEEN telling still parts from moving parts to thin out differences rests on the same idea.

However, at a live event the shape of the background does not move, but its color changes with the lighting. If the color changes, the background splats should also count as “changed". The more the lighting moves in the staging, the less the differences are expected to shrink.

If this is right, on the same stage a scene with fixed lighting should have a smaller size per frame. If the size does not change even with fixed lighting, then what mainly decides the size of the differences is the performers’ movement. None of the papers I read has results from that comparison.

Another wall is the making speed. To stream 30 frames a second live, one frame’s difference must be made in one thirtieth of a second (about 0.033 seconds). Even HiCoM, with the shortest training time among these papers, averages under 2 seconds. That is a gap of about 60 times at 2 seconds, about 150 times at QUEEN’s 5 seconds and about 360 times at 3DGStream’s 12 seconds (my calculation). This gap is probably one reason the examples in use today cluster around “record, then stream".

Sponsored

October 2026 addendum: streaming examples that came out in September

After I wrote this article, there were developments in streaming moving 3DGS. NEXIA Entertainment began commercial 4DGS streaming for live and sports footage, and ByteDance’s Multimedia Lab published a 4D media system for 6DoF video that handles both live and catch-up viewing. At IBC 2026, Nokia showed a demo of moving splats encoded with MPEG’s V3C and GS4, delivered over MPEG DASH and rendered on Samsung and XREAL devices. On the research side, MoQSplat (arXiv 2609.18624), which delivers 3DGS progressively over Media over QUIC, also appeared. I have tried none of them; this is what I read in Radiance Fields’ monthly roundup and each company’s announcements.

Source: Radiance Fields, “Gaussian Splatting in September 2026" and each company’s announcements (September 2026).

Sponsored

What I Haven’t Confirmed in This Article

  • I have not reproduced the papers’ numbers. Scenes, camera counts and GPUs differ between papers, and lining them up cannot show which is better
  • QUEEN’s GPU and dataset name and HiCoM’s comparison target are not stated in the abstracts I read, so I could not confirm them
  • The Mbps at 30 FPS and the size per million splats for each format are my simple calculations. Real streaming would involve further compression and thinning
  • How the DNE and Gracia stream is formatted, and how it holds its differences, was not public
  • I found no example of live-streaming with the 4DGS difference approach in commercial operation. That means I did not find one in what I checked
  • That lighting changes make the differences larger is this blog’s guess, with no measurement or paper to back it

Summary: how far has “send only what changed" come?

The approach of sending moving 3DGS as “one reference frame plus only what changed" took shape in 2024 papers. How differences are held varies by paper and falls into four types.

  • Hold how splats move in a small network (3DGStream)
  • Hold per-splat differences, thinned out (QUEEN)
  • Hold it as coherent motion (HiCoM)
  • Hold it as ordinary video (V3)

Per frame, about 0.5-7.6MB is reported. In the example compared under the same conditions, 3DGStream was about one sixth of making a full file every frame (47.1MB), and V3 was about one fourteenth to one fifteenth of that 3DGStream. At 30 FPS that is about 122-1,824 Mbps, still more than four times ordinary 4K live video (30 Mbps).

File formats are firming up for still 3DGS, while there is not yet a standard for moving 4DGS. Uses I found go as far as music-video production and streaming of recorded performances; I could not confirm real operation of live broadcast.

If you are someone whoWhat you can do next
Wants to see moving 3DGS firstPlay DNE and Gracia’s “Open" in a browser or on a Meta Quest 3
Wants to try the making side on your own GPUOf the papers, 3DGStream and V3 measure on an RTX 3090. Start by checking, in the two papers, the numbers closest to your GPU’s conditions
Wants to judge if it can be used for live broadcast workThere is no standard and no real operating example yet. It is safest to wait for what follows MPEG’s call (deadline 15 October 2026)
Is fine with still 3DGSNo need to chase the moving side. .spz or .sog is enough to make it small

In the papers, making the differences for moving 3DGS fits within ten-odd seconds per frame even on a consumer RTX 3090. In the world of video too, what can be done on your own PC’s GPU may be the question from here on.

Sites I consulted

  • 3DGStream (arXiv 2403.01444; HTML version with the comparison table; project page)
  • QUEEN (NeurIPS 2024; NVIDIA project page)
  • HiCoM (arXiv 2411.07541)
  • V3 (arXiv 2409.13648; HTML version)
  • PLY, SPZ and SOG formats (PlayCanvas); .splat format (antimatter15/splat); SuperSplat 3.0 release notes (the 990MB example)
  • SPZ 4 announcement (Niantic Spatial)
  • KHR_gaussian_splatting announcement and specification (Khronos Group)
  • Gaussian Splat Coding (MPEG); MPEG’s call for material (Radiance Fields)
  • A$AP Rocky “Helicopter" (Radiance Fields); DNE x Gracia “Open" (CG Channel); 4DV.ai x OBSBOT (OBSBOT, CineD); SplatLabs
  • YouTube live encoder settings (recommended bitrates): support.google.com/youtube/answer/2853702
Sponsored