Slot3R: Set-Associative Spatial Memory
for Streaming 3D Reconstruction

1AGI Lab Westlake University2University of Illinois Urbana-Champaign3Nanyang Technological University

* Equal contribution  ·  † Corresponding author

Video overview

Demo

Explore the reconstructions in 3D. Rotate, zoom, or replay the reconstruction from its RGB observations.

Open interactive demo

Slot3R is a training-free framework for streaming 3D reconstruction. Set-associative spatial memory preserves multiple feature states at a shared address, while sparse readout bounds the memory accessed by the decoder at each frame.

Abstract

Streaming 3D reconstruction must preserve evidence from each frame while processing an expanding scene online. Spatial memory is a natural fit because it organizes history by reconstructed 3D location. Yet Point3R uses spatial proximity both to associate a new observation with an existing memory entry and to decide whether to fuse it, conflating co-location with state identity. Because pointers summarize image patches, nearby pointers may encode distinct surfaces, viewpoints, or visibility conditions; averaging them can destroy complementary evidence before later frames disambiguate it. We argue that location should determine address, not whether observations must merge. Slot3R realizes this principle as a training-free, set-associative retrofit that lets multiple states coexist at a shared address while keeping the pretrained Point3R backbone frozen. A bounded sparse readout further decouples persistent storage from per-frame decoder access. At 300–500 sampled frames, Slot3R reduces Point3R's point-cloud accuracy error (Acc) by 57.1%–63.1% on 7Scenes and 64.0%–72.0% on NeuralRGBD, lowers Sim(3)-aligned absolute trajectory error (ATE) on all three pose benchmarks, and remains competitive on video-depth estimation. It completes all evaluated settings from 600 to 1000 sampled frames at about 19 FPS under the same protocol, whereas Point3R and InfiniteVGGT run out of memory at 800 frames and beyond.

Qualitative Comparison

We compare point-cloud reconstructions on four 500-frame NeuralRGBD sequences. Select a scene below to compare Point3R, Slot3R, and the ground truth. Click any image to enlarge it.

Point3R Baseline
Slot3R Ours
Ground truth Reference

Static point-cloud views from the paper. Select a scene to compare; click an image to enlarge.

Method

Given streaming RGB images, Slot3R maintains a set-associative spatial memory while keeping the pretrained Point3R backbone frozen. Each 3D address indexes a bucket containing multiple feature states. Confidence and feature similarity determine whether incoming observations are rejected, inserted, fused, or replaced. For each new frame, sparse spatial readout combines local memory with global anchors under a fixed budget of 640 tokens.

Overall framework of Slot3R. Spatial addressing is separated from state fusion, and persistent memory storage is decoupled from per-frame decoder access.

Results

3D Reconstruction

Across the 300–500-frame evaluations, Slot3R reduces Point3R’s accuracy error (Acc) by 57.1%–63.1% on 7Scenes and 64.0%–72.0% on NeuralRGBD, while improving normal consistency. The complete point-cloud table from the paper is reproduced below.

Point-cloud reconstruction on 7Scenes and NeuralRGBD at 300–500 sampled frames (k = 2).
MethodAcc ↓Comp ↓NC ↑Avg. FPS ↑
300400500300400500300400500
7Scenes
CUT3R0.13620.16970.19260.07440.10780.08940.55130.54470.540913.89
TTT3R†0.04020.05020.06510.02450.02630.02950.56470.55750.552213.68
StreamVGGTOOMOOMOOMOOMOOMOOMOOMOOMOOMOOM
STream3R0.09620.1013OOM0.03170.0426OOM0.57690.5744OOM5.82
InfiniteVGGT†0.04320.04210.03980.02690.02690.02360.56950.56490.56173.63
RetrieveVGGT0.02660.02760.02840.02750.02920.02940.50650.50630.50840.97
ZipMap-Stream0.10960.10570.10210.26440.30010.30880.52230.51800.51505.12
Spann3R†0.07650.07970.07070.02780.02750.01870.54640.54280.540241.80
Point3R0.06340.06380.07450.03470.03080.02650.55960.55410.551016.77
Ours (Core)0.02620.02740.02750.01540.01690.01580.57900.57890.579019.17
Ours-VPC-M0.02570.02660.02820.01580.01640.01630.57890.57860.579217.54
Ours-VPC-A0.02640.02690.02700.01590.01670.01630.57930.57860.579115.73
Δ Ours (Core)
vs. Point3R
↓ 58.7%↓ 57.1%↓ 63.1%↓ 55.6%↓ 45.1%↓ 40.4%↑ 3.5%↑ 4.5%↑ 5.1%↑ 14.3%
NeuralRGBD
CUT3R0.22600.31700.34710.08830.12450.15390.58460.56560.553917.51
TTT3R0.10580.14250.16660.03230.07750.10270.64150.62690.622117.40
StreamVGGTOOMOOMOOMOOMOOMOOMOOMOOMOOMOOM
STream3R0.06640.0828OOM0.01490.0211OOM0.68800.6935OOM5.81
InfiniteVGGT†0.05320.06510.06960.02440.03480.03700.64310.65030.64333.63
RetrieveVGGT0.04380.04560.03700.02480.03160.03970.61660.61580.51321.07
ZipMap-Stream0.17450.19850.21410.29620.44520.53050.52880.52880.52915.01
Spann3R†0.10020.14170.18290.02620.04470.04870.59530.58550.579435.37
Point3R0.11320.16650.21060.03770.05570.07060.61160.60390.602518.01
Ours (Core)0.04080.05710.05890.01510.02070.02360.66620.67430.672819.02
Ours-VPC-M0.04030.05790.06060.01500.02070.02180.66300.67110.669217.57
Ours-VPC-A0.04010.05610.06110.01470.02060.02420.66310.67110.669916.25
Δ Ours (Core)
vs. Point3R
↓ 64.0%↓ 65.7%↓ 72.0%↓ 59.9%↓ 62.8%↓ 66.6%↑ 8.9%↑ 11.7%↑ 11.7%↑ 5.6%

First Second Third mark the original paper’s top-three ranks; ties share and occupy ranks. Δ is the relative change from Point3R to Ours (Core). OOM denotes out-of-memory failure. Avg. FPS averages completed runs. † Reconstruction values are reported by RetrieveVGGT; FPS is measured in our setup. Ours denotes Slot3R.

Camera Pose Estimation

We compare Point3R with core Slot3R and the two viewpoint-guided pose conditioning variants. The table reports absolute trajectory error and relative translation and rotation errors; the final row shows the best Slot3R-family improvement over Point3R for each metric.

Camera pose estimation after Sim(3) alignment on 94 ScanNet scenes, 14 Sintel sequences, and 8 TUM-Dynamic sequences.
MethodScanNet (Static)SintelTUM-Dynamic
ATE RMSE ↓RPE trans ↓RPE rot ↓ATE RMSE ↓RPE trans ↓RPE rot ↓ATE RMSE ↓RPE trans ↓RPE rot ↓
Point3R0.11200.03101.16300.43600.13601.54500.06400.02800.6610
Slot3R (Core)0.08090.03221.42350.42430.12792.06870.04040.02600.7231
Slot3R-VPC-M0.08170.01651.05280.28320.05741.11540.03070.00590.4843
Slot3R-VPC-A0.07940.01621.11030.25390.05631.15290.03000.00570.4964
Δ Best Slot3R variant
vs. Point3R
↓ 29.1%↓ 47.7%↓ 9.5%↓ 41.8%↓ 58.6%↓ 27.8%↓ 53.1%↓ 79.6%↓ 26.7%

Bold marks the best Slot3R-family result for each metric. Δ reports the error reduction relative to Point3R: (Point3R − best Slot3R-family value) / Point3R × 100%. The best variant is selected independently per column: VPC-A for ATE and translation RPE, VPC-M for rotation RPE. Rotation RPE is measured in degrees.

Long-Sequence Reconstruction

Slot3R completes all evaluated settings from 600 to 1000 frames at about 19 FPS. Point3R and InfiniteVGGT run out of memory at 800 frames and beyond under the same protocol. The following plot shows reconstruction accuracy for selected baselines.

Single NVIDIA RTX 4090 (24 GB) · k = 1 · FPS excludes data loading and metric computation. Decoder and memory updates are causal; the evaluation harness encodes each sequence before recurrent decoding. Persistent memory can grow with scene coverage.

View chart data
Point-cloud Acc (lower is better)
Method600 frames700 frames800 frames900 frames1000 frames

BibTeX

arXiv:2610.12282
@misc{zhang2026slot3r,
  title  = {Slot3R: Set-Associative Spatial Memory
            for Streaming 3D Reconstruction},
  author = {Zhang, Xiyuan and Yang, Yanming and
            Xu, Kaiyuan and Li, Ruibo and Zhang, Chi},
  year   = {2026},
  eprint = {2610.12282},
  archivePrefix = {arXiv},
  primaryClass = {cs.CV},
  url    = {https://arxiv.org/abs/2610.12282}
}