Slot3R: Set-Associative Spatial Memory
for Streaming 3D Reconstruction
Demo
Explore the reconstructions in 3D. Rotate, zoom, or replay the reconstruction from its RGB observations.
Abstract
Streaming 3D reconstruction must preserve evidence from each frame while processing an expanding scene online. Spatial memory is a natural fit because it organizes history by reconstructed 3D location. Yet Point3R uses spatial proximity both to associate a new observation with an existing memory entry and to decide whether to fuse it, conflating co-location with state identity. Because pointers summarize image patches, nearby pointers may encode distinct surfaces, viewpoints, or visibility conditions; averaging them can destroy complementary evidence before later frames disambiguate it. We argue that location should determine address, not whether observations must merge. Slot3R realizes this principle as a training-free, set-associative retrofit that lets multiple states coexist at a shared address while keeping the pretrained Point3R backbone frozen. A bounded sparse readout further decouples persistent storage from per-frame decoder access. At 300–500 sampled frames, Slot3R reduces Point3R's point-cloud accuracy error (Acc) by 57.1%–63.1% on 7Scenes and 64.0%–72.0% on NeuralRGBD, lowers Sim(3)-aligned absolute trajectory error (ATE) on all three pose benchmarks, and remains competitive on video-depth estimation. It completes all evaluated settings from 600 to 1000 sampled frames at about 19 FPS under the same protocol, whereas Point3R and InfiniteVGGT run out of memory at 800 frames and beyond.
Qualitative Comparison
We compare point-cloud reconstructions on four 500-frame NeuralRGBD sequences. Select a scene below to compare Point3R, Slot3R, and the ground truth. Click any image to enlarge it.
Static point-cloud views from the paper. Select a scene to compare; click an image to enlarge.
Method
Given streaming RGB images, Slot3R maintains a set-associative spatial memory while keeping the pretrained Point3R backbone frozen. Each 3D address indexes a bucket containing multiple feature states. Confidence and feature similarity determine whether incoming observations are rejected, inserted, fused, or replaced. For each new frame, sparse spatial readout combines local memory with global anchors under a fixed budget of 640 tokens.
Results
3D Reconstruction
Across the 300–500-frame evaluations, Slot3R reduces Point3R’s accuracy error (Acc) by 57.1%–63.1% on 7Scenes and 64.0%–72.0% on NeuralRGBD, while improving normal consistency. The complete point-cloud table from the paper is reproduced below.
| Method | Acc ↓ | Comp ↓ | NC ↑ | Avg. FPS ↑ | ||||||
|---|---|---|---|---|---|---|---|---|---|---|
| 300 | 400 | 500 | 300 | 400 | 500 | 300 | 400 | 500 | ||
| 7Scenes | ||||||||||
| CUT3R | 0.1362 | 0.1697 | 0.1926 | 0.0744 | 0.1078 | 0.0894 | 0.5513 | 0.5447 | 0.5409 | 13.89 |
| TTT3R† | 0.0402 | 0.0502 | 0.0651 | 0.0245 | 0.0263 | 0.0295 | 0.5647 | 0.5575 | 0.5522 | 13.68 |
| StreamVGGT | OOM | OOM | OOM | OOM | OOM | OOM | OOM | OOM | OOM | OOM |
| STream3R | 0.0962 | 0.1013 | OOM | 0.0317 | 0.0426 | OOM | 0.5769 | 0.5744 | OOM | 5.82 |
| InfiniteVGGT† | 0.0432 | 0.0421 | 0.0398 | 0.0269 | 0.0269 | 0.0236 | 0.5695 | 0.5649 | 0.5617 | 3.63 |
| RetrieveVGGT | 0.0266 | 0.0276 | 0.0284 | 0.0275 | 0.0292 | 0.0294 | 0.5065 | 0.5063 | 0.5084 | 0.97 |
| ZipMap-Stream | 0.1096 | 0.1057 | 0.1021 | 0.2644 | 0.3001 | 0.3088 | 0.5223 | 0.5180 | 0.5150 | 5.12 |
| Spann3R† | 0.0765 | 0.0797 | 0.0707 | 0.0278 | 0.0275 | 0.0187 | 0.5464 | 0.5428 | 0.5402 | 41.80 |
| Point3R | 0.0634 | 0.0638 | 0.0745 | 0.0347 | 0.0308 | 0.0265 | 0.5596 | 0.5541 | 0.5510 | 16.77 |
| Ours (Core) | 0.0262 | 0.0274 | 0.0275 | 0.0154 | 0.0169 | 0.0158 | 0.5790 | 0.5789 | 0.5790 | 19.17 |
| Ours-VPC-M | 0.0257 | 0.0266 | 0.0282 | 0.0158 | 0.0164 | 0.0163 | 0.5789 | 0.5786 | 0.5792 | 17.54 |
| Ours-VPC-A | 0.0264 | 0.0269 | 0.0270 | 0.0159 | 0.0167 | 0.0163 | 0.5793 | 0.5786 | 0.5791 | 15.73 |
| Δ Ours (Core) vs. Point3R | ↓ 58.7% | ↓ 57.1% | ↓ 63.1% | ↓ 55.6% | ↓ 45.1% | ↓ 40.4% | ↑ 3.5% | ↑ 4.5% | ↑ 5.1% | ↑ 14.3% |
| NeuralRGBD | ||||||||||
| CUT3R | 0.2260 | 0.3170 | 0.3471 | 0.0883 | 0.1245 | 0.1539 | 0.5846 | 0.5656 | 0.5539 | 17.51 |
| TTT3R | 0.1058 | 0.1425 | 0.1666 | 0.0323 | 0.0775 | 0.1027 | 0.6415 | 0.6269 | 0.6221 | 17.40 |
| StreamVGGT | OOM | OOM | OOM | OOM | OOM | OOM | OOM | OOM | OOM | OOM |
| STream3R | 0.0664 | 0.0828 | OOM | 0.0149 | 0.0211 | OOM | 0.6880 | 0.6935 | OOM | 5.81 |
| InfiniteVGGT† | 0.0532 | 0.0651 | 0.0696 | 0.0244 | 0.0348 | 0.0370 | 0.6431 | 0.6503 | 0.6433 | 3.63 |
| RetrieveVGGT | 0.0438 | 0.0456 | 0.0370 | 0.0248 | 0.0316 | 0.0397 | 0.6166 | 0.6158 | 0.5132 | 1.07 |
| ZipMap-Stream | 0.1745 | 0.1985 | 0.2141 | 0.2962 | 0.4452 | 0.5305 | 0.5288 | 0.5288 | 0.5291 | 5.01 |
| Spann3R† | 0.1002 | 0.1417 | 0.1829 | 0.0262 | 0.0447 | 0.0487 | 0.5953 | 0.5855 | 0.5794 | 35.37 |
| Point3R | 0.1132 | 0.1665 | 0.2106 | 0.0377 | 0.0557 | 0.0706 | 0.6116 | 0.6039 | 0.6025 | 18.01 |
| Ours (Core) | 0.0408 | 0.0571 | 0.0589 | 0.0151 | 0.0207 | 0.0236 | 0.6662 | 0.6743 | 0.6728 | 19.02 |
| Ours-VPC-M | 0.0403 | 0.0579 | 0.0606 | 0.0150 | 0.0207 | 0.0218 | 0.6630 | 0.6711 | 0.6692 | 17.57 |
| Ours-VPC-A | 0.0401 | 0.0561 | 0.0611 | 0.0147 | 0.0206 | 0.0242 | 0.6631 | 0.6711 | 0.6699 | 16.25 |
| Δ Ours (Core) vs. Point3R | ↓ 64.0% | ↓ 65.7% | ↓ 72.0% | ↓ 59.9% | ↓ 62.8% | ↓ 66.6% | ↑ 8.9% | ↑ 11.7% | ↑ 11.7% | ↑ 5.6% |
First Second Third mark the original paper’s top-three ranks; ties share and occupy ranks. Δ is the relative change from Point3R to Ours (Core). OOM denotes out-of-memory failure. Avg. FPS averages completed runs. † Reconstruction values are reported by RetrieveVGGT; FPS is measured in our setup. Ours denotes Slot3R.
Camera Pose Estimation
We compare Point3R with core Slot3R and the two viewpoint-guided pose conditioning variants. The table reports absolute trajectory error and relative translation and rotation errors; the final row shows the best Slot3R-family improvement over Point3R for each metric.
| Method | ScanNet (Static) | Sintel | TUM-Dynamic | ||||||
|---|---|---|---|---|---|---|---|---|---|
| ATE RMSE ↓ | RPE trans ↓ | RPE rot ↓ | ATE RMSE ↓ | RPE trans ↓ | RPE rot ↓ | ATE RMSE ↓ | RPE trans ↓ | RPE rot ↓ | |
| Point3R | 0.1120 | 0.0310 | 1.1630 | 0.4360 | 0.1360 | 1.5450 | 0.0640 | 0.0280 | 0.6610 |
| Slot3R (Core) | 0.0809 | 0.0322 | 1.4235 | 0.4243 | 0.1279 | 2.0687 | 0.0404 | 0.0260 | 0.7231 |
| Slot3R-VPC-M | 0.0817 | 0.0165 | 1.0528 | 0.2832 | 0.0574 | 1.1154 | 0.0307 | 0.0059 | 0.4843 |
| Slot3R-VPC-A | 0.0794 | 0.0162 | 1.1103 | 0.2539 | 0.0563 | 1.1529 | 0.0300 | 0.0057 | 0.4964 |
| Δ Best Slot3R variant vs. Point3R | ↓ 29.1% | ↓ 47.7% | ↓ 9.5% | ↓ 41.8% | ↓ 58.6% | ↓ 27.8% | ↓ 53.1% | ↓ 79.6% | ↓ 26.7% |
Bold marks the best Slot3R-family result for each metric. Δ reports the error reduction relative to Point3R: (Point3R − best Slot3R-family value) / Point3R × 100%. The best variant is selected independently per column: VPC-A for ATE and translation RPE, VPC-M for rotation RPE. Rotation RPE is measured in degrees.
Long-Sequence Reconstruction
Slot3R completes all evaluated settings from 600 to 1000 frames at about 19 FPS. Point3R and InfiniteVGGT run out of memory at 800 frames and beyond under the same protocol. The following plot shows reconstruction accuracy for selected baselines.
Single NVIDIA RTX 4090 (24 GB) · k = 1 · FPS excludes data loading and metric computation. Decoder and memory updates are causal; the evaluation harness encodes each sequence before recurrent decoding. Persistent memory can grow with scene coverage.
View chart data
| Method | 600 frames | 700 frames | 800 frames | 900 frames | 1000 frames |
|---|
BibTeX
@misc{zhang2026slot3r,
title = {Slot3R: Set-Associative Spatial Memory
for Streaming 3D Reconstruction},
author = {Zhang, Xiyuan and Yang, Yanming and
Xu, Kaiyuan and Li, Ruibo and Zhang, Chi},
year = {2026},
eprint = {2610.12282},
archivePrefix = {arXiv},
primaryClass = {cs.CV},
url = {https://arxiv.org/abs/2610.12282}
}