International Peer-Reviewed JournalOpen AccessISSN 2456-8880
irejournals@gmail.com+91-7433024337

Home / Current Issue / Paper 1710307

1710307PublishedVol 9 · Issue 2

SpatialMoR-VGGT: Spatially Adaptive Efficient 3D Scene Reconstruction

Saksham Gupta Dr. Ramanjot Kaur

Subject area: Science,Engineering and Technology  ·  Area of research: Deep Learning

Abstract

We present SpatialMoR-VGGT, a novel framework that extends the Mixture-of-Recursions (MoR) paradigm to spatial reasoning in 3D vision tasks. While VGGT has demonstrated remarkable capabilities as a feed-forward transformer that directly infers all key 3D attributes of a scene?including camera parameters, point maps, depth maps, and point tracks?it processes all spatial regions with uniform computational depth. Our framework dynamically adjusts the recursion depth for different spatial regions of the scene, allocating more computational resources to complex areas while maintaining efficiency in simpler regions. This adaptation requires addressing fundamental differences between sequential token processing in language and spatially coherent processing in vision. We introduce spatially-aware routing mechanisms and KV caching strategies specifically designed for visual data, along with a balanced training objective that preserves spatial coherence while enabling adaptive computation. Through rigorous experimentation on standard 3D reconstruction benchmarks, we demonstrate that SpatialMoR-VGGT achieves comparable reconstruction quality to standard VGGT with 18-22% reduced computational requirements. This work establishes a foundation for adaptive computation in 3D vision tasks, with potential applications across AR/VR, robotics, and real-time 3D content creation.

Keywords

3D Reconstruction, Adaptive Computation, Recursive Transformers, Visual Geometry

How to cite this paper

Saksham Gupta, Dr. Ramanjot Kaur "SpatialMoR-VGGT: Spatially Adaptive Efficient 3D Scene Reconstruction" Iconic Research And Engineering Journals Volume 9 Issue 2 2025 Page 830-839
Saksham Gupta, Dr. Ramanjot Kaur "SpatialMoR-VGGT: Spatially Adaptive Efficient 3D Scene Reconstruction" Iconic Research And Engineering Journals, vol. 9, no. 2, Aug. 2025
Saksham Gupta, Dr. Ramanjot Kaur (2025). SpatialMoR-VGGT: Spatially Adaptive Efficient 3D Scene Reconstruction. Iconic Research And Engineering Journals, 9(2).
Saksham Gupta, Dr. Ramanjot Kaur "SpatialMoR-VGGT: Spatially Adaptive Efficient 3D Scene Reconstruction" Iconic Research And Engineering Journals, vol. 9, no. 2, Aug. 2025.
@article{1710307,
      author = {Saksham Gupta, Dr. Ramanjot Kaur},
      title = {SpatialMoR-VGGT: Spatially Adaptive Efficient 3D Scene Reconstruction},
      journal = {Iconic Research And Engineering Journals},
      year = {2025},
      volume = {9},
      number = {2},
      pages = {830-839},
      issn = {2456-8880},
      url = {https://www.irejournals.com/formatedpaper/1710307.pdf},
      abstract = {We present SpatialMoR-VGGT, a novel framework that extends the Mixture-of-Recursions (MoR) paradigm to spatial reasoning in 3D vision tasks. While VGGT has demonstrated remarkable capabilities as a feed-forward transformer that directly infers all key 3D attributes of a scene?including camera parameters, point maps, depth maps, and point tracks?it processes all spatial regions with uniform computational depth. Our framework dynamically adjusts the recursion depth for different spatial regions of the scene, allocating more computational resources to complex areas while maintaining efficiency in simpler regions. This adaptation requires addressing fundamental differences between sequential token processing in language and spatially coherent processing in vision. We introduce spatially-aware routing mechanisms and KV caching strategies specifically designed for visual data, along with a balanced training objective that preserves spatial coherence while enabling adaptive computation. Through rigorous experimentation on standard 3D reconstruction benchmarks, we demonstrate that SpatialMoR-VGGT achieves comparable reconstruction quality to standard VGGT with 18-22% reduced computational requirements. This work establishes a foundation for adaptive computation in 3D vision tasks, with potential applications across AR/VR, robotics, and real-time 3D content creation.},
      keywords = {3D Reconstruction, Adaptive Computation, Recursive Transformers, Visual Geometry},
      month = {August},
  }