Verified project record
Robbyant/lingbot-map
LingBot-Map is a feed-forward 3D foundation model for reconstructing scenes from streaming image or video data. It unifies coordinate grounding, dense geometric cues, and long-range drift correction to support efficient streaming 3D reconstruction workflows.
Project overview
LingBot-Map is a feed-forward 3D foundation model for reconstructing scenes from streaming image or video data. It unifies coordinate grounding, dense geometric cues, and long-range drift correction to support efficient streaming 3D reconstruction workflows.
- Project type
- Model Runtime · Image & Vision · Video
- Deployment
- Refer to project documentation
- License
- Apache-2.0
Best for
- Researchers and developers reconstructing 3D scenes from streaming image or video inputs who require a unified approach to coordinate grounding, dense geometric cues, and long-range drift correction.
- Users processing extended visual sequences exceeding 3000 frames who need windowed inference and flythrough rendering capabilities.
Key capabilities
- Architecturally unifies coordinate grounding, dense geometric cues, and long-range drift correction via anchor context, pose-reference window, and trajectory memory.
- Provides a browser-based viser viewer for interactive 3D visualization of reconstructed scenes.
- Processes long sequences to produce headless point-cloud flythrough MP4 videos, supporting windowed inference and extensive camera path configurations.
- Uses an ONNX sky segmentation model to filter out sky points from the reconstructed point cloud, improving visualization quality for outdoor scenes.
- Provides evaluation scripts and preprocessing tools for benchmarks like KITTI and Oxford Spires.
- Enables inference on long sequences (>3000 frames) by using a sliding window approach that resets the KV cache.
- Allows offloading per-frame predictions to CPU and reducing bidirectional scale frames to avoid out-of-memory issues.
Limitations and risks
- Performance degrades when the KV cache stores more than 320 views without a keyframe strategy.
- Maximum inference range is bounded by the longest distance seen during training unless state resetting is used.
- Users may encounter out-of-memory issues on limited GPUs.
Getting started
- Create a specific Conda environment and manually install PyTorch matching CUDA 12.8 before installing the model via pip.
- Download the required LingBot-Map model checkpoints and the skyseg.onnx ONNX sky segmentation model from HuggingFace or ModelScope.
- Run the demo using the command line interface by providing the model path and an image folder, such as running python demo.py with the model checkpoint and example/courthouse inputs.
Evidence and sources
- GitHub project description: A feed-forward 3D foundation model for reconstructing scenes from streaming data
- README: LingBot-Map has focused on: - **Geometric Context Transformer**: Architecturally unifies coordinate grounding, dense geometric cues, and long-range drift correction within a singl…
- README: [](https://arxiv.org/abs/2604.14141)
- README: **Geometric Context Transformer**: Architecturally unifies coordinate grounding, dense geometric cues, and long-range drift correction within a single streaming framework through…
- README: 📊 **Evaluation benchmark released**. We released the evaluation scripts for KITTI and Oxford Spires
AI 搜索
把需求说清楚,让项目选择更有依据
告诉我们你要解决什么、运行在哪里、哪些条件不能妥协。雷达会从已核验项目中给出主推荐、备选和采用前检查。
目标你最终想完成什么
环境本地、云端或现有技术栈
硬条件部署、界面、语言与 License
从一个真实需求开始点击只会填入搜索框,你可以继续修改
今日榜单
0