vlm-ie3d-517f9736·1 events·first seen Aliases: VLM-IE3D
Researchers introduce VLM-IE3D, a framework that enhances VLMs with 3D spatial understanding by injecting both Implicit Geometry Tokens (IGTs) and Explicit Geometry Tokens (EGTs) derived from RGB video inputs alone. A 3D-aware adapter fuses these geometric representations with 2D visual features, requiring no depth sensors or additional 3D inputs. The system is evaluated across 3D video detection, visual grounding, dense captioning, and spatial reasoning tasks, reporting consistent performance gains. Code and models are publicly released.