sm4rt-82fe91ff·1 events·first seen Aliases: SM4RT
SM4RT is a new transformer architecture for end-to-end 4D reconstruction from monocular RGB video that models scene motion as a structured set of rigid-body transformations in SE(3) rather than independent point-wise displacements. The method decomposes scene dynamics into a compact set of motion bases represented as temporal sequences of 6D twists, with per-pixel assignment weights ensuring physically coherent rigid-body trajectories. A parallel motion geometry encoder-decoder jointly infers 3D geometry, world-coordinate motion, and kinematic structure in a single forward pass. The work targets a known gap in Geometry Foundation Models, which have advanced monocular 3D reconstruction but struggled with dynamic 4D understanding.