scenebind-f29caa0f·1 events·first seen Aliases: SceneBind, SceneBind Matching
SceneBind is a new multimodal representation framework that jointly encodes semantic and 3D spatial structure across vision, audio, and language modalities. It represents scenes as semantic-spatial entities using global embeddings plus object-centric slots, addressing a gap in existing omni-modal encoders that handle 'what' but not 'where'. The authors introduce a binaural audio-visual dataset with spatial annotations and a matching scheme for cross-modal retrieval and object grounding, achieving state-of-the-art on scene/spatial retrieval with strong zero-shot transfer to audio-visual localization.