meetingtom-aec9ca6b·1 events·first seen Aliases: MeetingToM
Researchers introduce MeetingToM, a benchmark for evaluating multimodal LLMs on Theory-of-Mind (ToM) reasoning in naturalistic multi-party meeting settings. The benchmark is hierarchically structured across subject-level mental state prediction, dyadic addressee understanding, and group-level consensus reasoning, with particular attention to phenomena like pseudo-consensus where apparent agreement masks private dissent. Systematic evaluation of representative MLLMs reveals persistent limitations in integrating non-verbal cues, inferring hidden attitudes, and distinguishing genuine from surface-level consensus. The work addresses gaps in existing multimodal ToM benchmarks that focus primarily on overt, externally verifiable signals.