computational-humor-with-multimodal-llms-methods-datasets-evaluation-and-challenges-97813955·1 events·first seen Aliases: Computational Humor with Multimodal LLMs: Methods, Datasets, Evaluation, and Challenges
A new arXiv survey examines the state of visual humor understanding in AI systems, covering memes, cartoons, and comics as test cases for non-literal reasoning. The authors organize the literature into a capability hierarchy spanning recognition, interpretation, reasoning, and generation, and trace the field's shift from task-specific fusion models to large multimodal model approaches. Key barriers identified include shortcut-prone evaluation, limited cultural coverage, weak evidence grounding, and safety and ownership concerns.