storyteller-ecc58942·1 events·first seen Aliases: StoryTeller
StoryTeller is a training-free framework that enables video-language models to produce story-aware audio descriptions for blind and low-vision audiences by maintaining a verified narrative memory across scenes. The system requires only raw video and a movie title, using semantic filtering and VLM verification to accept only video-supported facts, with no fine-tuning or precomputed character banks needed. The authors also introduce StoryAD-QA, a QA-based benchmark for evaluating narrative coherence in generated audio descriptions. Experiments show improvements over baselines on standard AD benchmarks and diverse long-form videos.