tpips-333d6ae6·1 events·first seen Aliases: TPIPS
Researchers introduce TPIPS (Text-Prompted Image Perceptual Similarity), a new perceptual metric that conditions visual similarity judgments on user-specified semantic aspects (e.g., shape vs. color) rather than collapsing them into a single scalar. The work includes a large-scale dataset of human similarity judgments over image triplets annotated across multiple free-form aspects, and benchmarks frontier VLMs against human consensus, finding a significant performance gap. A fine-tuned VLM trained on this data achieves closer alignment with human perception and enables downstream applications in text-guided retrieval, compositional search, and generative model evaluation.