alibaba-damo-academy-9f10b4dd·3 events·first seen Aliases: Alibaba DAMO Academy
Researchers from Alibaba DAMO Academy introduce CamVLA, a Vision-Language-Action model that eliminates the need for explicit camera calibration during robot deployment. The model decouples manipulation controls from camera geometry by predicting camera-centric end-effector actions and a 6-DoF hand-eye matrix, composing them into robot base-frame actions via deterministic geometric transformation. Operating on a single monocular RGB image without depth or calibration data, CamVLA improves success rates across diverse unseen viewpoints in both simulation and real-world evaluations.
Researchers from Alibaba DAMO Academy introduce ClinHallu, a benchmark of 7,031 validated instances designed to identify where hallucinations originate within medical MLLM reasoning pipelines. Each instance is annotated with a structured reasoning trace decomposed into Visual Recognition, Knowledge Recall, and Reasoning Integration stages, with stage-replacement interventions to measure the causal impact of correcting each stage. The paper also demonstrates that trace-supervised fine-tuning reduces stage-wise hallucinations, offering both diagnostic and mitigation value for clinical AI systems.
Alibaba's Qwen team introduces OFA (One-For-All), a unified multimodal pretrained model designed to handle both understanding and generation tasks across multiple modalities within a single framework. The model is pretrained using instruction-based multitask pretraining to endow it with diverse capabilities. This work was published in late 2022 as part of the broader wave of generalist multimodal models. It represents an early effort toward a single model architecture capable of spanning vision, language, and cross-modal tasks.