charm-b3e5879f·2 events·first seen Aliases: CHARM
Researchers introduce CHARM, a graph foundation model designed for zero-shot transfer across multimodal graphs containing text, images, and other node-associated data. The model replaces isolated node representations with hierarchical graph contexts that capture cross-modal relations and map domain-specific patterns to shared high-level concepts, reducing the need for target-domain fine-tuning. A modality-aware encoder converts these representations into graph tokens fed to a large language model backbone. Experiments show consistent improvements on zero-shot multimodal graph tasks, addressing a gap where existing GNN-based GFMs require downstream adaptation and LLM-based methods are largely unimodal.
CHARM (Channel-Aware Representation Model) is a new Transformer-based architecture for general-purpose representation learning over heterogeneous multivariate time series. It integrates channel-level textual descriptions into a permutation-equivariant encoder trained with a Joint Embedding Predictive Architecture (JEPA) and a novel temporally stable embedding loss. The model achieves strong performance across anomaly detection, classification, and forecasting tasks using only a linear probe, with text descriptions primarily serving as channel identifiers enabling cross-dataset generalization.