sparse-autoencoders-encode-both-concepts-and-functions-the-downstream-geometry-of-feature-effects-9846f45c·1 events·first seen Aliases: Sparse Autoencoders Encode Both Concepts and Functions: The Downstream Geometry of Feature Effects
A new arXiv preprint introduces Feature-Effect Geometry Analysis (FEGA), an unsupervised framework that studies how sparse autoencoder (SAE) feature interventions affect model logits rather than examining internal feature geometry. The authors find that consistent one-dimensional causal effects are rare across SAE features, and distinguish 'value-like' features (tied to static factual attributes) from 'pointer-like' features (context-dependent operations), with the latter exhibiting diffuse, hard-to-steer effects. The key finding is that a feature can be interpretable and causally relevant without providing a stable steering direction, which has significant implications for mechanistic interpretability workflows that rely on SAE-based steering.