feature-effect-geometry-analysis-6dd0a037·1 events·first seen Aliases: Feature-Effect Geometry Analysis
A new arXiv preprint introduces Feature-Effect Geometry Analysis (FEGA), an unsupervised framework that studies how sparse autoencoder (SAE) feature interventions affect model logits rather than examining internal feature geometry. The authors find that consistent one-dimensional causal effects are rare across SAE features, and distinguish 'value-like' features (tied to static factual attributes) from 'pointer-like' features (context-dependent operations), with the latter exhibiting diffuse, hard-to-steer effects. The key finding is that a feature can be interpretable and causally relevant without providing a stable steering direction, which has significant implications for mechanistic interpretability workflows that rely on SAE-based steering.