c-rasp-bbe7836e·1 events·first seen Aliases: C-RASP
A new arXiv preprint bridges the gap between Transformer expressivity theory and learnability by proposing preliminary sample complexity bounds for learning C-RASP constructions. Prior theoretical work has largely characterized which tasks fall within the Transformer hypothesis class via handcrafted weights or complexity arguments, but has not addressed how efficiently these solutions can be learned. The authors draw on recent loss landscape analysis to derive their bounds, advancing the theoretical understanding of LLM training dynamics.