olmoe-1b-7b-79f279d9·2 events·first seen Aliases: OLMoE-1B-7B, OLMo-7B
A new arXiv paper investigates Super Weights — individual LLM parameters whose removal catastrophically degrades performance — and finds that their apparent importance does not translate into trainability. Training Super Weights in isolation (100 to 8,192 parameters) collapses accuracy to random-guessing on OLMo-1B and OLMo-7B, while training an equal number of randomly chosen parameters in the same layers improves over baseline. LoRA, which applies structured low-rank updates across entire layers, succeeds with only 0.16% of parameters, and constraining LoRA updates at Super Weight coordinates yields no benefit. The findings challenge the assumption that parameter importance implies parameter trainability and suggest effective fine-tuning requires structured decompositions over full layers rather than targeted sparse updates.
MobileMoE introduces a family of on-device MoE language models with 0.3–0.9B active parameters and 1.3–5.3B total parameters, targeting mobile deployment under memory and compute constraints. The authors derive an on-device MoE scaling law identifying a sweet spot of moderate sparsity with fine-grained and shared experts, then train models through a four-stage recipe including quantization-aware training on open-source data. Across 14 benchmarks, MobileMoE matches or exceeds leading dense on-device LLMs with 2–4× fewer inference FLOPs, and delivers 1.8–3.8× faster prefill and 2.2–3.4× faster decode than dense baselines on commodity smartphones at comparable INT4 memory.