fewnerd-dfdb1ca7·1 events·first seen Aliases: FewNERD
A new arXiv preprint introduces FSE, a span-based NER model using a dual-expert architecture to address catastrophic forgetting in continual learning settings. A shared 'fast expert' filters unlikely spans at the token level, while task-specific 'slow experts' handle span classification, reducing per-task learning burden. The method also introduces a length-decay negative sampling strategy to handle span imbalance. Experiments on OntoNotes and FewNERD synthetic datasets show state-of-the-art results in continual NER scenarios.