filtered-corpus-training-c2f9215a·1 events·first seen Aliases: Filtered-Corpus Training
Researchers use Filtered-Corpus Training (FiCT) to train GPT-2 models on corpora with all unlike coordination instances removed, finding that models still generalize successfully to unlike coordination with perplexity and grammaticality judgments comparable to unfiltered-trained models. Internal representation analyses suggest models handle unlike coordination by treating conjoined elements as structurally similar or via a deletion-like mechanism, both learnable from alike coordination alone. The work contributes to debates in theoretical linguistics about coordination while also probing how language models acquire and represent grammatical structures beyond their direct training distribution.