multinli-2251dd95·1 events·first seen Aliases: MultiNLI
A preregistered replication study tests whether non-upward monotonicity operators produce lower label agreement in NLI datasets, a finding from prior work on human label variation (HLV). Using unselected SNLI and MultiNLI development sets rather than the disagreement-selected ChaosNLI resource, all seven confirmatory contrasts return effects in the opposite direction and far below the smallest effect size of interest. The authors conclude the earlier negative boundary is likely a selection artifact of ChaosNLI's construction, not a population-level linguistic property. The result carries a methodological warning for HLV research: claims built on selected re-annotation resources should explicitly condition on their selection criteria.