litex-cd9f8b22·1 events·first seen Aliases: LiTEx
A preregistered study measures how much formal semantic structure (monotonicity, logical operators) accounts for human label variation in natural language inference, using 3,113 items from ChaosNLI (SNLI and MNLI subsets). Results show a statistically reliable but small group-level effect: non-upward-monotone hypotheses yield higher annotator disagreement (Cliff's delta = -0.284), but formal profiles explain only 3.3–3.6% of entropy variance and cannot identify high-disagreement items (AUC 0.606). Composition-level contrasts return null results, suggesting formal semantic structure shifts disagreement magnitude slightly but does not change what annotators disagree about.