A new arXiv paper analyzes how AI systems reinforce dominant language ideologies that privilege Inner Circle (Global North) English norms and marginalize World Englishes, tracing this reproduction across training data, design protocols, evaluation benchmarks, and user feedback. The paper uses the public controversy over AI-associated vocabulary (e.g., the word 'delve') to illustrate how Global North speakers police English norms of Global South users. It identifies a 'standardisation paradox' where generative AI simultaneously homogenizes English toward standard forms while potentially pluralizing it through diverse corpora. The authors argue for more inclusive AI design approaches that recognize the plurality of Englishes.
A structured scholarly dialogue among five sociolinguists from World Englishes fields examines how generative AI tools affect linguistic inclusivity in academic writing and publishing. The paper argues that GenAI tends to reinforce dominant language norms and marginalize minoritized English varieties while also holding potential to democratize writing access. Contributors call for equity-informed policies, critical AI literacy, and inclusive co-design in GenAI development.
A preprint from arXiv proposes applying literary disciplines — comparative literature, narratology, critical theory, and world literature — as a framework for building more culturally literate AI systems. The essay argues that LLMs currently enact a 'massive, automated, and monolingual' form of cultural encounter and that structural monolingualism is a core problem. It develops a layered framework addressing global AI textuality through macrostructure, circulation, and untranslatability.
A new arXiv preprint develops a formal mathematical framework to analyze how widespread LLM-assisted writing may reduce population-level linguistic diversity, a phenomenon the authors term 'linguistic monoculture.' The paper models authors and LLMs as distributions over linguistic features and analyzes three interaction mechanisms—fixed shared models, recursively updated shared models, and personalized models—characterizing equilibria and convergence rates for each. A key finding is that individually rational authors may over-conform relative to the social optimum because they do not internalize the value of their own distinctiveness, creating a negative externality. The 'price of monoculture' can grow without bound when distinctiveness dominates authenticity in the utility model.
New research suggests that large language models not only inherit human biases from training data but can also develop novel biases of their own when used in hiring contexts. The study raises concerns about AI résumé screening systems that operate before any human review. This adds to a growing body of evidence that LLM-based hiring tools may produce unfair outcomes in ways that are difficult to anticipate or audit.
A new arXiv preprint proposes a theoretical framework for understanding NLP work on culture as a 'material-discursive practice,' drawing on Karen Barad's concept of the agential cut to argue that model, data, annotation, and evaluation choices actively shape the cultural phenomena they purport to measure. The author illustrates this through six case studies involving television and film dialogue analysis, including examination of how LLMs erase cultural markers, attune to historical material, and exercise agency in agentic workflows. The paper calls for a theory-driven, empirically rigorous, and culturally contingent research program that treats methodological choices as ethical commitments. This is primarily a philosophy-of-science and methodology contribution to the cultural NLP subfield.
A new arXiv paper introduces a controlled evaluation framework to disentangle language proficiency from culture-specific knowledge access in LLMs. Using real-world cultural questions across 13 locales and ~80 models, the authors apply item response theory to show that while English dominates on culture-agnostic questions, local languages yield a consistent knowledge-access advantage on culture-specific questions once proficiency differences are factored out. The finding challenges the common interpretation that weaker local-language accuracy implies weaker cultural knowledge, and has implications for how multilingual and regionally-aligned models are evaluated.
A new arXiv preprint introduces a multilingual evaluation framework using 414 proverbs across 15 languages to assess whether LLMs preserve culturally grounded meaning when generating narratives. Using four LLMs to produce 13k narratives, the study finds that cross-lingual prompting preserves proverb-level semantic meaning but systematically redistributes agency, social positioning, and narrative structure. Strong inter-model convergence across architectures suggests multilingual LLMs rely on shared semantic abstractions. The authors argue that semantic similarity metrics alone overestimate cultural preservation in multilingual evaluations.
Researchers introduce DiaLLM, a framework that continually pretrains three open-weight LLM families on the International Corpus of English to study dialect adaptation across Australian, Indian, and Northern British English. The study finds that dialectal robustness (understanding) and generation are dissociated: benchmarks are shaped by pretraining and SFT, while alignment reshapes generation in ways benchmarks fail to capture. A key finding is that the alignment method most aggressively optimizing dialectal reward is not preferred by human evaluators, revealing a reward-quality gap. Code, checkpoints, and preference datasets are released.