A new arXiv preprint introduces a computational framework for quantifying creativity as selective transformation across lexical, semantic, conceptual, structural, and narrative levels of literary texts. Drawing on sociological theories of imitation from Tarde and Baldwin, the authors apply directional alignment and calibrated similarity measures to historically documented literary relationships. The work produces 'transformation profiles' that characterize where imitation persists and where creative divergence occurs across text pairs.
This paper empirically validates a creative quality metric from a companion work (Calibrated Surprise, Zou & Xu 2026a) under strict low-resource conditions: ~100 expert chain-of-thought annotations and a small base model. The authors introduce Creative Quality Alignment (CQA) as a class of engineering methods and identify a systematic bias in public alignment datasets toward craft knowledge, with weak coverage of audience modeling and reality-logic. A theoretical argument based on 'architectural duality' in single conditional distribution LLMs is offered to explain why so few examples suffice, distinguishing the result from purely empirical findings like LIMA.
A preprint from arXiv proposes applying literary disciplines — comparative literature, narratology, critical theory, and world literature — as a framework for building more culturally literate AI systems. The essay argues that LLMs currently enact a 'massive, automated, and monolingual' form of cultural encounter and that structural monolingualism is a core problem. It develops a layered framework addressing global AI textuality through macrostructure, circulation, and untranslatability.
A new arXiv preprint introduces the Human Creativity Benchmark (HCB), which collects 15,000 professional judgments across five creative domains and three workflow phases to evaluate creative AI. The benchmark explicitly separates 'convergence' (shared professional standards) from 'divergence' (legitimate taste variation), arguing that collapsing these into a single quality metric discards actionable information. Key findings include that convergence concentrates on verifiable dimensions like technical correctness, while divergence concentrates on aesthetic direction and conceptual risk, and that no model excels uniformly across all workflow phases.
DEFINED is a computational framework for automated creativity assessment in debate scenarios, operationalizing creativity through an eight-dimensional hierarchical metric system implemented via a pretrained autoregressive language model with a hierarchical scoring head. The system addresses data scarcity through constrained data augmentation and mixed-granularity training from limited expert-annotated data. It outperforms prompt-based LLM evaluators and existing debate scoring methods on authentic competition data. The work is relevant to AI evaluation methodology and the broader question of whether LLMs can reliably assess complex human cognitive outputs.
A new arXiv preprint introduces Metaphor Tracer, a method for scoring token positions in a single forward pass on two properties — an 'aggregator' (whether a position consolidates the whole text) and a 'differentiator' (whether other tokens are transiently carried into its subspace) — without any training. The approach is validated across three unrelated models using engineered registers and psychoanalyst-annotated clinical transcripts as ground truth, achieving 34/36 cell matches on the latter. The work argues that structural value is relational (a property of a token's place in a specific text) rather than essentialist (a property of its vector alone), and finds that instruction tuning raises discourse-reading fidelity without affecting type-level transfer.
A new arXiv preprint introduces Semantic Reference Frames (SemRF), a formal framework for analyzing how language model computation evolves across transformer layers via the residual stream. The method uses anchor-based coordinates with pseudo-inverse tying to separate semantic measurement from residual dynamics, enabling stable cross-layer comparison without measurement drift. The framework defines layerwise semantic trajectories, Voronoi-based coarse cell assignments, and a minimum-action canonical trace, with theoretical links to parameter efficiency and knowledge density. This is a mechanistic interpretability contribution aimed at understanding internal model computation.
A new arXiv preprint develops a formal mathematical framework to analyze how widespread LLM-assisted writing may reduce population-level linguistic diversity, a phenomenon the authors term 'linguistic monoculture.' The paper models authors and LLMs as distributions over linguistic features and analyzes three interaction mechanisms—fixed shared models, recursively updated shared models, and personalized models—characterizing equilibria and convergence rates for each. A key finding is that individually rational authors may over-conform relative to the social optimum because they do not internalize the value of their own distinctiveness, creating a negative externality. The 'price of monoculture' can grow without bound when distinctiveness dominates authenticity in the utility model.
A new arXiv preprint introduces a multilingual evaluation framework using 414 proverbs across 15 languages to assess whether LLMs preserve culturally grounded meaning when generating narratives. Using four LLMs to produce 13k narratives, the study finds that cross-lingual prompting preserves proverb-level semantic meaning but systematically redistributes agency, social positioning, and narrative structure. Strong inter-model convergence across architectures suggests multilingual LLMs rely on shared semantic abstractions. The authors argue that semantic similarity metrics alone overestimate cultural preservation in multilingual evaluations.