A new arXiv paper proposes an evaluation framework for LLMs used as social simulators, arguing that matching final outcomes is insufficient — simulators must also replicate the underlying rationale-derived reasoning paths. Using a 94-person sunscreen concept test, the authors map open-ended human rationales into signed 'reason states' and show that LLM-simulated reasons, while superficially plausible, frequently mirror the stimulus material rather than recovering the respondent's actual acceptance or rejection logic. The framework provides an interpretable audit for whether LLM social simulators are reasoning correctly, not just producing correct-looking outputs.
A new arXiv paper proposes a three-dimensional framework for understanding how LLMs revise moral judgments under social pressure, drawing parallels to human social psychology. Across three studies, the authors find that models' judgment updates are structured by the distance of an incoming view from the model's initial position, source attribution, and coalition pressure. The work reframes sycophancy not as a one-dimensional failure but as one expression of a broader social-influence-driven updating process. The framework aims to provide principled criteria for distinguishing constructive belief revision from sycophantic compliance in alignment-relevant contexts.
Researchers introduce a counterfactual context revision framework to audit how LLMs simulate individual users' stances in online discussions. By applying controlled text-only and multimodal (meme-based) revisions to conversational contexts, they measure how readily simulated stances shift in response to semantically independent changes. Results show effective and robust stance transitions across both revision types and polarization-preference mechanisms, raising concerns about whether LLM simulations reflect genuine user-specific beliefs or are highly context-sensitive artifacts. The work contributes an evaluation framework and highlights risks of using LLMs to model online opinion dynamics.
An opinion paper from arXiv argues that LLM self-explanations — natural language rationalizations of model decisions — can be plausible and actionable even when they do not faithfully reflect the model's underlying reasoning. The authors critique standard XAI evaluation protocols for self-explanations and propose guidelines covering plausibility, faithfulness, and a third criterion: actionability. The paper reframes the self-explanation debate away from faithfulness as the sole standard toward practical utility for diverse stakeholders.
A new arXiv preprint investigates whether LLMs can replicate systematic human decision-making biases in route choice scenarios without explicitly specifying cumulative prospect theory (CPT) parameters. The authors design a behavioral evaluation framework and find that LLMs reproduce non-rational choice patterns consistent with CPT effects under uncertainty. The findings suggest LLMs could serve as a scalable substitute for survey-based parameter calibration in agent-based behavioral simulations, addressing a longstanding bottleneck in large-scale human decision modeling.
Researchers conducted a population-matching experiment evaluating 25 LLMs on conditional inference tasks across four languages, comparing model behavior to matched human populations. The study finds that LLMs function as accurate semantic operators but systematically fail to capture pragmatic enrichments—context-sensitive inferences beyond literal logical meaning—that humans apply effortlessly. Model performance on pragmatic reasoning is not predicted by open vs. closed weights, training orientation, or architecture type, suggesting pragmatic reasoning remains an emergent and unreliable capability. The findings contribute to ongoing debates about whether LLMs reason like humans or merely approximate surface-level linguistic patterns.
A new arXiv paper evaluates human participants and 25 LLMs on commonsense causal reasoning tasks, finding similar error patterns in both groups. The authors identify specific attention heads driving LLM responses that implement pattern-matching, and show these heads can predict human reasoning errors caused by superficially irrelevant prompt details. The findings challenge the common assumption that human reasoning relies on principled abstract world models while LLMs merely pattern-match, suggesting both may share a more unified cognitive mechanism.
A systematic study audits whether converting instruction-tuned LLMs into reasoning models via SFT, RL-based post-training, or distillation preserves alignment behaviors such as safe refusal, bias avoidance, and privacy protection. Across six trustworthiness dimensions, the authors find consistent alignment regressions—including increased toxicity, amplified stereotyping, miscalibrated refusal, and privacy leakage—even as reasoning benchmark scores improve. The regressions are quantified via KL divergence from the instruction-tuned baseline, suggesting behavioral drift is a systematic byproduct of reasoning post-training. The paper argues trustworthiness metrics should be reported alongside reasoning capability gains.
A preprint from arXiv demonstrates that an LLM pipeline can automate reproducibility assessments of published social and behavioral science studies, recovering original effect sizes in 41% of cases (vs. 34% for human reanalysts) and reaching the same qualitative conclusion in 96% of cases (vs. 74% for humans). The study evaluated 76 published studies with predefined claims. The results suggest LLMs could serve as a scalable tool for systematic auditing of empirical research, addressing the resource-intensive nature of traditional reproducibility efforts.