A new arXiv preprint proposes a two-component methodology for tracking AI trustworthiness across a system's lifecycle: a formal framework that learns interpretable trustworthiness levels from measurable dimensions using decision trees, and a governance procedure for design-time labeling, post-deployment monitoring, reassessment, and reporting. The approach introduces diagnostics such as boundary margins and profile drift to detect when a deployed system's trustworthiness has changed materially. The work targets the gap between high-level AI governance frameworks and the concrete, auditable documentation needed for regulatory conformity.
A preprint from arXiv argues that AI safety discourse over-indexes on visible, model-centric failures while neglecting quieter systemic risks in deployed socio-technical systems. The authors propose a five-layer diagnostic framework covering epistemic, control, temporal, organizational, and ecosystem integrity. The paper identifies under-recognized risk patterns including uncertainty laundering, prompt injection, memory poisoning, evaluation deception, and model collapse, and calls for a shift from model-centric evaluation toward socio-technical reliability. It concludes with design, governance, and research agenda recommendations.
A new arXiv preprint proposes a methodology for deriving harmonized capability thresholds across frontier AI companies, addressing the current inconsistency in published thresholds that makes third-party verification and cross-company comparison difficult. The authors cover three risk domains: cyber misuse, biological misuse, and automated AI R&D, using expected harm modeling for the first two and observed AI progress rates for the third. The work explicitly flags a potential race-to-the-bottom dynamic in safety standards when thresholds are not harmonized, and identifies empirical gaps in existing approaches.
A systematic study audits whether converting instruction-tuned LLMs into reasoning models via SFT, RL-based post-training, or distillation preserves alignment behaviors such as safe refusal, bias avoidance, and privacy protection. Across six trustworthiness dimensions, the authors find consistent alignment regressions—including increased toxicity, amplified stereotyping, miscalibrated refusal, and privacy leakage—even as reasoning benchmark scores improve. The regressions are quantified via KL divergence from the instruction-tuned baseline, suggesting behavioral drift is a systematic byproduct of reasoning post-training. The paper argues trustworthiness metrics should be reported alongside reasoning capability gains.
This commentary piece argues that as AI-generated advice becomes more consequential, users need systematic methods to evaluate AI reliability and quality—analogous to a job interview process. The author proposes frameworks for assessing AI outputs before trusting them for important decisions. The piece addresses the practical challenge of calibrating trust in AI systems across different use cases.
Eticas presents a structured AI auditing framework that bridges risk cataloging to executable audit methodology, demonstrated end-to-end on PII leakage testing against GPT-4-0314. The taxonomy organizes 76 active subcategories across 10 categories with mappings to 18 external frameworks, and is published under CC BY 4.0 with SKOS/JSON-LD distributions. The key contribution is an operationalization layer that converts named risks into measurable, severity-graded findings — addressing a gap the authors identify across at least 74 existing AI risk taxonomies. The PII leakage demonstration shows disclosure rates ranging from 0% to 84% under adversarial conditioning, graded as SYSTEMIC severity.
A new arXiv survey covers 1,250 papers (2024–2026) on AI self-improvement, proposing a two-axis taxonomy distinguishing what is improved (behavior, policy, evaluator, or research process) from the degree of loop closure (human-in-the-loop to fully closed). The authors construct a verification hierarchy for self-evaluation signals—from formal verifiers (strongest) to intrinsic self-assessment (weakest)—and find that demonstrated self-improvement strength tracks this hierarchy while failure modes (self-confirming loops, model collapse, diversity collapse) arise from its violations. The paper argues that 'research direction-setting' remains the key bottleneck keeping humans in the loop, and identifies governance-grade measurement of self-improvement as the most underpopulated niche in the field. The work connects technical RSI limits to safety and governance concerns raised by frontier labs experimenting with closed-loop AI research.
OpenAI contributed to a multi-stakeholder report co-authored by 58 researchers across 30 organizations, including Mila, CSET, and the Schwartz Reisman Institute. The report identifies 10 mechanisms for improving the verifiability of claims about AI systems. These tools are intended to help developers demonstrate safety, security, fairness, and privacy properties, while enabling policymakers and civil society to evaluate AI development processes.
A new arXiv preprint proposes a Bayesian inference and decision-audit framework for interpreting public AI evaluation archives (LiveBench, Open LLM Leaderboard v2, LMArena, GAIA, tau-bench) as longitudinal time series rather than terminal leaderboards. The paper demonstrates that a single terminal snapshot is compatible with multiple distinct performance histories, yielding ambiguous timing estimates for reaching capability ceilings. A candidate selection-aware frontier model is shown to fail synthetic recovery, objective-archive prediction, preference transfer, and uncertainty calibration, with fixed audit gates rejecting its stronger claims. The work proposes an archive-and-adjudication protocol to reconstruct evaluation histories and falsify unsupported frontier capability claims.