A new arXiv preprint proposes a methodology for deriving harmonized capability thresholds across frontier AI companies, addressing the current inconsistency in published thresholds that makes third-party verification and cross-company comparison difficult. The authors cover three risk domains: cyber misuse, biological misuse, and automated AI R&D, using expected harm modeling for the first two and observed AI progress rates for the third. The work explicitly flags a potential race-to-the-bottom dynamic in safety standards when thresholds are not harmonized, and identifies empirical gaps in existing approaches.
Anthropic released a significant revision to its Responsible Scaling Policy (RSP), its risk governance framework for managing catastrophic risks from frontier AI. The update introduces two explicit capability thresholds—autonomous AI R&D and CBRN weapons uplift—that trigger mandatory upgrades to AI Safety Level (ASL) standards, with current models operating under ASL-2. New elements include safety-case-inspired documentation processes, internal governance stress-testing, and external expert input mechanisms, drawing on risk management practices from high-consequence industries like biosafety.
Anthropic announced a funded initiative to source third-party evaluations measuring advanced AI capabilities and safety risks, with priority areas including cybersecurity, CBRN threats, model autonomy, national security risks, social manipulation, and misalignment. The initiative is tied to Anthropic's Responsible Scaling Policy and AI Safety Level (ASL) framework, aiming to address a gap between demand and supply of high-quality safety-relevant evals. Proposals are solicited via an application form, with Anthropic framing the effort as benefiting the broader AI safety ecosystem rather than just internal use.
Anthropic released its Responsible Scaling Policy (RSP), a formal framework of technical and organizational protocols for managing catastrophic risks from increasingly capable AI systems. The policy introduces AI Safety Levels (ASL-1 through ASL-5+), modeled on US biosafety level standards, requiring progressively stricter safety, security, and operational standards as models become more capable. Current Claude models are classified as ASL-2; ASL-3 triggers stricter deployment constraints including adversarial red-teaming requirements. The policy has been approved by Anthropic's board and is intended as a template for industry-wide adoption.
OpenAI and Anthropic conducted a first-of-its-kind cross-lab safety evaluation, testing each other's frontier models across dimensions including misalignment, instruction following, hallucinations, and jailbreaking resistance. The collaboration represents a novel form of inter-lab safety research cooperation. Findings highlight both progress and ongoing challenges in AI safety, and establish a potential template for future cross-organizational evaluations.
Anthropic published a policy proposal calling for a targeted AI transparency framework applicable at federal, state, or international levels, targeting only the largest frontier AI developers (suggested thresholds: ~$100M annual revenue or ~$1B R&D/capex). The framework would require covered developers to publicly disclose a Secure Development Framework covering CBRN and misalignment risks, publish system cards at deployment, self-certify compliance, and face legal liability for false statements. The proposal is explicitly lightweight and flexible, designed to avoid prescriptive standards while creating accountability mechanisms and whistleblower protections during the period before comprehensive safety standards are established.
Dario Amodei delivered prepared remarks at the UK AI Safety Summit (November 2023) explaining Anthropic's Responsible Scaling Policy (RSP), which was the first such policy published by a major AI lab. The RSP introduces AI Safety Levels (ASL-1 through ASL-4), modeled on biosafety level frameworks, with capability thresholds triggering mandatory safeguards before further training or deployment. Key implementation lessons include deep executive involvement, integrating RSP requirements into product roadmaps, and formal accountability through Anthropic's board and Long Term Benefit Trust. The remarks outline specific ASL-3 requirements around CBRN misuse prevention and security, and preview ASL-4 criteria involving near-human autonomy or becoming a primary source of global security threats.
A preprint from arXiv argues that AI safety discourse over-indexes on visible, model-centric failures while neglecting quieter systemic risks in deployed socio-technical systems. The authors propose a five-layer diagnostic framework covering epistemic, control, temporal, organizational, and ecosystem integrity. The paper identifies under-recognized risk patterns including uncertainty laundering, prompt injection, memory poisoning, evaluation deception, and model collapse, and calls for a shift from model-centric evaluation toward socio-technical reliability. It concludes with design, governance, and research agenda recommendations.
OpenAI published a post describing its use of independent experts to evaluate frontier AI systems through third-party testing. The initiative aims to strengthen safety validation, verify safeguards, and increase transparency around capability and risk assessments. The announcement signals a continued push toward external accountability mechanisms for frontier model evaluation.