harmonizing-ai-safety-thresholds-340cffa4·1 events·first seen Aliases: Harmonizing AI Safety Thresholds
A new arXiv preprint proposes a methodology for deriving harmonized capability thresholds across frontier AI companies, addressing the current inconsistency in published thresholds that makes third-party verification and cross-company comparison difficult. The authors cover three risk domains: cyber misuse, biological misuse, and automated AI R&D, using expected harm modeling for the first two and observed AI progress rates for the third. The work explicitly flags a potential race-to-the-bottom dynamic in safety standards when thresholds are not harmonized, and identifies empirical gaps in existing approaches.