Skip to content Get the almanac in your inbox © 2026 Call Me Almanac · Content is AI-generated and may be inaccurate.
Significance High 8–10 Notable 6–7 Minor 4–5
OpenAI published results on multiple long-standing open problems across mathematics and theoretical computer science, covering areas including geometry, cryptography, and complexity theory. The announcement comes directly from OpenAI's blog, suggesting these are AI-assisted or AI-driven mathematical discoveries. This is potentially significant as a demonstration of frontier AI capability in formal reasoning and mathematical research.
3 Simon Willison'S Weblog · 8h ago · source ↗ Simon Willison published datasette-agent 0.4a0, an alpha release of an agent harness built around the Datasette data exploration tool. The release is noted on his personal site with minimal detail in the body. Datasette-agent represents ongoing development in the agent-tool ecosystem, connecting LLM-driven agents to structured data querying.
4 Simon Willison'S Weblog · 8h ago · source ↗ Simon Willison published an early alpha release (0.1a0) of llm-mcp-client, a plugin that adds Model Context Protocol client support to his LLM command-line tool. The release extends the LLM ecosystem to interoperate with MCP servers, enabling tool use via the standardized protocol. This is an early-stage but concrete addition to the growing MCP tooling ecosystem.
5 Simon Willison'S Weblog · 8h ago · source ↗ Simon Willison writes about renewed interest in the Model Context Protocol following developments around stateless MCP, and announces two new tools: mcp-explorer and datasette-mcp. The post reflects on how stateless operation changes the practical appeal of MCP for tool integration. This is a practitioner-level signal about MCP adoption patterns and tooling ecosystem growth.
DeepSeek released V4-Flash 0731, a new model variant, on July 31, 2026. The Latent Space AINews digest notes this as the sole significant development of the day. The release extends DeepSeek's V4 model family with a Flash-tier (likely faster/cheaper) variant.
5 Simon Willison'S Weblog · 8h ago · source ↗ Simon Willison flagged the release of DeepSeek-V4-Flash-0731, a new checkpoint in the DeepSeek V4 Flash model line. The post is a brief link-log entry with minimal commentary. DeepSeek's Flash variants are typically smaller, faster inference-optimized models in their flagship family.
OpenAI banned accounts associated with a previously unreported Russia-origin influence operation dubbed 'Bad Grammar', which used AI to generate English- and Russian-language comments on Telegram targeting narratives around Ukraine, Moldova, the Baltic states, and US politics. This is an OpenAI threat intelligence disclosure documenting a covert influence operation leveraging their models. The case adds to the growing body of evidence that AI tools are being operationalized for state-linked information warfare.
OpenAI terminated accounts associated with 'Spamouflage', a PRC-linked covert influence operation that used OpenAI tools to research social media activity, generate posts, and debug a previously unreported website. The action represents OpenAI's enforcement against state-linked misuse of its AI systems for information operations. This is notable as a documented case of a major lab detecting and disrupting foreign influence activity leveraging generative AI.
OpenAI identified and banned accounts associated with the Russia-origin 'Doppelganger' influence operation, which used OpenAI tools to generate anti-Ukraine social media comments, translations, and website copy in multiple languages. The action is part of OpenAI's ongoing effort to detect and disrupt malicious uses of its AI systems. This case illustrates the use of generative AI in state-linked information operations.
OpenAI banned accounts linked to a previously unreported Israel-origin covert influence operation dubbed 'Zero Zeno,' which used AI to generate political content including anti-Hamas, anti-Qatar, pro-Israel, anti-BJP, and pro-Histadrut messaging. The operation represents a case of AI-assisted influence activity disrupted by a major AI provider. This is part of OpenAI's ongoing effort to detect and disrupt malicious uses of its platform.
OpenAI terminated accounts associated with IUVM, an Iran-origin influence operation that used AI tools to generate and translate pro-Iran, anti-Israel, and anti-US website content. The action is part of OpenAI's ongoing effort to disrupt malicious uses of its platform. This case documents a concrete instance of AI-assisted state-linked information operations being detected and disrupted.
OpenAI identified and banned accounts linked to the Iran-affiliated threat actor STORM-0817, which was using OpenAI's models to debug Android malware, scrape social media platforms, and translate offensive tooling. The disclosure is part of OpenAI's ongoing threat intelligence reporting on state-linked misuse of AI systems. This case illustrates concrete adversarial use of frontier AI for cyberoffense and surveillance tooling.
OpenAI identified and banned accounts likely belonging to SweetSpecter, a suspected China-based threat actor, after detecting use of AI tools to research vulnerabilities, write malicious code, and support spear-phishing campaigns. This represents a concrete case of a nation-state-linked adversary operationalizing frontier AI for offensive cyber operations. OpenAI's disclosure is part of its ongoing effort to disrupt malicious uses of its platform.
OpenAI identified and banned accounts associated with CyberAv3ngers, an Iran-linked threat actor, that were using OpenAI's models to research industrial control systems, default credentials, and potential targets. The action is part of OpenAI's ongoing effort to disrupt malicious uses of its AI platform. This case is notable as an example of state-linked actors leveraging frontier AI for critical infrastructure reconnaissance.
OpenAI banned a likely US-based account that used AI to generate a fake ChatGPT error message falsely claiming to detect 'Russian troll' activity. The fabricated screenshot was designed to spread disinformation by impersonating OpenAI's platform. This is part of OpenAI's ongoing effort to disrupt malicious uses of its AI systems.
OpenAI banned a cluster of accounts linked to a Russia-origin influence operation dubbed 'Stop News,' which used AI tools to generate multilingual articles and social media posts targeting audiences in Ukraine and Western countries. The operation represents a documented case of AI being weaponized for coordinated inauthentic behavior at scale. OpenAI's disclosure is part of its ongoing effort to detect and disrupt malicious uses of its platform.
OpenAI disrupted and banned a cluster of accounts, dubbed 'A2Z', that were using AI to generate multilingual influence content targeting elections, the Ukraine conflict, and political topics across multiple platforms. The operation is part of OpenAI's ongoing effort to detect and remove coordinated inauthentic behavior from its services. This is a concrete case study in AI-enabled influence operations and platform enforcement.
OpenAI identified and banned a cluster of accounts linked to an Iran-origin influence operation designated STORM-2035, which used AI to generate election-related content targeting US and UK audiences. The content was distributed across multiple sites as part of a coordinated influence campaign. This is a concrete case of AI-enabled information operations being disrupted by a frontier lab.
OpenAI banned accounts that used its models via an Israel-based startup to generate conversations and distribute links to gambling sites on X (formerly Twitter). The operation represents a coordinated misuse of AI-generated content for spam and influence purposes. The disclosure is part of OpenAI's ongoing series of reports on disrupting malicious uses of its models.
OpenAI identified and banned accounts originating from Rwanda that were using AI to generate partisan political comments ahead of the country's elections. The action is part of OpenAI's ongoing effort to disrupt malicious uses of its AI systems. This represents a concrete enforcement case of AI-enabled influence operations targeting an electoral context.
OpenAI identified and banned accounts that were using its AI systems to generate abusive reports and complaints targeting Vietnamese public figures and platforms. The action is documented in a threat report focused on malicious use of AI for coordinated harassment or censorship campaigns. This represents a concrete enforcement case of AI misuse for politically motivated targeting.
OpenAI identified and banned accounts that were using its AI systems to generate comments criticizing a Russian anti-corruption foundation and associated figures. The operation represents a documented case of AI-enabled influence operations being disrupted by a frontier lab. This is relevant to AI misuse tracking and platform safety enforcement.
OpenAI banned accounts likely originating from China that were using AI to generate English-language social media posts and Spanish-language articles as part of an influence operation dubbed 'Sponsored Discontent.' The operation represents a documented case of AI-assisted information operations being detected and disrupted by a frontier lab. This is notable as a concrete example of AI misuse at scale and OpenAI's active role in detecting and countering such activity.
OpenAI identified and banned accounts apparently originating in Cambodia that were using its AI to translate and generate conversations for romance-baiting and investment scam (pig butchering) workflows. The action represents a concrete enforcement case of AI being weaponized for large-scale fraud operations. This is a notable safety/misuse incident from a tier-1 source.
OpenAI identified and banned accounts linked to Iranian influence networks IUVM and STORM-2035 that were using AI to generate articles and social media posts as part of coordinated influence operations. The action represents a cross-platform disruption effort targeting state-linked misuse of AI-generated content. This is a concrete example of AI being weaponized for information operations and of a major lab taking enforcement action against such abuse.
Datasette
llm-mcp-client
Simon Willison
OpenAI
OpenAI
OpenAI
OpenAI
any topic Agent and Tool Ecosystem AI Safety Research Alignment and RLHF Enterprise Deployment Patterns Evaluation and Benchmarking Frontier Model Releases Inference Economics Long Context Evolution Multimodal Progress Open Weights Progress Regulatory Developments Training Infrastructure
any topic Agent and Tool Ecosystem AI Safety Research Alignment and RLHF Enterprise Deployment Patterns Evaluation and Benchmarking Frontier Model Releases Inference Economics Long Context Evolution Multimodal Progress Open Weights Progress Regulatory Developments Training Infrastructure
any entity 1-Lipschitz Neural Networks on Hadamard Manifolds 1.58-bit quantization 10 USC 3252 10x Genomics 12-Factor App 1B-scale language models 1M-token context 1X 2-Parameter Logistic IRT Model 2026 FIFA World Cup 2026 State of AI Traffic and Cyberthreat Benchmark Report 2D-RoPE 2WikiMultiHopQA 32B language model 3C3H 3D asset generation 3D Gaussian Splatting 3D Medical Image Classification 3D optical flow 3D point cloud generation 3D tracking 3D-Aware VLMs with Implicit and Explicit Geometries 3D-Fit 3D-RoPE 3E Memory Module 3LM 404 Media 4chan 4D reconstruction 5PILS 6-PACK 6G Radio Access Network 7B language model 8B autoregressive language model A Blueprint for Equilibrium-Based Differentiable Continuous-Variable Thermodynamic Computing A Causal Model of Theory of Mind in Conflict for Artificial Intelligence A Comparative Study of Deep Learning Architectures for Multi-Horizon Behavioural Forecasting for Mobile Health A Computational Audit of Demographic Association Encoding in ClinicalBERT Language Predictions A Constrained Natural-Language Interface for Variational Multi-Physics Finite Element Simulations in FEniCS A Continual Validation, Updating, and Decision-Making Framework for Self-Adaptive Digital Twins via Robust Model Predictive Control A Diffusion Approximation for Temporal-Difference Learning with Linear Features under Markovian Noise A Factorial Study of Synthetic Data Generation for Low-Resource Machine Translation using Grammar Books A History-Aware Visually Grounded Critic for Computer Use Agents A Human-in-the-Loop Corpus for LLM-Based Simplification of Scientific Summaries A Komi-Yazva–Russian Parallel Corpus and Evaluation Protocol for Zero- and Few-Shot LLM Translation A Learning-Rate-Gated Failure of GRPO in a Small Language and Vision-Language Model Web Agent A Methodology for Auditable Trustworthiness Levels in AI Lifecycle Governance A Multi-Agent System for Autonomous, Fine-Tuning-Free Clinical Symptom Detection: Development and Validation Study A Multi-Domain Benchmark for Detecting AI-Generated Text-Rich Images from GPT-Image-2 A Physics-Informed Fourier-Wavelet Transformer for Multiscale Computational Fluid Dynamics Surrogate Modeling A Practical Investigation of Training-free Relaxed Speculative Decoding A Resource for Enthymeme Detection in Controversial Political Discourse A Self-Evolving Agent for Longitudinal Personal Health Management A sleep-like consolidation mechanism for LLMs A Study of Temporal Fusion Strategies for Named Entity Recognition in Historical Texts A Systematic Approach for Selecting Trajectories for Data Augmentation A Taxonomy of Conceptual Alignment in Human-Robot Dialogue A Training-Free Mixture-of-Agents Framework for Multi-Document Summarization using LLMs and Knowledge Graphs A Tree-of-Thoughts Inspired Hybrid Approach for Legal Case Judgement Summarization using LLMs A Unifying Lens on Supervised Fine-Tuning Through Target Distribution Design A24 A2A Protocol A2C A2Z A3C AA-Briefcase AA-Omniscience AAAI Aardvark AARRI-Bench AASIST AB-UPT Abanca ABBEL ABC-Bench Abeba Birhane abhigyanpatwari abjadai Ableton abliteration Abridge Abstraction Gap Accelerate Accelerated Decentralized Stochastic Gradient Descent for Strongly Convex Optimization Accelerating Masked Diffusion Large Language Models: A Survey of Efficient Inference Techniques Accenture Accenture Anthropic Business Group accumulated message effect Accuracy and Satisfaction in Multi-Turn LLM Dialogues for NFR Assessment ACE ACEBench-Agent AceReason-14B Achieving Precise Text-To-Cypher Via Grounded Knowledge Graph Data Generation ACKTR ACL ACL Anthology ACL-Verbatim Acoustic Cue Alignment in Audio Language Models for Speech Emotion Recognition ACROS Act2Answer Action Chunking Transformer ACTION-BED Action-BED: Task-Driven Bayesian Experimental Design with Singly Intractable Objectives Action-Dependent Factorized Baselines Actionable Activation Directions for Detecting and Mitigating Emergent Misalignment Across Language Model Families Activation Atlases activation capping Activation Oracles activation patching Activation Steering active learning Active Listening Active Offline-to-Online Reinforcement Learning ActiveSAM ActiveVision ACTS Ada AdaCodec AdaFlash AdaGrad AdaJEPA AdaJEPA: An Adaptive Latent World Model Adalat AI AdaLook Adam AdamO AdamW AdaPrefix-GRPO ADAPT-GQE ADAPT-VQE adapter fine-tuning Adaptive Clip Policy Optimization Adaptive Data Scheduling Adaptive Depth Sparse Framework Adaptive Gated Feedback Optimization Adaptive Multi-Resolution Procedural Knowledge Compression for Large Language Models Adaptive Multi-Step Lookahead Decoding for Diffusion Language Models Adaptive Parallel Reasoning Adaptive Self-Debiasing adaptive thinking Adaptive Turn-Taking for Real-time Multi-Party Voice Agents adaptive-rank instantiation ADAS AdaSR ADEME ADHD Adithya S K Aditi Krishnapriyan Aditya (co-lead author) Aditya G. Parameswaran Adjacent Contrastive Reasoning ADK Python adk-samples Adobe Adobe Creative Cloud Advanced AI Scaling Framework Advanced Light Source Advanced Matrix Extensions (AMX) Advanced Photon Source Advanced Voice Mode AdvancedMathBench Advancing Omnimodal Embodied Agents from Isolated Skills to Everyday Physical Autonomy AdvBench AdversaBench Adversarial Attacks on Neural Network Policies Adversarial Creation and Detection of AI-Generated Social Bot Content adversarial examples adversarial pragmatics Adversarial Pragmatics for AI Safety Evaluation: A Benchmark for Instruction Conflict, Embedded Commands, and Policy Ambiguity adversarial refinement adversarial robustness adversarial training AdvGRPO Aeneas Aerial Image Dataset (AID) Affinity by Canva AFM 3 AFM 3 Cloud Pro AFM 3 Core Advanced AfriSenti AG-MG Parallel Corpus AG-UI AG-UI Protocol AGC-Bench AGC-Judge Agda AGDO Age of Empires II age prediction Agence France-Presse Agent 365 agent cloud Agent Cognitive Redundancy Ratio Agent Communication Protocol Agent Development Kit agent ecologies Agent Governance Toolkit agent harness Agent JIT Compilation Agent Reasoning Evaluation (ARE) Agent Runtime agent sandboxing agent-device agent-native Agent-Native Immune System Agent-Native Immune System: Architecture, Taxonomy, and Engineering agent-native telemetry Agent-Reach Agent-S agent-skills agent-task efficiency agent-teams-ai agent-to-agent evaluation protocol Agent-to-Agent Protocol (A2A) agent-toolkit-for-aws AgentBeats AgentBeats: Agentifying Agent Assessment for Openness, Standardization, and Reproducibility AgentBoard AgentCL AgentDojo Agentforce agentic AI Agentic AI Foundation Agentic AI Pipelines Agentic AI Systems Agentic Chain-of-Thought Steering Agentic Chain-of-Thought Steering for Efficient and Controllable LLM Reasoning Agentic CLEAR agentic coding Agentic Commerce Protocol Agentic Environment Engineering for Large Language Models: A Survey of Environment Modeling, Synthesis, Evaluation, and Application Agentic generation of verifiable rules for deterministic, self-expanding reaction classification Agentic Permissions Policy Algebra for Taint Confinement in LLM Agents agentic proving Agentic RAG agentic re-optimization framework Agentic Resource Discovery Agentic RL Agentic Self-Instruct Agentic System Monitoring Methodology Agentic Technical Debt agentic vulnerabilities AgenticRL AgenticSTS AgentKit AgentMap agentmemory AgentMob Agentopia Agentopia: Long-Term Life Simulation and Learning in Agent Societies Agents in the Wild: Where Research Meets Deployment Agents-A1 agents-cli Agents-K1 Agents.js AGENTS.md Agents' Last Exam AgentScope AGENTSERVESIM agentskills.io AgentSnare AgentSpec AgentsView AgentWorldBench AGI AGI (Artificial General Intelligence) AGI cognitive framework AGI economy Agility Robotics Agon AGORA agreement attraction AHA-WAM Ahead-of-Time Compilation Ahmad Osman AI Agents AI agents that matter AI alignment AI and Compute AI Andrew AI as Normal Technology AI biosecurity risk assessment AI Clinical Copilot AI Co-Clinician AI control AI Control Roadmap AI Cybersecurity Threat Evaluation Framework AI data sovereignty AI Developer Conference 2025 AI Engineer NYC AI Engineer World's Fair AI Existential Risk AI Exposure Scores: what they measure, what they miss, and what comes next AI for Citizens AI for Game Development AI for Math Initiative AI for Science AI for Science program AI Fund AI Healthcare Chatbots as Information Infrastructure: A Large-Scale Study of User-Reported Breakdowns AI image verification AI inference carbon emissions AI leaderboards AI Liability Directive AI Literacy & Creator Collective AI Mathematical Olympiad AI Overview AI Panic Blog AI Persuasive Framing in Collective Dilemmas AI Registry AI Reproducibility Benchmark AI Risk Management Framework AI Safety Fund AI Safety Level (ASL) AI Safety Level Standards AI Safety via Debate AI scaling laws AI Sheets AI Snake Oil AI systems and the reproduction of (standard) language ideologies in World Englishes AI translation of literary texts is "fine", but readers still prefer human translations AI Verify Foundation AI vs. AI AI Watermark Evidence Fails Forensic Readiness: An Empirical Evaluation AI-accelerated End-to-End Framework for Rapid Professional Upskilling ai-agent-book AI-assisted human evaluation AI-assisted red teaming AI-Assisted Systematization for Evaluating GenAI Systems AI-assisted theorem proving ai-boost/awesome-harness-engineering AI-driven constraint reasoning AI-Driven Life Cycle (AI-DLC) AI-Engineering-Coach ai-hedge-fund AI-MO AI-PAVE-Br AI-Trader AI's Capability in Assisting Scientific Research in Physics, Astrophysics, and Cosmology II: Project Planning and Proposal Evaluation AI@HHMI AI21 Labs AIA Labs AIcrowd Aider AIDev dataset aidlc-workflows Aiera AIEWF AILuminate AIME AIME 2025 AIME 2026 AIME2024 AIME24 AIME25 AIME26 aiming-lab AIMO Interpretability Challenge AIMO Progress Prize AIMS AINews AionUi Air India AIR: Adaptive Interleaved Reasoning with Code in MLLMs Aira Explorer AiraXiv Airbus AirflowAttack AIriskEval-edu Demo AISPA AISPA: User-Centric System Prompt Auditing for Large Language Model Applications aisuite Akshat Bubna Akshay Nathan Albert Gu Alberta AI Academy Alberta Machine Intelligence Institute Alberti ALE-Bench ALER-TI Aletheia Alex Hernandez Alex Kearns-Apuya Alex L. Zhang Alex Lupsasca Alex Rives Alexa+ Alexandr Wang AlexNet ALFRED Alfred Lin ALFWorld Algonauts 2025 Algorithmic and Minimax Complexities in Kernel Bandits Algorithmic Impermeability algorithmic monoculture Algorithmic Monocultures in Hiring Ali Essam Ghareeb Alibaba Alibaba Cloud Alibaba Cloud Model Studio Alibaba DAMO Academy Alibaba Qwen Alibaba Qwen Team ALIGN AlignAtt AlignAtt4LLM ALIGNBEAM alignment auditing alignment faking Alignment Research Center alignment tampering alignment tax Alireza Rezvani All Hands AI All Smoke, No Alarm: Oracle Signals in Agent-Authored Test Code All Tech Is Human all-edge F1 Allam Allegro Allegro Hand Allen Institute Allen Institute for AI AllenAI Allied for Startups ALMANAC ALOHA AlpacaEval 2 Alpamayo R1 Alphabet AlphaEarth Foundations AlphaEvolve AlphaFold AlphaFold2 AlphaFold3 AlphaGenome AlphaGo AlphaOracle AlphaStar AlphaTensor alpic-ai Altana Altera Alternating Token-Weighted Unlearning Altimeter Capital Always-On Evaluation Protocol (AOEP-v0) Always-OnAgents: A Survey of Persistent Memory, State, and Governance in LLM Agents ALX Alyah AM-Parser AMALIA Amar Subramanya AMARIS Amazon Amazon Beauty Amazon Bedrock Amazon Bedrock AgentCore Amazon Graviton Amazon Reviews Dataset Amazon SageMaker Amazon SageMaker Studio Amazon Trainium2 Amazon Web Services Ambient Diffusion Policy AMC12 AMD AMD Instinct AMD Instinct MI300 AMD MI300 AMD MI355X AMEL American Federation of Teachers American Technology Council Ami Vora AMIA Amir Bar amortized variational inference AMP Amrith Setlur AMRS aMUSEd An Agency-Transferring Model-Free Policy Enhancement Technique An Empirical Analysis of Factual Errors in Human-Written Text and its Application An Evaluation Framework for Structured Audio Captions Validated by Controlled Perturbations An Exact Instrument for State Usage in Selective State-Space Models, and the Input-Driven Migration It Reveals Analogical Deep Research Analogical Deep Research: Retrieving and Integrating Historical Analogies for Foresight Analysis Analysing Self-Harm Representations in Language Models: a Cross-Architecture Study AnalyticGeo7K Analytics-Everywhere-Lab Anatomy of Post-Training: Using Interpretability to Characterize Data and Shape the Learning Signal ANCESTRA anchoring effect Andon Labs Andreessen Horowitz Andrej Karpathy Andrew Kelley Andrew Ng Andrew Qu Andrew Wagenmaker Anduril Industries Andy Beam Andy Jassy Andyyyy64 Angelini Pharma AngelSpec Anglophone retrieval bias Angluin's condition anisotropic BAO damping Anjney Midha Anna's Archive Annapurna Labs annotator disagreement anomalyco AnomalyShapeNet ANSI National Accreditation Board AnswerSheet ANSYR ANTAP Anthology Fund Anthony Albanese Anthropic Anthropic Academy Anthropic Advanced AI Framework Anthropic Agent SDK Anthropic Beneficial Deployments Anthropic Claude Code Anthropic Economic Advisory Council Anthropic Economic Futures Programme Anthropic Economic Futures Research Fund Anthropic Economic Index Anthropic Economic Policy Framework Anthropic Fable 5 Anthropic for Startups Anthropic Interviewer Anthropic Korea Anthropic Labs Anthropic Long-Term Benefit Trust Anthropic Mythos Anthropic National Security and Public Sector Advisory Council Anthropic Partner Academy Anthropic Policy Frontier Red Team Anthropic Public Record Anthropic Responsible Scaling Policy Anthropic Safeguards Team Anthropic Terms of Service Anthropic Threat Intelligence Report August 2025 Anthropic Usage Policy Anthropic v. Department of War anthropics/skills anthropomorphism in AI Anti-Drift Rectification Anti-Periodic Positional Encoding: Möbius Boundary Conditions Make In-Context Retrieval Reliable Antigravity Antoine Zambelli Anton Korinek Any-Dimensional Learning by Sampling AnyGroundBench AnyLanguageModel AnyMo Anysphere Anytime-Valid E-Process anytime-valid sequential testing AP2 Apache 2.0 Apache Arrow Apache Ossie Apache Software Foundation Apache Spark Apart Research APE Apertus 70B Apertus-8B-Instruct-2509 APEX-Accounting APEX-Accounting APEX-Agents-AA aphasia API knowledge boundary probing APO: Unsupervised Atomic Policy Optimization for 3D Structure Prediction of Atomic Systems Apollo Apollo Global Management Apollo Research APPA (Agentic Permissions Policy Algebra) Apparent Psychological Profiles of Large Language Models are Largely a Measurement Artifact Appearance Pointers -- Multimodal Region Control of Diffusion Transformers Appia Foundation Apple Apple Foundation Models 3 Apple Foundation Models Framework Apple Health Apple Neural Engine Apple Silicon Applied Intuition apply_patch APPO: Agentic Procedural Policy Optimization Approximate DP APPS Apps SDK AppWorld Apriel-H1 AprielGuard APS-Bench APS-RAG apurvsinghgautam AR military headset Ar9av AraBERT Arabic Arabic Instruction Following Eval (IFEval) Arabic LLMs ArabiGEE AraGen ARC Evals ARC Framework Arc Institute ARC Prize Foundation Arc Virtual Cell Challenge ARC-AGI Archon Arco Education ARDY Are Performance-Optimization Benchmarks Reliably Measuring Coding Agents? Are We Ready For An Agent-Native Memory System? AREA Arena AI Arena Code Arena Leaderboard Arena Search Arena-Hard Arena.ai Code Arena WebDev ArenaHard Argentina Argilla Argilla 2.0 Argonne National Laboratory Argus ariadng ARIS Aristotle API Arithmetic Pedagogy for Language Models Arize AI Arize Phoenix Arm Arm AGI CPU Armin Ronacher ArMIS Army Special Operations Command Arnuv Tandon ArogyaBodha ArogyaSutra ARPA-H ARROW ART (Agent Reinforcement Trainer) Arthur Mensch Artifact Artifacts Artificial Analysis Artificial Analysis Big Bench Audio Artificial Analysis Coding Agent Index Artificial Analysis Conversational Dynamics Artificial Analysis Intelligence Index Artificial Analysis Intelligence Leaderboard Artificial Analysis LLM Performance Leaderboard Artificial Analysis Text to Image Artificial Analysis Text to Image Leaderboard Artificial Bee Colony Artificial Epanorthosis: Why large language models overuse a classical rhetorical figure, and how to mitigate it ArXiv arXiv:2602.05394 ASAM ASAP 2.0 ASAP++ Ascend SuperPOD Ashwyn Sharma Asia-Pacific ASIMOV-2.0 AskChem AskChem-Bench AskDocs ASL-2 ASL-3 ASML ASML Holding NV Aspect-Based Sentiment Analysis ASRD AssemblyAI AssemblyAI Universal Assessing the Operational Impact of Poisoning Attacks over Augmented 3D Point Cloud Public Datasets for Connected and Autonomous Vehicles AssetOpsBench assistant axis Assistants API Assisted Generation ASSISTments Association for Computational Linguistics Associative Recurrent Memory Transformer Astera Institute ASTRA Astral AstrBot AstrBotDevs ASVspoof 5 asynchronous inference Asynchronous Noise Schedule Atari ATDev ATE-Bench Atla Atlas Atlas H&E-TME ATLAS Open Data 13 TeV ATLAS: Active Theory Learning for Automated Science Atlassian AToken Atomic Policy Optimization ATOMIC2020 Atoms of Thought (EEG Microstates paper) Attack Success Rate Attainable Utility Preservation Attention Amnesia in Hybrid LLMs: When CoT Fine-Tuning Breaks Long-Range Recall, and How to Fix It attention entropy Attention Expansion: Enhancing Keyphrase Extraction from Long Documents with Attention-Augmented Contextualized Embeddings attention head circuit Attention Residuals Attention Sorting Attention submodule Attention-based MIL Attractor States Emerge in Multi-Turn LLM Conversations AUC Audex Audio Interaction Model Audio MultiChallenge Audio-Flamingo-3 Audio-LLM Audio-Native Speech Recognition with a Frozen Discrete-Diffusion Language Model AudioCards AudioDER AudioLDM 2 AudioVAE2 Auditing Protocol-Level Shortcuts in Large Audio Language Model Judges for Speech Evaluation AUDITS Augment Augment Code AUPRC AuRA AUROC Australia AI Safety Institute Australia National AI Plan Australian Emergency Department triage notes Australian Government Australian National University Auto Benchmark Audit (ABA) Autodata Autodesk Fusion AutoDex autoDream AutoForest Autoformer AutoGPTQ AutoJourn AUTokenizer AutoLab automated AI research Automated Background Swapping Automated Compliance Mapping in Cloud Security with Domain-Adapted Sentence Transformers Automated Discovery Has No Universally Superior Harness Automated Extraction of Techno-Economic Data from 76,000 Energy System Studies Automated grading of Linux/bash examinations using large language models: a four-level cognitive taxonomy approach automated mechanistic interpretability automated red teaming Automated Reference Verification System Automated reproducibility assessments in the social and behavioral sciences using large language models automated test suite automated theorem proving Automatic Domain Randomization Automatic Post-Editing (APE) Automatic Speech Recognition AutomationBench-AA AutoMem autonomous driving AUTOPILOT-VQA Autoregressive Boltzmann Generators autoregressive transformer autoresearch AutoResearchClaw AutoRound AutoSkillHarm AutoSynthesis avatarin Avenger Labs Aviral Kumar awesome-claude-code awesome-copilot awesome-llm-apps AWQ AWS AWS Bedrock AWS GovCloud AWS Inferentia2 AWS Labs AWS Marketplace AWS Neuron AWS Neuron SDK AWS Trainium AxDafny Axel Backlund Axiom Math Axios Axolotl AXPO Aya Aya Expanse Aya Vision Aya-Expanse-8B-Base Azure Azure AI Studio Azure DevOps Azure Foundry Azure OCR B³D-RWKV BabelJudge BABI-Long BABILong BabyCL Backbone-as-Architect backdoor attack Backdoor Circuit Analysis (Language-Switching) Backdoor Unlearning Generalization: A Path Toward the Removal of Unknown Triggers in LLMs backnotprop backpropagation through time Backstory Bad Grammar Bad Scientist BadWAM BadWAM: When World-Action Models Dream Right but Act Wrong Bag of Words (BoW) BAGEL BAGEL-7B Baidu Baidu Translate Baifeng Shi Baiyu Chen ball balancing ball balancing task Bamba BamiBERT Bandit Bang-v3 BanglaBERT Banque des Territoires Barcelona Supercomputing Center Language Technologies bare metal sandboxes Bargaining Scenarios Dataset Bark Base Baseline Agent Baseline-Log Physical Separation Baseten BashArena basic-memory basicmachines-co Basque Center on Cognition, Brain, and Language Basque language baulab Bayesian decision theory Bayesian Inference and Decision Audits for Public Archives of Frontier AI Evaluations Bayesian Multiobjective Optimization Bayesian Optimal Experimental Design Bayesian Optimization BayesPO BayLing-Duplex BBC News BBQ BBVA BC Protocol BCG X BCR BDD100K Be My Eyes Be My Tutor: On-Policy Co-Distillation for Mutual LLM Improvement via Peer Feedback BEA-Dialogue Beam Search Bebop Bechiri and Lanasri [2026] Becker Friedman Institute for Economics Before You Think: System 0, AI-Mediated Cognition and Cognitive Colonization behavior cloning behavior trees with learning-enabled components behavioral fine-tuning Behavioral Trajectory Tracking Framework behavioral-gradient validator Behavioral-SafetyBench Behavox BEIR Belebele Belief-Space Safety Filter (BeliefSF) BeliefTrack Ben Bernanke Ben Mann Ben Parr Ben Thompson Benchling Benchmark Agent Benchmark Everything Everywhere All at Once Benchmarking Multimodal Large Language Models for Scientific Visualization Literacy benchmaxxing BenCzechMark Bending Spoons Beneficial Deployments Benevity BenevolentAI Bengali Bengaluru BenHalluEval BenHalluScore Benjamin Chang Benjamin Krause Berkeley AI Research (BAIR) Berkeley Artificial Intelligence Research Berkshire Hathaway Berlin Database of Emotional Speech (EMO-DB) BerriAI BERT BERT-base BERT-F1 BERTopic BERTScore BERTurk BERTweet BES (trained models) Bessemer Venture Partners Best-of-N Best-of-N Sampling best@k Better Call GRPO Beyond a Single Judge: Simulating Social Persona Panels for Generative UI Evaluation Beyond Accuracy: Community Perspectives on Machine Translation Beyond Adam: SOAP and Muon for Faster, Label-Efficient Training of Machine Learning Interatomic Potentials Beyond Aggregate Risk: Role-Stratified Conformal Risk Control for LLM Tool Calls Beyond Attack-Success Rate: Action-Graded Severity Scale for Tool-Using AI Agents Beyond Benchmarks: Exposing the Hidden Crisis in Bangla Hate Speech Detection Beyond Fixed Representations: The Vocabulary and Verifier Gaps in Open-Ended AI Beyond Fully Random Masking: Attention-Guided Denoising and Optimization for Diffusion Language Models Beyond Function Calling: Benchmarking Tool-Using Agents under Tool-Environment Unreliability Beyond Global Replanning: Hierarchical Recovery for Cross-Device Agent Systems Beyond Independent Labels: Schwartz-Geometry Decoding for Human Value Detection Beyond Negative-Ridge Endpoints: Mixed-Sign Spectral Regularization via Negative-Shifted Gradient Descent Beyond Runtime Enforcement: Shield Synthesis as Defensibility Analysis for Adversarial Networks Beyond Scale and Generation: Understanding Language Model-based Entity Matching Beyond Sentiment: Structured Information Extraction from Financial News Beyond Success Rate: Cost-Aware Evaluation of Offensive and Defensive Security Agents Beyond Sufficiency: Time Series Explanation with Counterfactual Necessity Beyond Surface Forms: A Comprehensive, Mechanism-Oriented Taxonomy of Indirect Linguistic Encoding for LLM-Based Coded Language Detection Beyond Sycophancy: Structured Resistance and Compliance in LLM Moral Reasoning Beyond task performance: Decoding bioacoustic embeddings with speech features Beyond the Commitment Boundary: Probing Epiphenomenal Chain-of-Thought in Large Reasoning Models Beyond the Hard Budget: Sparsity Regularizers for More Interpretable Top-k Sparse Autoencoders Beyond the Leaderboard: Design Lessons for Trustworthy Multimodal VQA Beyond Third-Person Audits: Situated Interaction Auditing for User-Centered LLM Bias Research Beyond Tokenization: Direct Timestep Embedding and Contrastive Alignment for Time-Series Question Answering Beyond Uniform Tokens: Adaptive Compression for Time Series Language Models BeyondUncertainty BFCL BFCL Multi-Turn BFCL-V3 BFCLv3 BFW schedule hint BGE BGE-M3 BGL bi-manual robotic manipulation Bias Benchmark for Question Answering Bias Leaves a Gradient Trail Biconvex Optimization bidirectional attention Bidirectional Evolutionary Search Bielik Big Bench Big Bench Audio Big-Math BigBird BigCode BigCodeArena BigCodeBench Bigger or Cheaper? Scale and Quantization Effects on Uncertainty Signals in Vision-Language Models Under Image Degradation BigLaw Bench BigScience Bill & Melinda Gates Foundation Bill Dally Bill Gurley Bill McDermott Binarized Neural Network binary quantization BINEVAL Bing binning semiring bioacoustic monitoring bioacoustics BioASQ 14b BioBERT Biographical
any entity 1-Lipschitz Neural Networks on Hadamard Manifolds 1.58-bit quantization 10 USC 3252 10x Genomics 12-Factor App 1B-scale language models 1M-token context 1X 2-Parameter Logistic IRT Model 2026 FIFA World Cup 2026 State of AI Traffic and Cyberthreat Benchmark Report 2D-RoPE 2WikiMultiHopQA 32B language model 3C3H 3D asset generation 3D Gaussian Splatting 3D Medical Image Classification 3D optical flow 3D point cloud generation 3D tracking 3D-Aware VLMs with Implicit and Explicit Geometries 3D-Fit 3D-RoPE 3E Memory Module 3LM 404 Media 4chan 4D reconstruction 5PILS 6-PACK 6G Radio Access Network 7B language model 8B autoregressive language model A Blueprint for Equilibrium-Based Differentiable Continuous-Variable Thermodynamic Computing A Causal Model of Theory of Mind in Conflict for Artificial Intelligence A Comparative Study of Deep Learning Architectures for Multi-Horizon Behavioural Forecasting for Mobile Health A Computational Audit of Demographic Association Encoding in ClinicalBERT Language Predictions A Constrained Natural-Language Interface for Variational Multi-Physics Finite Element Simulations in FEniCS A Continual Validation, Updating, and Decision-Making Framework for Self-Adaptive Digital Twins via Robust Model Predictive Control A Diffusion Approximation for Temporal-Difference Learning with Linear Features under Markovian Noise A Factorial Study of Synthetic Data Generation for Low-Resource Machine Translation using Grammar Books A History-Aware Visually Grounded Critic for Computer Use Agents A Human-in-the-Loop Corpus for LLM-Based Simplification of Scientific Summaries A Komi-Yazva–Russian Parallel Corpus and Evaluation Protocol for Zero- and Few-Shot LLM Translation A Learning-Rate-Gated Failure of GRPO in a Small Language and Vision-Language Model Web Agent A Methodology for Auditable Trustworthiness Levels in AI Lifecycle Governance A Multi-Agent System for Autonomous, Fine-Tuning-Free Clinical Symptom Detection: Development and Validation Study A Multi-Domain Benchmark for Detecting AI-Generated Text-Rich Images from GPT-Image-2 A Physics-Informed Fourier-Wavelet Transformer for Multiscale Computational Fluid Dynamics Surrogate Modeling A Practical Investigation of Training-free Relaxed Speculative Decoding A Resource for Enthymeme Detection in Controversial Political Discourse A Self-Evolving Agent for Longitudinal Personal Health Management A sleep-like consolidation mechanism for LLMs A Study of Temporal Fusion Strategies for Named Entity Recognition in Historical Texts A Systematic Approach for Selecting Trajectories for Data Augmentation A Taxonomy of Conceptual Alignment in Human-Robot Dialogue A Training-Free Mixture-of-Agents Framework for Multi-Document Summarization using LLMs and Knowledge Graphs A Tree-of-Thoughts Inspired Hybrid Approach for Legal Case Judgement Summarization using LLMs A Unifying Lens on Supervised Fine-Tuning Through Target Distribution Design A24 A2A Protocol A2C A2Z A3C AA-Briefcase AA-Omniscience AAAI Aardvark AARRI-Bench AASIST AB-UPT Abanca ABBEL ABC-Bench Abeba Birhane abhigyanpatwari abjadai Ableton abliteration Abridge Abstraction Gap Accelerate Accelerated Decentralized Stochastic Gradient Descent for Strongly Convex Optimization Accelerating Masked Diffusion Large Language Models: A Survey of Efficient Inference Techniques Accenture Accenture Anthropic Business Group accumulated message effect Accuracy and Satisfaction in Multi-Turn LLM Dialogues for NFR Assessment ACE ACEBench-Agent AceReason-14B Achieving Precise Text-To-Cypher Via Grounded Knowledge Graph Data Generation ACKTR ACL ACL Anthology ACL-Verbatim Acoustic Cue Alignment in Audio Language Models for Speech Emotion Recognition ACROS Act2Answer Action Chunking Transformer ACTION-BED Action-BED: Task-Driven Bayesian Experimental Design with Singly Intractable Objectives Action-Dependent Factorized Baselines Actionable Activation Directions for Detecting and Mitigating Emergent Misalignment Across Language Model Families Activation Atlases activation capping Activation Oracles activation patching Activation Steering active learning Active Listening Active Offline-to-Online Reinforcement Learning ActiveSAM ActiveVision ACTS Ada AdaCodec AdaFlash AdaGrad AdaJEPA AdaJEPA: An Adaptive Latent World Model Adalat AI AdaLook Adam AdamO AdamW AdaPrefix-GRPO ADAPT-GQE ADAPT-VQE adapter fine-tuning Adaptive Clip Policy Optimization Adaptive Data Scheduling Adaptive Depth Sparse Framework Adaptive Gated Feedback Optimization Adaptive Multi-Resolution Procedural Knowledge Compression for Large Language Models Adaptive Multi-Step Lookahead Decoding for Diffusion Language Models Adaptive Parallel Reasoning Adaptive Self-Debiasing adaptive thinking Adaptive Turn-Taking for Real-time Multi-Party Voice Agents adaptive-rank instantiation ADAS AdaSR ADEME ADHD Adithya S K Aditi Krishnapriyan Aditya (co-lead author) Aditya G. Parameswaran Adjacent Contrastive Reasoning ADK Python adk-samples Adobe Adobe Creative Cloud Advanced AI Scaling Framework Advanced Light Source Advanced Matrix Extensions (AMX) Advanced Photon Source Advanced Voice Mode AdvancedMathBench Advancing Omnimodal Embodied Agents from Isolated Skills to Everyday Physical Autonomy AdvBench AdversaBench Adversarial Attacks on Neural Network Policies Adversarial Creation and Detection of AI-Generated Social Bot Content adversarial examples adversarial pragmatics Adversarial Pragmatics for AI Safety Evaluation: A Benchmark for Instruction Conflict, Embedded Commands, and Policy Ambiguity adversarial refinement adversarial robustness adversarial training AdvGRPO Aeneas Aerial Image Dataset (AID) Affinity by Canva AFM 3 AFM 3 Cloud Pro AFM 3 Core Advanced AfriSenti AG-MG Parallel Corpus AG-UI AG-UI Protocol AGC-Bench AGC-Judge Agda AGDO Age of Empires II age prediction Agence France-Presse Agent 365 agent cloud Agent Cognitive Redundancy Ratio Agent Communication Protocol Agent Development Kit agent ecologies Agent Governance Toolkit agent harness Agent JIT Compilation Agent Reasoning Evaluation (ARE) Agent Runtime agent sandboxing agent-device agent-native Agent-Native Immune System Agent-Native Immune System: Architecture, Taxonomy, and Engineering agent-native telemetry Agent-Reach Agent-S agent-skills agent-task efficiency agent-teams-ai agent-to-agent evaluation protocol Agent-to-Agent Protocol (A2A) agent-toolkit-for-aws AgentBeats AgentBeats: Agentifying Agent Assessment for Openness, Standardization, and Reproducibility AgentBoard AgentCL AgentDojo Agentforce agentic AI Agentic AI Foundation Agentic AI Pipelines Agentic AI Systems Agentic Chain-of-Thought Steering Agentic Chain-of-Thought Steering for Efficient and Controllable LLM Reasoning Agentic CLEAR agentic coding Agentic Commerce Protocol Agentic Environment Engineering for Large Language Models: A Survey of Environment Modeling, Synthesis, Evaluation, and Application Agentic generation of verifiable rules for deterministic, self-expanding reaction classification Agentic Permissions Policy Algebra for Taint Confinement in LLM Agents agentic proving Agentic RAG agentic re-optimization framework Agentic Resource Discovery Agentic RL Agentic Self-Instruct Agentic System Monitoring Methodology Agentic Technical Debt agentic vulnerabilities AgenticRL AgenticSTS AgentKit AgentMap agentmemory AgentMob Agentopia Agentopia: Long-Term Life Simulation and Learning in Agent Societies Agents in the Wild: Where Research Meets Deployment Agents-A1 agents-cli Agents-K1 Agents.js AGENTS.md Agents' Last Exam AgentScope AGENTSERVESIM agentskills.io AgentSnare AgentSpec AgentsView AgentWorldBench AGI AGI (Artificial General Intelligence) AGI cognitive framework AGI economy Agility Robotics Agon AGORA agreement attraction AHA-WAM Ahead-of-Time Compilation Ahmad Osman AI Agents AI agents that matter AI alignment AI and Compute AI Andrew AI as Normal Technology AI biosecurity risk assessment AI Clinical Copilot AI Co-Clinician AI control AI Control Roadmap AI Cybersecurity Threat Evaluation Framework AI data sovereignty AI Developer Conference 2025 AI Engineer NYC AI Engineer World's Fair AI Existential Risk AI Exposure Scores: what they measure, what they miss, and what comes next AI for Citizens AI for Game Development AI for Math Initiative AI for Science AI for Science program AI Fund AI Healthcare Chatbots as Information Infrastructure: A Large-Scale Study of User-Reported Breakdowns AI image verification AI inference carbon emissions AI leaderboards AI Liability Directive AI Literacy & Creator Collective AI Mathematical Olympiad AI Overview AI Panic Blog AI Persuasive Framing in Collective Dilemmas AI Registry AI Reproducibility Benchmark AI Risk Management Framework AI Safety Fund AI Safety Level (ASL) AI Safety Level Standards AI Safety via Debate AI scaling laws AI Sheets AI Snake Oil AI systems and the reproduction of (standard) language ideologies in World Englishes AI translation of literary texts is "fine", but readers still prefer human translations AI Verify Foundation AI vs. AI AI Watermark Evidence Fails Forensic Readiness: An Empirical Evaluation AI-accelerated End-to-End Framework for Rapid Professional Upskilling ai-agent-book AI-assisted human evaluation AI-assisted red teaming AI-Assisted Systematization for Evaluating GenAI Systems AI-assisted theorem proving ai-boost/awesome-harness-engineering AI-driven constraint reasoning AI-Driven Life Cycle (AI-DLC) AI-Engineering-Coach ai-hedge-fund AI-MO AI-PAVE-Br AI-Trader AI's Capability in Assisting Scientific Research in Physics, Astrophysics, and Cosmology II: Project Planning and Proposal Evaluation AI@HHMI AI21 Labs AIA Labs AIcrowd Aider AIDev dataset aidlc-workflows Aiera AIEWF AILuminate AIME AIME 2025 AIME 2026 AIME2024 AIME24 AIME25 AIME26 aiming-lab AIMO Interpretability Challenge AIMO Progress Prize AIMS AINews AionUi Air India AIR: Adaptive Interleaved Reasoning with Code in MLLMs Aira Explorer AiraXiv Airbus AirflowAttack AIriskEval-edu Demo AISPA AISPA: User-Centric System Prompt Auditing for Large Language Model Applications aisuite Akshat Bubna Akshay Nathan Albert Gu Alberta AI Academy Alberta Machine Intelligence Institute Alberti ALE-Bench ALER-TI Aletheia Alex Hernandez Alex Kearns-Apuya Alex L. Zhang Alex Lupsasca Alex Rives Alexa+ Alexandr Wang AlexNet ALFRED Alfred Lin ALFWorld Algonauts 2025 Algorithmic and Minimax Complexities in Kernel Bandits Algorithmic Impermeability algorithmic monoculture Algorithmic Monocultures in Hiring Ali Essam Ghareeb Alibaba Alibaba Cloud Alibaba Cloud Model Studio Alibaba DAMO Academy Alibaba Qwen Alibaba Qwen Team ALIGN AlignAtt AlignAtt4LLM ALIGNBEAM alignment auditing alignment faking Alignment Research Center alignment tampering alignment tax Alireza Rezvani All Hands AI All Smoke, No Alarm: Oracle Signals in Agent-Authored Test Code All Tech Is Human all-edge F1 Allam Allegro Allegro Hand Allen Institute Allen Institute for AI AllenAI Allied for Startups ALMANAC ALOHA AlpacaEval 2 Alpamayo R1 Alphabet AlphaEarth Foundations AlphaEvolve AlphaFold AlphaFold2 AlphaFold3 AlphaGenome AlphaGo AlphaOracle AlphaStar AlphaTensor alpic-ai Altana Altera Alternating Token-Weighted Unlearning Altimeter Capital Always-On Evaluation Protocol (AOEP-v0) Always-OnAgents: A Survey of Persistent Memory, State, and Governance in LLM Agents ALX Alyah AM-Parser AMALIA Amar Subramanya AMARIS Amazon Amazon Beauty Amazon Bedrock Amazon Bedrock AgentCore Amazon Graviton Amazon Reviews Dataset Amazon SageMaker Amazon SageMaker Studio Amazon Trainium2 Amazon Web Services Ambient Diffusion Policy AMC12 AMD AMD Instinct AMD Instinct MI300 AMD MI300 AMD MI355X AMEL American Federation of Teachers American Technology Council Ami Vora AMIA Amir Bar amortized variational inference AMP Amrith Setlur AMRS aMUSEd An Agency-Transferring Model-Free Policy Enhancement Technique An Empirical Analysis of Factual Errors in Human-Written Text and its Application An Evaluation Framework for Structured Audio Captions Validated by Controlled Perturbations An Exact Instrument for State Usage in Selective State-Space Models, and the Input-Driven Migration It Reveals Analogical Deep Research Analogical Deep Research: Retrieving and Integrating Historical Analogies for Foresight Analysis Analysing Self-Harm Representations in Language Models: a Cross-Architecture Study AnalyticGeo7K Analytics-Everywhere-Lab Anatomy of Post-Training: Using Interpretability to Characterize Data and Shape the Learning Signal ANCESTRA anchoring effect Andon Labs Andreessen Horowitz Andrej Karpathy Andrew Kelley Andrew Ng Andrew Qu Andrew Wagenmaker Anduril Industries Andy Beam Andy Jassy Andyyyy64 Angelini Pharma AngelSpec Anglophone retrieval bias Angluin's condition anisotropic BAO damping Anjney Midha Anna's Archive Annapurna Labs annotator disagreement anomalyco AnomalyShapeNet ANSI National Accreditation Board AnswerSheet ANSYR ANTAP Anthology Fund Anthony Albanese Anthropic Anthropic Academy Anthropic Advanced AI Framework Anthropic Agent SDK Anthropic Beneficial Deployments Anthropic Claude Code Anthropic Economic Advisory Council Anthropic Economic Futures Programme Anthropic Economic Futures Research Fund Anthropic Economic Index Anthropic Economic Policy Framework Anthropic Fable 5 Anthropic for Startups Anthropic Interviewer Anthropic Korea Anthropic Labs Anthropic Long-Term Benefit Trust Anthropic Mythos Anthropic National Security and Public Sector Advisory Council Anthropic Partner Academy Anthropic Policy Frontier Red Team Anthropic Public Record Anthropic Responsible Scaling Policy Anthropic Safeguards Team Anthropic Terms of Service Anthropic Threat Intelligence Report August 2025 Anthropic Usage Policy Anthropic v. Department of War anthropics/skills anthropomorphism in AI Anti-Drift Rectification Anti-Periodic Positional Encoding: Möbius Boundary Conditions Make In-Context Retrieval Reliable Antigravity Antoine Zambelli Anton Korinek Any-Dimensional Learning by Sampling AnyGroundBench AnyLanguageModel AnyMo Anysphere Anytime-Valid E-Process anytime-valid sequential testing AP2 Apache 2.0 Apache Arrow Apache Ossie Apache Software Foundation Apache Spark Apart Research APE Apertus 70B Apertus-8B-Instruct-2509 APEX-Accounting APEX-Accounting APEX-Agents-AA aphasia API knowledge boundary probing APO: Unsupervised Atomic Policy Optimization for 3D Structure Prediction of Atomic Systems Apollo Apollo Global Management Apollo Research APPA (Agentic Permissions Policy Algebra) Apparent Psychological Profiles of Large Language Models are Largely a Measurement Artifact Appearance Pointers -- Multimodal Region Control of Diffusion Transformers Appia Foundation Apple Apple Foundation Models 3 Apple Foundation Models Framework Apple Health Apple Neural Engine Apple Silicon Applied Intuition apply_patch APPO: Agentic Procedural Policy Optimization Approximate DP APPS Apps SDK AppWorld Apriel-H1 AprielGuard APS-Bench APS-RAG apurvsinghgautam AR military headset Ar9av AraBERT Arabic Arabic Instruction Following Eval (IFEval) Arabic LLMs ArabiGEE AraGen ARC Evals ARC Framework Arc Institute ARC Prize Foundation Arc Virtual Cell Challenge ARC-AGI Archon Arco Education ARDY Are Performance-Optimization Benchmarks Reliably Measuring Coding Agents? Are We Ready For An Agent-Native Memory System? AREA Arena AI Arena Code Arena Leaderboard Arena Search Arena-Hard Arena.ai Code Arena WebDev ArenaHard Argentina Argilla Argilla 2.0 Argonne National Laboratory Argus ariadng ARIS Aristotle API Arithmetic Pedagogy for Language Models Arize AI Arize Phoenix Arm Arm AGI CPU Armin Ronacher ArMIS Army Special Operations Command Arnuv Tandon ArogyaBodha ArogyaSutra ARPA-H ARROW ART (Agent Reinforcement Trainer) Arthur Mensch Artifact Artifacts Artificial Analysis Artificial Analysis Big Bench Audio Artificial Analysis Coding Agent Index Artificial Analysis Conversational Dynamics Artificial Analysis Intelligence Index Artificial Analysis Intelligence Leaderboard Artificial Analysis LLM Performance Leaderboard Artificial Analysis Text to Image Artificial Analysis Text to Image Leaderboard Artificial Bee Colony Artificial Epanorthosis: Why large language models overuse a classical rhetorical figure, and how to mitigate it ArXiv arXiv:2602.05394 ASAM ASAP 2.0 ASAP++ Ascend SuperPOD Ashwyn Sharma Asia-Pacific ASIMOV-2.0 AskChem AskChem-Bench AskDocs ASL-2 ASL-3 ASML ASML Holding NV Aspect-Based Sentiment Analysis ASRD AssemblyAI AssemblyAI Universal Assessing the Operational Impact of Poisoning Attacks over Augmented 3D Point Cloud Public Datasets for Connected and Autonomous Vehicles AssetOpsBench assistant axis Assistants API Assisted Generation ASSISTments Association for Computational Linguistics Associative Recurrent Memory Transformer Astera Institute ASTRA Astral AstrBot AstrBotDevs ASVspoof 5 asynchronous inference Asynchronous Noise Schedule Atari ATDev ATE-Bench Atla Atlas Atlas H&E-TME ATLAS Open Data 13 TeV ATLAS: Active Theory Learning for Automated Science Atlassian AToken Atomic Policy Optimization ATOMIC2020 Atoms of Thought (EEG Microstates paper) Attack Success Rate Attainable Utility Preservation Attention Amnesia in Hybrid LLMs: When CoT Fine-Tuning Breaks Long-Range Recall, and How to Fix It attention entropy Attention Expansion: Enhancing Keyphrase Extraction from Long Documents with Attention-Augmented Contextualized Embeddings attention head circuit Attention Residuals Attention Sorting Attention submodule Attention-based MIL Attractor States Emerge in Multi-Turn LLM Conversations AUC Audex Audio Interaction Model Audio MultiChallenge Audio-Flamingo-3 Audio-LLM Audio-Native Speech Recognition with a Frozen Discrete-Diffusion Language Model AudioCards AudioDER AudioLDM 2 AudioVAE2 Auditing Protocol-Level Shortcuts in Large Audio Language Model Judges for Speech Evaluation AUDITS Augment Augment Code AUPRC AuRA AUROC Australia AI Safety Institute Australia National AI Plan Australian Emergency Department triage notes Australian Government Australian National University Auto Benchmark Audit (ABA) Autodata Autodesk Fusion AutoDex autoDream AutoForest Autoformer AutoGPTQ AutoJourn AUTokenizer AutoLab automated AI research Automated Background Swapping Automated Compliance Mapping in Cloud Security with Domain-Adapted Sentence Transformers Automated Discovery Has No Universally Superior Harness Automated Extraction of Techno-Economic Data from 76,000 Energy System Studies Automated grading of Linux/bash examinations using large language models: a four-level cognitive taxonomy approach automated mechanistic interpretability automated red teaming Automated Reference Verification System Automated reproducibility assessments in the social and behavioral sciences using large language models automated test suite automated theorem proving Automatic Domain Randomization Automatic Post-Editing (APE) Automatic Speech Recognition AutomationBench-AA AutoMem autonomous driving AUTOPILOT-VQA Autoregressive Boltzmann Generators autoregressive transformer autoresearch AutoResearchClaw AutoRound AutoSkillHarm AutoSynthesis avatarin Avenger Labs Aviral Kumar awesome-claude-code awesome-copilot awesome-llm-apps AWQ AWS AWS Bedrock AWS GovCloud AWS Inferentia2 AWS Labs AWS Marketplace AWS Neuron AWS Neuron SDK AWS Trainium AxDafny Axel Backlund Axiom Math Axios Axolotl AXPO Aya Aya Expanse Aya Vision Aya-Expanse-8B-Base Azure Azure AI Studio Azure DevOps Azure Foundry Azure OCR B³D-RWKV BabelJudge BABI-Long BABILong BabyCL Backbone-as-Architect backdoor attack Backdoor Circuit Analysis (Language-Switching) Backdoor Unlearning Generalization: A Path Toward the Removal of Unknown Triggers in LLMs backnotprop backpropagation through time Backstory Bad Grammar Bad Scientist BadWAM BadWAM: When World-Action Models Dream Right but Act Wrong Bag of Words (BoW) BAGEL BAGEL-7B Baidu Baidu Translate Baifeng Shi Baiyu Chen ball balancing ball balancing task Bamba BamiBERT Bandit Bang-v3 BanglaBERT Banque des Territoires Barcelona Supercomputing Center Language Technologies bare metal sandboxes Bargaining Scenarios Dataset Bark Base Baseline Agent Baseline-Log Physical Separation Baseten BashArena basic-memory basicmachines-co Basque Center on Cognition, Brain, and Language Basque language baulab Bayesian decision theory Bayesian Inference and Decision Audits for Public Archives of Frontier AI Evaluations Bayesian Multiobjective Optimization Bayesian Optimal Experimental Design Bayesian Optimization BayesPO BayLing-Duplex BBC News BBQ BBVA BC Protocol BCG X BCR BDD100K Be My Eyes Be My Tutor: On-Policy Co-Distillation for Mutual LLM Improvement via Peer Feedback BEA-Dialogue Beam Search Bebop Bechiri and Lanasri [2026] Becker Friedman Institute for Economics Before You Think: System 0, AI-Mediated Cognition and Cognitive Colonization behavior cloning behavior trees with learning-enabled components behavioral fine-tuning Behavioral Trajectory Tracking Framework behavioral-gradient validator Behavioral-SafetyBench Behavox BEIR Belebele Belief-Space Safety Filter (BeliefSF) BeliefTrack Ben Bernanke Ben Mann Ben Parr Ben Thompson Benchling Benchmark Agent Benchmark Everything Everywhere All at Once Benchmarking Multimodal Large Language Models for Scientific Visualization Literacy benchmaxxing BenCzechMark Bending Spoons Beneficial Deployments Benevity BenevolentAI Bengali Bengaluru BenHalluEval BenHalluScore Benjamin Chang Benjamin Krause Berkeley AI Research (BAIR) Berkeley Artificial Intelligence Research Berkshire Hathaway Berlin Database of Emotional Speech (EMO-DB) BerriAI BERT BERT-base BERT-F1 BERTopic BERTScore BERTurk BERTweet BES (trained models) Bessemer Venture Partners Best-of-N Best-of-N Sampling best@k Better Call GRPO Beyond a Single Judge: Simulating Social Persona Panels for Generative UI Evaluation Beyond Accuracy: Community Perspectives on Machine Translation Beyond Adam: SOAP and Muon for Faster, Label-Efficient Training of Machine Learning Interatomic Potentials Beyond Aggregate Risk: Role-Stratified Conformal Risk Control for LLM Tool Calls Beyond Attack-Success Rate: Action-Graded Severity Scale for Tool-Using AI Agents Beyond Benchmarks: Exposing the Hidden Crisis in Bangla Hate Speech Detection Beyond Fixed Representations: The Vocabulary and Verifier Gaps in Open-Ended AI Beyond Fully Random Masking: Attention-Guided Denoising and Optimization for Diffusion Language Models Beyond Function Calling: Benchmarking Tool-Using Agents under Tool-Environment Unreliability Beyond Global Replanning: Hierarchical Recovery for Cross-Device Agent Systems Beyond Independent Labels: Schwartz-Geometry Decoding for Human Value Detection Beyond Negative-Ridge Endpoints: Mixed-Sign Spectral Regularization via Negative-Shifted Gradient Descent Beyond Runtime Enforcement: Shield Synthesis as Defensibility Analysis for Adversarial Networks Beyond Scale and Generation: Understanding Language Model-based Entity Matching Beyond Sentiment: Structured Information Extraction from Financial News Beyond Success Rate: Cost-Aware Evaluation of Offensive and Defensive Security Agents Beyond Sufficiency: Time Series Explanation with Counterfactual Necessity Beyond Surface Forms: A Comprehensive, Mechanism-Oriented Taxonomy of Indirect Linguistic Encoding for LLM-Based Coded Language Detection Beyond Sycophancy: Structured Resistance and Compliance in LLM Moral Reasoning Beyond task performance: Decoding bioacoustic embeddings with speech features Beyond the Commitment Boundary: Probing Epiphenomenal Chain-of-Thought in Large Reasoning Models Beyond the Hard Budget: Sparsity Regularizers for More Interpretable Top-k Sparse Autoencoders Beyond the Leaderboard: Design Lessons for Trustworthy Multimodal VQA Beyond Third-Person Audits: Situated Interaction Auditing for User-Centered LLM Bias Research Beyond Tokenization: Direct Timestep Embedding and Contrastive Alignment for Time-Series Question Answering Beyond Uniform Tokens: Adaptive Compression for Time Series Language Models BeyondUncertainty BFCL BFCL Multi-Turn BFCL-V3 BFCLv3 BFW schedule hint BGE BGE-M3 BGL bi-manual robotic manipulation Bias Benchmark for Question Answering Bias Leaves a Gradient Trail Biconvex Optimization bidirectional attention Bidirectional Evolutionary Search Bielik Big Bench Big Bench Audio Big-Math BigBird BigCode BigCodeArena BigCodeBench Bigger or Cheaper? Scale and Quantization Effects on Uncertainty Signals in Vision-Language Models Under Image Degradation BigLaw Bench BigScience Bill & Melinda Gates Foundation Bill Dally Bill Gurley Bill McDermott Binarized Neural Network binary quantization BINEVAL Bing binning semiring bioacoustic monitoring bioacoustics BioASQ 14b BioBERT Biographical