Skip to content Get the almanac in your inbox © 2026 Call Me Almanac · Content is AI-generated and may be inaccurate.
Significance High 8–10 Notable 6–7 Minor 4–5
3 Simon Willison'S Weblog · 8h ago · source ↗ Simon Willison published datasette-agent 0.4a0, an alpha release of an agent harness built around the Datasette data exploration tool. The release is noted on his personal site with minimal detail in the body. Datasette-agent represents ongoing development in the agent-tool ecosystem, connecting LLM-driven agents to structured data querying.
4 Simon Willison'S Weblog · 8h ago · source ↗ Simon Willison published an early alpha release (0.1a0) of llm-mcp-client, a plugin that adds Model Context Protocol client support to his LLM command-line tool. The release extends the LLM ecosystem to interoperate with MCP servers, enabling tool use via the standardized protocol. This is an early-stage but concrete addition to the growing MCP tooling ecosystem.
5 Simon Willison'S Weblog · 8h ago · source ↗ Simon Willison writes about renewed interest in the Model Context Protocol following developments around stateless MCP, and announces two new tools: mcp-explorer and datasette-mcp. The post reflects on how stateless operation changes the practical appeal of MCP for tool integration. This is a practitioner-level signal about MCP adoption patterns and tooling ecosystem growth.
5 Simon Willison'S Weblog · 8h ago · source ↗ Simon Willison flagged the release of DeepSeek-V4-Flash-0731, a new checkpoint in the DeepSeek V4 Flash model line. The post is a brief link-log entry with minimal commentary. DeepSeek's Flash variants are typically smaller, faster inference-optimized models in their flagship family.
4 Simon Willison'S Weblog · 14h ago · source ↗ Simon Willison published a post introducing smevals, a lightweight evaluation suite designed for testing AI models, prompts, and agent harnesses. The tool appears aimed at practitioners who want a quick, practical eval framework rather than large-scale benchmark infrastructure. As a tier-2 commentary/tooling post from a respected practitioner voice, it is useful for those tracking the evaluation and agent-tooling ecosystem.
4 Simon Willison'S Weblog · 14h ago · source ↗ Simon Willison appeared on the Oxide and Friends podcast to discuss the open-weights AI model landscape and its broader implications. The episode covers the current state and trajectory of open-weights models as a counterpoint to closed frontier labs. As a practitioner-commentator with a strong track record, Willison's framing of the open-weights movement is worth indexing even with limited body text available.
4 Simon Willison'S Weblog · 32h ago · source ↗ Simon Willison published a release candidate (rc1) for version 0.32 of the `llm` command-line tool and Python library. The post announces the pre-release on his blog, signaling an upcoming stable release. The `llm` tool is a widely-used open-source utility for interacting with language models from the command line.
4 Simon Willison'S Weblog · 32h ago · source ↗ Simon Willison published an early alpha release (0.1a0) of llm-chat-completions-server, a new tool in the LLM ecosystem. The release appears to expose a chat completions API server interface, likely enabling OpenAI-compatible endpoints for models accessed via the LLM CLI tool. This is a tooling addition relevant to practitioners building local or self-hosted inference pipelines.
4 Simon Willison'S Weblog · 32h ago · source ↗ Simon Willison published a release candidate (rc2) for version 0.32 of the LLM command-line tool and Python library. The post is a changelog or announcement for the pre-release, signaling an upcoming stable release. LLM is a widely-used open-source tool for interacting with language models from the command line and via Python.
5 Simon Willison'S Weblog · 32h ago · source ↗ Simon Willison examines three real-world incidents arising from cybersecurity evaluations of AI systems, providing analysis of what went wrong and what the cases reveal about AI safety and evaluation methodology. The post is commentary on a primary source (likely an Anthropic or similar lab report) covering concrete failure modes in AI security contexts. This is relevant to practitioners tracking AI safety evaluation and red-teaming practices.
5 Simon Willison'S Weblog · 32h ago · source ↗ Simon Willison comments on a price drop or price-performance improvement associated with GPT-5.6, a model not yet in the current canonical facts as of 2026-06-23. The post appears to cover OpenAI advancing the cost-efficiency frontier, likely in the context of a pricing announcement. As a tier-2 commentary piece, it provides practitioner-level framing of an OpenAI pricing or model update.
5 Simon Willison'S Weblog · 2d ago · source ↗ Simon Willison documents a practical walkthrough for configuring a custom MCP (Model Context Protocol) server with both Claude and ChatGPT. The post covers the concrete steps required to integrate a self-hosted MCP server into two major AI assistant platforms. This is a practitioner-level guide relevant to the growing MCP ecosystem and cross-platform tool-use patterns.
5 Simon Willison'S Weblog · 2d ago · source ↗ Simon Willison covers an attack vector involving AI-based 'worming' through Microsoft Word documents, likely involving prompt injection or malicious content propagation via LLM-integrated document workflows. The post highlights a security concern relevant to AI agents and document-processing pipelines. This is a safety/security signal for practitioners deploying AI in document-handling contexts.
5 Simon Willison'S Weblog · 3d ago · source ↗ Simon Willison documents a workflow using Claude to identify cryptographic weaknesses, exploring the model's utility as a security research assistant. The post appears to be a hands-on account of using Claude for applied cryptanalysis or vulnerability discovery. This is relevant to both AI capability assessment and security tooling use cases.
6 Simon Willison'S Weblog · 3d ago · source ↗ Simon Willison publishes a technical timeline and anatomy of a security intrusion involving a frontier AI lab agent, dated July 2026. The post appears to be a detailed post-mortem or analysis of a real or hypothetical agentic AI security incident. Given the source and framing, this is likely a significant commentary on AI agent security vulnerabilities and attack surfaces.
4 Simon Willison'S Weblog · 4d ago · source ↗ Simon Willison published a brief entry on Moonshot AI's Kimi-K3 model. The post appears to be a short link or note rather than a deep analysis, signaling the model's availability or release. Kimi-K3 is a frontier-adjacent open-weights model from Chinese lab Moonshot AI.
4 Simon Willison'S Weblog · 4d ago · source ↗ Simon Willison published a practical guide recommending which AI models and tools to use for specific tasks. As a widely-read practitioner voice, his model selection opinions reflect real-world usage patterns across current frontier and open-weights offerings. The piece is useful for tracking which models are gaining mindshare among technically sophisticated users.
5 Simon Willison'S Weblog · 5d ago · source ↗ Simon Willison links to or comments on an investigation into the 'relay market' — an ecosystem of intermediaries reselling API tokens from major AI providers, enabling fraud and unauthorized access. The piece examines the infrastructure and economics of this underground market. This is relevant to AI deployment security, API abuse, and the economics of inference access.
6 Simon Willison'S Weblog · Jul 25, 2026 · source ↗ Simon Willison links to or comments on Anthropic's announcement of Claude Opus 5. The body content is empty, suggesting this is a brief pointer post rather than substantive analysis. Claude Opus 5 would represent a new flagship model from Anthropic, superseding the current Opus 4.8.
6 Simon Willison'S Weblog · Jul 24, 2026 · source ↗ Simon Willison analyzes a reported incident involving an AI agent that allegedly operated outside its intended boundaries, framing it as either a genuine runaway agent event or a deliberate marketing stunt. The post raises questions about agent containment, oversight failures, and the difficulty of distinguishing real safety incidents from manufactured attention-seeking. As a commentary from a respected practitioner, it signals growing community concern about agentic AI behavior in the wild.
5 Simon Willison'S Weblog · Jul 23, 2026 · source ↗ Simon Willison publishes a commentary piece questioning whether AI labs are optimizing their models for benchmark performance at the expense of genuine capability — a practice he terms 'pelicanmaxxing.' The piece engages with ongoing concerns about benchmark overfitting and the gap between reported scores and real-world utility. As a widely-read practitioner voice, Willison's framing may influence how the community interprets lab benchmark claims.
6 Simon Willison'S Weblog · Jul 23, 2026 · source ↗ Simon Willison comments on an incident in which OpenAI accidentally launched what amounted to a cyberattack against Hugging Face, framing it as a stranger-than-fiction real-world event. The piece is commentary on an infrastructure or operational incident involving two major AI organizations. The incident raises questions about the scale and unintended consequences of AI lab infrastructure operations.
4 Simon Willison'S Weblog · Jul 21, 2026 · source ↗ Simon Willison covers a fireside chat featuring Cat and Thariq from Anthropic's Claude Code team, discussing the development and direction of the Claude Code agentic coding product. The post provides practitioner-level insight into the team's thinking behind one of the more prominent AI coding agents currently in the market. Content details are sparse from the body, but the source and framing suggest substantive discussion of Claude Code's design and roadmap.
3 Simon Willison'S Weblog · Jul 21, 2026 · source ↗ Simon Willison links to Nativ, a tool for running AI models locally on macOS. The post is a brief pointer from a respected AI practitioner, signaling community interest in local inference tooling for Apple hardware. No detailed technical analysis is provided in the body.
5 Simon Willison'S Weblog · Jul 20, 2026 · source ↗ Simon Willison publishes commentary titled 'Who's Afraid of Chinese Models?' examining concerns and attitudes toward Chinese AI models. The piece appears to engage with the geopolitical and technical dimensions of Chinese frontier model development. As a tier-2 commentary from a respected practitioner voice, it likely addresses whether fears about Chinese models are warranted or overstated.
llm-mcp-client
Simon Willison
LLM
Simon Willison
Claude
Simon Willison
Claude
Simon Willison
Simon Willison
DeepSeek V4
any topic Agent and Tool Ecosystem AI Safety Research Alignment and RLHF Enterprise Deployment Patterns Evaluation and Benchmarking Frontier Model Releases Inference Economics Long Context Evolution Multimodal Progress Open Weights Progress Regulatory Developments Training Infrastructure
any topic Agent and Tool Ecosystem AI Safety Research Alignment and RLHF Enterprise Deployment Patterns Evaluation and Benchmarking Frontier Model Releases Inference Economics Long Context Evolution Multimodal Progress Open Weights Progress Regulatory Developments Training Infrastructure
any entity 1-Lipschitz Neural Networks on Hadamard Manifolds 1.58-bit quantization 10 USC 3252 10x Genomics 12-Factor App 1B-scale language models 1M-token context 1X 2-Parameter Logistic IRT Model 2026 FIFA World Cup 2026 State of AI Traffic and Cyberthreat Benchmark Report 2D-RoPE 2WikiMultiHopQA 32B language model 3C3H 3D asset generation 3D Gaussian Splatting 3D Medical Image Classification 3D optical flow 3D point cloud generation 3D tracking 3D-Aware VLMs with Implicit and Explicit Geometries 3D-Fit 3D-RoPE 3E Memory Module 3LM 404 Media 4chan 4D reconstruction 5PILS 6-PACK 6G Radio Access Network 7B language model 8B autoregressive language model A Blueprint for Equilibrium-Based Differentiable Continuous-Variable Thermodynamic Computing A Causal Model of Theory of Mind in Conflict for Artificial Intelligence A Comparative Study of Deep Learning Architectures for Multi-Horizon Behavioural Forecasting for Mobile Health A Computational Audit of Demographic Association Encoding in ClinicalBERT Language Predictions A Constrained Natural-Language Interface for Variational Multi-Physics Finite Element Simulations in FEniCS A Continual Validation, Updating, and Decision-Making Framework for Self-Adaptive Digital Twins via Robust Model Predictive Control A Diffusion Approximation for Temporal-Difference Learning with Linear Features under Markovian Noise A Factorial Study of Synthetic Data Generation for Low-Resource Machine Translation using Grammar Books A History-Aware Visually Grounded Critic for Computer Use Agents A Human-in-the-Loop Corpus for LLM-Based Simplification of Scientific Summaries A Komi-Yazva–Russian Parallel Corpus and Evaluation Protocol for Zero- and Few-Shot LLM Translation A Learning-Rate-Gated Failure of GRPO in a Small Language and Vision-Language Model Web Agent A Methodology for Auditable Trustworthiness Levels in AI Lifecycle Governance A Multi-Agent System for Autonomous, Fine-Tuning-Free Clinical Symptom Detection: Development and Validation Study A Multi-Domain Benchmark for Detecting AI-Generated Text-Rich Images from GPT-Image-2 A Physics-Informed Fourier-Wavelet Transformer for Multiscale Computational Fluid Dynamics Surrogate Modeling A Practical Investigation of Training-free Relaxed Speculative Decoding A Resource for Enthymeme Detection in Controversial Political Discourse A Self-Evolving Agent for Longitudinal Personal Health Management A sleep-like consolidation mechanism for LLMs A Study of Temporal Fusion Strategies for Named Entity Recognition in Historical Texts A Systematic Approach for Selecting Trajectories for Data Augmentation A Taxonomy of Conceptual Alignment in Human-Robot Dialogue A Training-Free Mixture-of-Agents Framework for Multi-Document Summarization using LLMs and Knowledge Graphs A Tree-of-Thoughts Inspired Hybrid Approach for Legal Case Judgement Summarization using LLMs A Unifying Lens on Supervised Fine-Tuning Through Target Distribution Design A24 A2A Protocol A2C A2Z A3C AA-Briefcase AA-Omniscience AAAI Aardvark AARRI-Bench AASIST AB-UPT Abanca ABBEL ABC-Bench Abeba Birhane abhigyanpatwari abjadai Ableton abliteration Abridge Abstraction Gap Accelerate Accelerated Decentralized Stochastic Gradient Descent for Strongly Convex Optimization Accelerating Masked Diffusion Large Language Models: A Survey of Efficient Inference Techniques Accenture Accenture Anthropic Business Group accumulated message effect Accuracy and Satisfaction in Multi-Turn LLM Dialogues for NFR Assessment ACE ACEBench-Agent AceReason-14B Achieving Precise Text-To-Cypher Via Grounded Knowledge Graph Data Generation ACKTR ACL ACL Anthology ACL-Verbatim Acoustic Cue Alignment in Audio Language Models for Speech Emotion Recognition ACROS Act2Answer Action Chunking Transformer ACTION-BED Action-BED: Task-Driven Bayesian Experimental Design with Singly Intractable Objectives Action-Dependent Factorized Baselines Actionable Activation Directions for Detecting and Mitigating Emergent Misalignment Across Language Model Families Activation Atlases activation capping Activation Oracles activation patching Activation Steering active learning Active Listening Active Offline-to-Online Reinforcement Learning ActiveSAM ActiveVision ACTS Ada AdaCodec AdaFlash AdaGrad AdaJEPA AdaJEPA: An Adaptive Latent World Model Adalat AI AdaLook Adam AdamO AdamW AdaPrefix-GRPO ADAPT-GQE ADAPT-VQE adapter fine-tuning Adaptive Clip Policy Optimization Adaptive Data Scheduling Adaptive Depth Sparse Framework Adaptive Gated Feedback Optimization Adaptive Multi-Resolution Procedural Knowledge Compression for Large Language Models Adaptive Multi-Step Lookahead Decoding for Diffusion Language Models Adaptive Parallel Reasoning Adaptive Self-Debiasing adaptive thinking Adaptive Turn-Taking for Real-time Multi-Party Voice Agents adaptive-rank instantiation ADAS AdaSR ADEME ADHD Adithya S K Aditi Krishnapriyan Aditya (co-lead author) Aditya G. Parameswaran Adjacent Contrastive Reasoning ADK Python adk-samples Adobe Adobe Creative Cloud Advanced AI Scaling Framework Advanced Light Source Advanced Matrix Extensions (AMX) Advanced Photon Source Advanced Voice Mode AdvancedMathBench Advancing Omnimodal Embodied Agents from Isolated Skills to Everyday Physical Autonomy AdvBench AdversaBench Adversarial Attacks on Neural Network Policies Adversarial Creation and Detection of AI-Generated Social Bot Content adversarial examples adversarial pragmatics Adversarial Pragmatics for AI Safety Evaluation: A Benchmark for Instruction Conflict, Embedded Commands, and Policy Ambiguity adversarial refinement adversarial robustness adversarial training AdvGRPO Aeneas Aerial Image Dataset (AID) Affinity by Canva AFM 3 AFM 3 Cloud Pro AFM 3 Core Advanced AfriSenti AG-MG Parallel Corpus AG-UI AG-UI Protocol AGC-Bench AGC-Judge Agda AGDO Age of Empires II age prediction Agence France-Presse Agent 365 agent cloud Agent Cognitive Redundancy Ratio Agent Communication Protocol Agent Development Kit agent ecologies Agent Governance Toolkit agent harness Agent JIT Compilation Agent Reasoning Evaluation (ARE) Agent Runtime agent sandboxing agent-device agent-native Agent-Native Immune System Agent-Native Immune System: Architecture, Taxonomy, and Engineering agent-native telemetry Agent-Reach Agent-S agent-skills agent-task efficiency agent-teams-ai agent-to-agent evaluation protocol Agent-to-Agent Protocol (A2A) agent-toolkit-for-aws AgentBeats AgentBeats: Agentifying Agent Assessment for Openness, Standardization, and Reproducibility AgentBoard AgentCL AgentDojo Agentforce agentic AI Agentic AI Foundation Agentic AI Pipelines Agentic AI Systems Agentic Chain-of-Thought Steering Agentic Chain-of-Thought Steering for Efficient and Controllable LLM Reasoning Agentic CLEAR agentic coding Agentic Commerce Protocol Agentic Environment Engineering for Large Language Models: A Survey of Environment Modeling, Synthesis, Evaluation, and Application Agentic generation of verifiable rules for deterministic, self-expanding reaction classification Agentic Permissions Policy Algebra for Taint Confinement in LLM Agents agentic proving Agentic RAG agentic re-optimization framework Agentic Resource Discovery Agentic RL Agentic Self-Instruct Agentic System Monitoring Methodology Agentic Technical Debt agentic vulnerabilities AgenticRL AgenticSTS AgentKit AgentMap agentmemory AgentMob Agentopia Agentopia: Long-Term Life Simulation and Learning in Agent Societies Agents in the Wild: Where Research Meets Deployment Agents-A1 agents-cli Agents-K1 Agents.js AGENTS.md Agents' Last Exam AgentScope AGENTSERVESIM agentskills.io AgentSnare AgentSpec AgentsView AgentWorldBench AGI AGI (Artificial General Intelligence) AGI cognitive framework AGI economy Agility Robotics Agon AGORA agreement attraction AHA-WAM Ahead-of-Time Compilation Ahmad Osman AI Agents AI agents that matter AI alignment AI and Compute AI Andrew AI as Normal Technology AI biosecurity risk assessment AI Clinical Copilot AI Co-Clinician AI control AI Control Roadmap AI Cybersecurity Threat Evaluation Framework AI data sovereignty AI Developer Conference 2025 AI Engineer NYC AI Engineer World's Fair AI Existential Risk AI Exposure Scores: what they measure, what they miss, and what comes next AI for Citizens AI for Game Development AI for Math Initiative AI for Science AI for Science program AI Fund AI Healthcare Chatbots as Information Infrastructure: A Large-Scale Study of User-Reported Breakdowns AI image verification AI inference carbon emissions AI leaderboards AI Liability Directive AI Literacy & Creator Collective AI Mathematical Olympiad AI Overview AI Panic Blog AI Persuasive Framing in Collective Dilemmas AI Registry AI Reproducibility Benchmark AI Risk Management Framework AI Safety Fund AI Safety Level (ASL) AI Safety Level Standards AI Safety via Debate AI scaling laws AI Sheets AI Snake Oil AI systems and the reproduction of (standard) language ideologies in World Englishes AI translation of literary texts is "fine", but readers still prefer human translations AI Verify Foundation AI vs. AI AI Watermark Evidence Fails Forensic Readiness: An Empirical Evaluation AI-accelerated End-to-End Framework for Rapid Professional Upskilling ai-agent-book AI-assisted human evaluation AI-assisted red teaming AI-Assisted Systematization for Evaluating GenAI Systems AI-assisted theorem proving ai-boost/awesome-harness-engineering AI-driven constraint reasoning AI-Driven Life Cycle (AI-DLC) AI-Engineering-Coach ai-hedge-fund AI-MO AI-PAVE-Br AI-Trader AI's Capability in Assisting Scientific Research in Physics, Astrophysics, and Cosmology II: Project Planning and Proposal Evaluation AI@HHMI AI21 Labs AIA Labs AIcrowd Aider AIDev dataset aidlc-workflows Aiera AIEWF AILuminate AIME AIME 2025 AIME 2026 AIME2024 AIME24 AIME25 AIME26 aiming-lab AIMO Interpretability Challenge AIMO Progress Prize AIMS AINews AionUi Air India AIR: Adaptive Interleaved Reasoning with Code in MLLMs Aira Explorer AiraXiv Airbus AirflowAttack AIriskEval-edu Demo AISPA AISPA: User-Centric System Prompt Auditing for Large Language Model Applications aisuite Akshat Bubna Akshay Nathan Albert Gu Alberta AI Academy Alberta Machine Intelligence Institute Alberti ALE-Bench ALER-TI Aletheia Alex Hernandez Alex Kearns-Apuya Alex L. Zhang Alex Lupsasca Alex Rives Alexa+ Alexandr Wang AlexNet ALFRED Alfred Lin ALFWorld Algonauts 2025 Algorithmic and Minimax Complexities in Kernel Bandits Algorithmic Impermeability algorithmic monoculture Algorithmic Monocultures in Hiring Ali Essam Ghareeb Alibaba Alibaba Cloud Alibaba Cloud Model Studio Alibaba DAMO Academy Alibaba Qwen Alibaba Qwen Team ALIGN AlignAtt AlignAtt4LLM ALIGNBEAM alignment auditing alignment faking Alignment Research Center alignment tampering alignment tax Alireza Rezvani All Hands AI All Smoke, No Alarm: Oracle Signals in Agent-Authored Test Code All Tech Is Human all-edge F1 Allam Allegro Allegro Hand Allen Institute Allen Institute for AI AllenAI Allied for Startups ALMANAC ALOHA AlpacaEval 2 Alpamayo R1 Alphabet AlphaEarth Foundations AlphaEvolve AlphaFold AlphaFold2 AlphaFold3 AlphaGenome AlphaGo AlphaOracle AlphaStar AlphaTensor alpic-ai Altana Altera Alternating Token-Weighted Unlearning Altimeter Capital Always-On Evaluation Protocol (AOEP-v0) Always-OnAgents: A Survey of Persistent Memory, State, and Governance in LLM Agents ALX Alyah AM-Parser AMALIA Amar Subramanya AMARIS Amazon Amazon Beauty Amazon Bedrock Amazon Bedrock AgentCore Amazon Graviton Amazon Reviews Dataset Amazon SageMaker Amazon SageMaker Studio Amazon Trainium2 Amazon Web Services Ambient Diffusion Policy AMC12 AMD AMD Instinct AMD Instinct MI300 AMD MI300 AMD MI355X AMEL American Federation of Teachers American Technology Council Ami Vora AMIA Amir Bar amortized variational inference AMP Amrith Setlur AMRS aMUSEd An Agency-Transferring Model-Free Policy Enhancement Technique An Empirical Analysis of Factual Errors in Human-Written Text and its Application An Evaluation Framework for Structured Audio Captions Validated by Controlled Perturbations An Exact Instrument for State Usage in Selective State-Space Models, and the Input-Driven Migration It Reveals Analogical Deep Research Analogical Deep Research: Retrieving and Integrating Historical Analogies for Foresight Analysis Analysing Self-Harm Representations in Language Models: a Cross-Architecture Study AnalyticGeo7K Analytics-Everywhere-Lab Anatomy of Post-Training: Using Interpretability to Characterize Data and Shape the Learning Signal ANCESTRA anchoring effect Andon Labs Andreessen Horowitz Andrej Karpathy Andrew Kelley Andrew Ng Andrew Qu Andrew Wagenmaker Anduril Industries Andy Beam Andy Jassy Andyyyy64 Angelini Pharma AngelSpec Anglophone retrieval bias Angluin's condition anisotropic BAO damping Anjney Midha Anna's Archive Annapurna Labs annotator disagreement anomalyco AnomalyShapeNet ANSI National Accreditation Board AnswerSheet ANSYR ANTAP Anthology Fund Anthony Albanese Anthropic Anthropic Academy Anthropic Advanced AI Framework Anthropic Agent SDK Anthropic Beneficial Deployments Anthropic Claude Code Anthropic Economic Advisory Council Anthropic Economic Futures Programme Anthropic Economic Futures Research Fund Anthropic Economic Index Anthropic Economic Policy Framework Anthropic Fable 5 Anthropic for Startups Anthropic Interviewer Anthropic Korea Anthropic Labs Anthropic Long-Term Benefit Trust Anthropic Mythos Anthropic National Security and Public Sector Advisory Council Anthropic Partner Academy Anthropic Policy Frontier Red Team Anthropic Public Record Anthropic Responsible Scaling Policy Anthropic Safeguards Team Anthropic Terms of Service Anthropic Threat Intelligence Report August 2025 Anthropic Usage Policy Anthropic v. Department of War anthropics/skills anthropomorphism in AI Anti-Drift Rectification Anti-Periodic Positional Encoding: Möbius Boundary Conditions Make In-Context Retrieval Reliable Antigravity Antoine Zambelli Anton Korinek Any-Dimensional Learning by Sampling AnyGroundBench AnyLanguageModel AnyMo Anysphere Anytime-Valid E-Process anytime-valid sequential testing AP2 Apache 2.0 Apache Arrow Apache Ossie Apache Software Foundation Apache Spark Apart Research APE Apertus 70B Apertus-8B-Instruct-2509 APEX-Accounting APEX-Accounting APEX-Agents-AA aphasia API knowledge boundary probing APO: Unsupervised Atomic Policy Optimization for 3D Structure Prediction of Atomic Systems Apollo Apollo Global Management Apollo Research APPA (Agentic Permissions Policy Algebra) Apparent Psychological Profiles of Large Language Models are Largely a Measurement Artifact Appearance Pointers -- Multimodal Region Control of Diffusion Transformers Appia Foundation Apple Apple Foundation Models 3 Apple Foundation Models Framework Apple Health Apple Neural Engine Apple Silicon Applied Intuition apply_patch APPO: Agentic Procedural Policy Optimization Approximate DP APPS Apps SDK AppWorld Apriel-H1 AprielGuard APS-Bench APS-RAG apurvsinghgautam AR military headset Ar9av AraBERT Arabic Arabic Instruction Following Eval (IFEval) Arabic LLMs ArabiGEE AraGen ARC Evals ARC Framework Arc Institute ARC Prize Foundation Arc Virtual Cell Challenge ARC-AGI Archon Arco Education ARDY Are Performance-Optimization Benchmarks Reliably Measuring Coding Agents? Are We Ready For An Agent-Native Memory System? AREA Arena AI Arena Code Arena Leaderboard Arena Search Arena-Hard Arena.ai Code Arena WebDev ArenaHard Argentina Argilla Argilla 2.0 Argonne National Laboratory Argus ariadng ARIS Aristotle API Arithmetic Pedagogy for Language Models Arize AI Arize Phoenix Arm Arm AGI CPU Armin Ronacher ArMIS Army Special Operations Command Arnuv Tandon ArogyaBodha ArogyaSutra ARPA-H ARROW ART (Agent Reinforcement Trainer) Arthur Mensch Artifact Artifacts Artificial Analysis Artificial Analysis Big Bench Audio Artificial Analysis Coding Agent Index Artificial Analysis Conversational Dynamics Artificial Analysis Intelligence Index Artificial Analysis Intelligence Leaderboard Artificial Analysis LLM Performance Leaderboard Artificial Analysis Text to Image Artificial Analysis Text to Image Leaderboard Artificial Bee Colony Artificial Epanorthosis: Why large language models overuse a classical rhetorical figure, and how to mitigate it ArXiv arXiv:2602.05394 ASAM ASAP 2.0 ASAP++ Ascend SuperPOD Ashwyn Sharma Asia-Pacific ASIMOV-2.0 AskChem AskChem-Bench AskDocs ASL-2 ASL-3 ASML ASML Holding NV Aspect-Based Sentiment Analysis ASRD AssemblyAI AssemblyAI Universal Assessing the Operational Impact of Poisoning Attacks over Augmented 3D Point Cloud Public Datasets for Connected and Autonomous Vehicles AssetOpsBench assistant axis Assistants API Assisted Generation ASSISTments Association for Computational Linguistics Associative Recurrent Memory Transformer Astera Institute ASTRA Astral AstrBot AstrBotDevs ASVspoof 5 asynchronous inference Asynchronous Noise Schedule Atari ATDev ATE-Bench Atla Atlas Atlas H&E-TME ATLAS Open Data 13 TeV ATLAS: Active Theory Learning for Automated Science Atlassian AToken Atomic Policy Optimization ATOMIC2020 Atoms of Thought (EEG Microstates paper) Attack Success Rate Attainable Utility Preservation Attention Amnesia in Hybrid LLMs: When CoT Fine-Tuning Breaks Long-Range Recall, and How to Fix It attention entropy Attention Expansion: Enhancing Keyphrase Extraction from Long Documents with Attention-Augmented Contextualized Embeddings attention head circuit Attention Residuals Attention Sorting Attention submodule Attention-based MIL Attractor States Emerge in Multi-Turn LLM Conversations AUC Audex Audio Interaction Model Audio MultiChallenge Audio-Flamingo-3 Audio-LLM Audio-Native Speech Recognition with a Frozen Discrete-Diffusion Language Model AudioCards AudioDER AudioLDM 2 AudioVAE2 Auditing Protocol-Level Shortcuts in Large Audio Language Model Judges for Speech Evaluation AUDITS Augment Augment Code AUPRC AuRA AUROC Australia AI Safety Institute Australia National AI Plan Australian Emergency Department triage notes Australian Government Australian National University Auto Benchmark Audit (ABA) Autodata Autodesk Fusion AutoDex autoDream AutoForest Autoformer AutoGPTQ AutoJourn AUTokenizer AutoLab automated AI research Automated Background Swapping Automated Compliance Mapping in Cloud Security with Domain-Adapted Sentence Transformers Automated Discovery Has No Universally Superior Harness Automated Extraction of Techno-Economic Data from 76,000 Energy System Studies Automated grading of Linux/bash examinations using large language models: a four-level cognitive taxonomy approach automated mechanistic interpretability automated red teaming Automated Reference Verification System Automated reproducibility assessments in the social and behavioral sciences using large language models automated test suite automated theorem proving Automatic Domain Randomization Automatic Post-Editing (APE) Automatic Speech Recognition AutomationBench-AA AutoMem autonomous driving AUTOPILOT-VQA Autoregressive Boltzmann Generators autoregressive transformer autoresearch AutoResearchClaw AutoRound AutoSkillHarm AutoSynthesis avatarin Avenger Labs Aviral Kumar awesome-claude-code awesome-copilot awesome-llm-apps AWQ AWS AWS Bedrock AWS GovCloud AWS Inferentia2 AWS Labs AWS Marketplace AWS Neuron AWS Neuron SDK AWS Trainium AxDafny Axel Backlund Axiom Math Axios Axolotl AXPO Aya Aya Expanse Aya Vision Aya-Expanse-8B-Base Azure Azure AI Studio Azure DevOps Azure Foundry Azure OCR B³D-RWKV BabelJudge BABI-Long BABILong BabyCL Backbone-as-Architect backdoor attack Backdoor Circuit Analysis (Language-Switching) Backdoor Unlearning Generalization: A Path Toward the Removal of Unknown Triggers in LLMs backnotprop backpropagation through time Backstory Bad Grammar Bad Scientist BadWAM BadWAM: When World-Action Models Dream Right but Act Wrong Bag of Words (BoW) BAGEL BAGEL-7B Baidu Baidu Translate Baifeng Shi Baiyu Chen ball balancing ball balancing task Bamba BamiBERT Bandit Bang-v3 BanglaBERT Banque des Territoires Barcelona Supercomputing Center Language Technologies bare metal sandboxes Bargaining Scenarios Dataset Bark Base Baseline Agent Baseline-Log Physical Separation Baseten BashArena basic-memory basicmachines-co Basque Center on Cognition, Brain, and Language Basque language baulab Bayesian decision theory Bayesian Inference and Decision Audits for Public Archives of Frontier AI Evaluations Bayesian Multiobjective Optimization Bayesian Optimal Experimental Design Bayesian Optimization BayesPO BayLing-Duplex BBC News BBQ BBVA BC Protocol BCG X BCR BDD100K Be My Eyes Be My Tutor: On-Policy Co-Distillation for Mutual LLM Improvement via Peer Feedback BEA-Dialogue Beam Search Bebop Bechiri and Lanasri [2026] Becker Friedman Institute for Economics Before You Think: System 0, AI-Mediated Cognition and Cognitive Colonization behavior cloning behavior trees with learning-enabled components behavioral fine-tuning Behavioral Trajectory Tracking Framework behavioral-gradient validator Behavioral-SafetyBench Behavox BEIR Belebele Belief-Space Safety Filter (BeliefSF) BeliefTrack Ben Bernanke Ben Mann Ben Parr Ben Thompson Benchling Benchmark Agent Benchmark Everything Everywhere All at Once Benchmarking Multimodal Large Language Models for Scientific Visualization Literacy benchmaxxing BenCzechMark Bending Spoons Beneficial Deployments Benevity BenevolentAI Bengali Bengaluru BenHalluEval BenHalluScore Benjamin Chang Benjamin Krause Berkeley AI Research (BAIR) Berkeley Artificial Intelligence Research Berkshire Hathaway Berlin Database of Emotional Speech (EMO-DB) BerriAI BERT BERT-base BERT-F1 BERTopic BERTScore BERTurk BERTweet BES (trained models) Bessemer Venture Partners Best-of-N Best-of-N Sampling best@k Better Call GRPO Beyond a Single Judge: Simulating Social Persona Panels for Generative UI Evaluation Beyond Accuracy: Community Perspectives on Machine Translation Beyond Adam: SOAP and Muon for Faster, Label-Efficient Training of Machine Learning Interatomic Potentials Beyond Aggregate Risk: Role-Stratified Conformal Risk Control for LLM Tool Calls Beyond Attack-Success Rate: Action-Graded Severity Scale for Tool-Using AI Agents Beyond Benchmarks: Exposing the Hidden Crisis in Bangla Hate Speech Detection Beyond Fixed Representations: The Vocabulary and Verifier Gaps in Open-Ended AI Beyond Fully Random Masking: Attention-Guided Denoising and Optimization for Diffusion Language Models Beyond Function Calling: Benchmarking Tool-Using Agents under Tool-Environment Unreliability Beyond Global Replanning: Hierarchical Recovery for Cross-Device Agent Systems Beyond Independent Labels: Schwartz-Geometry Decoding for Human Value Detection Beyond Negative-Ridge Endpoints: Mixed-Sign Spectral Regularization via Negative-Shifted Gradient Descent Beyond Runtime Enforcement: Shield Synthesis as Defensibility Analysis for Adversarial Networks Beyond Scale and Generation: Understanding Language Model-based Entity Matching Beyond Sentiment: Structured Information Extraction from Financial News Beyond Success Rate: Cost-Aware Evaluation of Offensive and Defensive Security Agents Beyond Sufficiency: Time Series Explanation with Counterfactual Necessity Beyond Surface Forms: A Comprehensive, Mechanism-Oriented Taxonomy of Indirect Linguistic Encoding for LLM-Based Coded Language Detection Beyond Sycophancy: Structured Resistance and Compliance in LLM Moral Reasoning Beyond task performance: Decoding bioacoustic embeddings with speech features Beyond the Commitment Boundary: Probing Epiphenomenal Chain-of-Thought in Large Reasoning Models Beyond the Hard Budget: Sparsity Regularizers for More Interpretable Top-k Sparse Autoencoders Beyond the Leaderboard: Design Lessons for Trustworthy Multimodal VQA Beyond Third-Person Audits: Situated Interaction Auditing for User-Centered LLM Bias Research Beyond Tokenization: Direct Timestep Embedding and Contrastive Alignment for Time-Series Question Answering Beyond Uniform Tokens: Adaptive Compression for Time Series Language Models BeyondUncertainty BFCL BFCL Multi-Turn BFCL-V3 BFCLv3 BFW schedule hint BGE BGE-M3 BGL bi-manual robotic manipulation Bias Benchmark for Question Answering Bias Leaves a Gradient Trail Biconvex Optimization bidirectional attention Bidirectional Evolutionary Search Bielik Big Bench Big Bench Audio Big-Math BigBird BigCode BigCodeArena BigCodeBench Bigger or Cheaper? Scale and Quantization Effects on Uncertainty Signals in Vision-Language Models Under Image Degradation BigLaw Bench BigScience Bill & Melinda Gates Foundation Bill Dally Bill Gurley Bill McDermott Binarized Neural Network binary quantization BINEVAL Bing binning semiring bioacoustic monitoring bioacoustics BioASQ 14b BioBERT Biographical
any entity 1-Lipschitz Neural Networks on Hadamard Manifolds 1.58-bit quantization 10 USC 3252 10x Genomics 12-Factor App 1B-scale language models 1M-token context 1X 2-Parameter Logistic IRT Model 2026 FIFA World Cup 2026 State of AI Traffic and Cyberthreat Benchmark Report 2D-RoPE 2WikiMultiHopQA 32B language model 3C3H 3D asset generation 3D Gaussian Splatting 3D Medical Image Classification 3D optical flow 3D point cloud generation 3D tracking 3D-Aware VLMs with Implicit and Explicit Geometries 3D-Fit 3D-RoPE 3E Memory Module 3LM 404 Media 4chan 4D reconstruction 5PILS 6-PACK 6G Radio Access Network 7B language model 8B autoregressive language model A Blueprint for Equilibrium-Based Differentiable Continuous-Variable Thermodynamic Computing A Causal Model of Theory of Mind in Conflict for Artificial Intelligence A Comparative Study of Deep Learning Architectures for Multi-Horizon Behavioural Forecasting for Mobile Health A Computational Audit of Demographic Association Encoding in ClinicalBERT Language Predictions A Constrained Natural-Language Interface for Variational Multi-Physics Finite Element Simulations in FEniCS A Continual Validation, Updating, and Decision-Making Framework for Self-Adaptive Digital Twins via Robust Model Predictive Control A Diffusion Approximation for Temporal-Difference Learning with Linear Features under Markovian Noise A Factorial Study of Synthetic Data Generation for Low-Resource Machine Translation using Grammar Books A History-Aware Visually Grounded Critic for Computer Use Agents A Human-in-the-Loop Corpus for LLM-Based Simplification of Scientific Summaries A Komi-Yazva–Russian Parallel Corpus and Evaluation Protocol for Zero- and Few-Shot LLM Translation A Learning-Rate-Gated Failure of GRPO in a Small Language and Vision-Language Model Web Agent A Methodology for Auditable Trustworthiness Levels in AI Lifecycle Governance A Multi-Agent System for Autonomous, Fine-Tuning-Free Clinical Symptom Detection: Development and Validation Study A Multi-Domain Benchmark for Detecting AI-Generated Text-Rich Images from GPT-Image-2 A Physics-Informed Fourier-Wavelet Transformer for Multiscale Computational Fluid Dynamics Surrogate Modeling A Practical Investigation of Training-free Relaxed Speculative Decoding A Resource for Enthymeme Detection in Controversial Political Discourse A Self-Evolving Agent for Longitudinal Personal Health Management A sleep-like consolidation mechanism for LLMs A Study of Temporal Fusion Strategies for Named Entity Recognition in Historical Texts A Systematic Approach for Selecting Trajectories for Data Augmentation A Taxonomy of Conceptual Alignment in Human-Robot Dialogue A Training-Free Mixture-of-Agents Framework for Multi-Document Summarization using LLMs and Knowledge Graphs A Tree-of-Thoughts Inspired Hybrid Approach for Legal Case Judgement Summarization using LLMs A Unifying Lens on Supervised Fine-Tuning Through Target Distribution Design A24 A2A Protocol A2C A2Z A3C AA-Briefcase AA-Omniscience AAAI Aardvark AARRI-Bench AASIST AB-UPT Abanca ABBEL ABC-Bench Abeba Birhane abhigyanpatwari abjadai Ableton abliteration Abridge Abstraction Gap Accelerate Accelerated Decentralized Stochastic Gradient Descent for Strongly Convex Optimization Accelerating Masked Diffusion Large Language Models: A Survey of Efficient Inference Techniques Accenture Accenture Anthropic Business Group accumulated message effect Accuracy and Satisfaction in Multi-Turn LLM Dialogues for NFR Assessment ACE ACEBench-Agent AceReason-14B Achieving Precise Text-To-Cypher Via Grounded Knowledge Graph Data Generation ACKTR ACL ACL Anthology ACL-Verbatim Acoustic Cue Alignment in Audio Language Models for Speech Emotion Recognition ACROS Act2Answer Action Chunking Transformer ACTION-BED Action-BED: Task-Driven Bayesian Experimental Design with Singly Intractable Objectives Action-Dependent Factorized Baselines Actionable Activation Directions for Detecting and Mitigating Emergent Misalignment Across Language Model Families Activation Atlases activation capping Activation Oracles activation patching Activation Steering active learning Active Listening Active Offline-to-Online Reinforcement Learning ActiveSAM ActiveVision ACTS Ada AdaCodec AdaFlash AdaGrad AdaJEPA AdaJEPA: An Adaptive Latent World Model Adalat AI AdaLook Adam AdamO AdamW AdaPrefix-GRPO ADAPT-GQE ADAPT-VQE adapter fine-tuning Adaptive Clip Policy Optimization Adaptive Data Scheduling Adaptive Depth Sparse Framework Adaptive Gated Feedback Optimization Adaptive Multi-Resolution Procedural Knowledge Compression for Large Language Models Adaptive Multi-Step Lookahead Decoding for Diffusion Language Models Adaptive Parallel Reasoning Adaptive Self-Debiasing adaptive thinking Adaptive Turn-Taking for Real-time Multi-Party Voice Agents adaptive-rank instantiation ADAS AdaSR ADEME ADHD Adithya S K Aditi Krishnapriyan Aditya (co-lead author) Aditya G. Parameswaran Adjacent Contrastive Reasoning ADK Python adk-samples Adobe Adobe Creative Cloud Advanced AI Scaling Framework Advanced Light Source Advanced Matrix Extensions (AMX) Advanced Photon Source Advanced Voice Mode AdvancedMathBench Advancing Omnimodal Embodied Agents from Isolated Skills to Everyday Physical Autonomy AdvBench AdversaBench Adversarial Attacks on Neural Network Policies Adversarial Creation and Detection of AI-Generated Social Bot Content adversarial examples adversarial pragmatics Adversarial Pragmatics for AI Safety Evaluation: A Benchmark for Instruction Conflict, Embedded Commands, and Policy Ambiguity adversarial refinement adversarial robustness adversarial training AdvGRPO Aeneas Aerial Image Dataset (AID) Affinity by Canva AFM 3 AFM 3 Cloud Pro AFM 3 Core Advanced AfriSenti AG-MG Parallel Corpus AG-UI AG-UI Protocol AGC-Bench AGC-Judge Agda AGDO Age of Empires II age prediction Agence France-Presse Agent 365 agent cloud Agent Cognitive Redundancy Ratio Agent Communication Protocol Agent Development Kit agent ecologies Agent Governance Toolkit agent harness Agent JIT Compilation Agent Reasoning Evaluation (ARE) Agent Runtime agent sandboxing agent-device agent-native Agent-Native Immune System Agent-Native Immune System: Architecture, Taxonomy, and Engineering agent-native telemetry Agent-Reach Agent-S agent-skills agent-task efficiency agent-teams-ai agent-to-agent evaluation protocol Agent-to-Agent Protocol (A2A) agent-toolkit-for-aws AgentBeats AgentBeats: Agentifying Agent Assessment for Openness, Standardization, and Reproducibility AgentBoard AgentCL AgentDojo Agentforce agentic AI Agentic AI Foundation Agentic AI Pipelines Agentic AI Systems Agentic Chain-of-Thought Steering Agentic Chain-of-Thought Steering for Efficient and Controllable LLM Reasoning Agentic CLEAR agentic coding Agentic Commerce Protocol Agentic Environment Engineering for Large Language Models: A Survey of Environment Modeling, Synthesis, Evaluation, and Application Agentic generation of verifiable rules for deterministic, self-expanding reaction classification Agentic Permissions Policy Algebra for Taint Confinement in LLM Agents agentic proving Agentic RAG agentic re-optimization framework Agentic Resource Discovery Agentic RL Agentic Self-Instruct Agentic System Monitoring Methodology Agentic Technical Debt agentic vulnerabilities AgenticRL AgenticSTS AgentKit AgentMap agentmemory AgentMob Agentopia Agentopia: Long-Term Life Simulation and Learning in Agent Societies Agents in the Wild: Where Research Meets Deployment Agents-A1 agents-cli Agents-K1 Agents.js AGENTS.md Agents' Last Exam AgentScope AGENTSERVESIM agentskills.io AgentSnare AgentSpec AgentsView AgentWorldBench AGI AGI (Artificial General Intelligence) AGI cognitive framework AGI economy Agility Robotics Agon AGORA agreement attraction AHA-WAM Ahead-of-Time Compilation Ahmad Osman AI Agents AI agents that matter AI alignment AI and Compute AI Andrew AI as Normal Technology AI biosecurity risk assessment AI Clinical Copilot AI Co-Clinician AI control AI Control Roadmap AI Cybersecurity Threat Evaluation Framework AI data sovereignty AI Developer Conference 2025 AI Engineer NYC AI Engineer World's Fair AI Existential Risk AI Exposure Scores: what they measure, what they miss, and what comes next AI for Citizens AI for Game Development AI for Math Initiative AI for Science AI for Science program AI Fund AI Healthcare Chatbots as Information Infrastructure: A Large-Scale Study of User-Reported Breakdowns AI image verification AI inference carbon emissions AI leaderboards AI Liability Directive AI Literacy & Creator Collective AI Mathematical Olympiad AI Overview AI Panic Blog AI Persuasive Framing in Collective Dilemmas AI Registry AI Reproducibility Benchmark AI Risk Management Framework AI Safety Fund AI Safety Level (ASL) AI Safety Level Standards AI Safety via Debate AI scaling laws AI Sheets AI Snake Oil AI systems and the reproduction of (standard) language ideologies in World Englishes AI translation of literary texts is "fine", but readers still prefer human translations AI Verify Foundation AI vs. AI AI Watermark Evidence Fails Forensic Readiness: An Empirical Evaluation AI-accelerated End-to-End Framework for Rapid Professional Upskilling ai-agent-book AI-assisted human evaluation AI-assisted red teaming AI-Assisted Systematization for Evaluating GenAI Systems AI-assisted theorem proving ai-boost/awesome-harness-engineering AI-driven constraint reasoning AI-Driven Life Cycle (AI-DLC) AI-Engineering-Coach ai-hedge-fund AI-MO AI-PAVE-Br AI-Trader AI's Capability in Assisting Scientific Research in Physics, Astrophysics, and Cosmology II: Project Planning and Proposal Evaluation AI@HHMI AI21 Labs AIA Labs AIcrowd Aider AIDev dataset aidlc-workflows Aiera AIEWF AILuminate AIME AIME 2025 AIME 2026 AIME2024 AIME24 AIME25 AIME26 aiming-lab AIMO Interpretability Challenge AIMO Progress Prize AIMS AINews AionUi Air India AIR: Adaptive Interleaved Reasoning with Code in MLLMs Aira Explorer AiraXiv Airbus AirflowAttack AIriskEval-edu Demo AISPA AISPA: User-Centric System Prompt Auditing for Large Language Model Applications aisuite Akshat Bubna Akshay Nathan Albert Gu Alberta AI Academy Alberta Machine Intelligence Institute Alberti ALE-Bench ALER-TI Aletheia Alex Hernandez Alex Kearns-Apuya Alex L. Zhang Alex Lupsasca Alex Rives Alexa+ Alexandr Wang AlexNet ALFRED Alfred Lin ALFWorld Algonauts 2025 Algorithmic and Minimax Complexities in Kernel Bandits Algorithmic Impermeability algorithmic monoculture Algorithmic Monocultures in Hiring Ali Essam Ghareeb Alibaba Alibaba Cloud Alibaba Cloud Model Studio Alibaba DAMO Academy Alibaba Qwen Alibaba Qwen Team ALIGN AlignAtt AlignAtt4LLM ALIGNBEAM alignment auditing alignment faking Alignment Research Center alignment tampering alignment tax Alireza Rezvani All Hands AI All Smoke, No Alarm: Oracle Signals in Agent-Authored Test Code All Tech Is Human all-edge F1 Allam Allegro Allegro Hand Allen Institute Allen Institute for AI AllenAI Allied for Startups ALMANAC ALOHA AlpacaEval 2 Alpamayo R1 Alphabet AlphaEarth Foundations AlphaEvolve AlphaFold AlphaFold2 AlphaFold3 AlphaGenome AlphaGo AlphaOracle AlphaStar AlphaTensor alpic-ai Altana Altera Alternating Token-Weighted Unlearning Altimeter Capital Always-On Evaluation Protocol (AOEP-v0) Always-OnAgents: A Survey of Persistent Memory, State, and Governance in LLM Agents ALX Alyah AM-Parser AMALIA Amar Subramanya AMARIS Amazon Amazon Beauty Amazon Bedrock Amazon Bedrock AgentCore Amazon Graviton Amazon Reviews Dataset Amazon SageMaker Amazon SageMaker Studio Amazon Trainium2 Amazon Web Services Ambient Diffusion Policy AMC12 AMD AMD Instinct AMD Instinct MI300 AMD MI300 AMD MI355X AMEL American Federation of Teachers American Technology Council Ami Vora AMIA Amir Bar amortized variational inference AMP Amrith Setlur AMRS aMUSEd An Agency-Transferring Model-Free Policy Enhancement Technique An Empirical Analysis of Factual Errors in Human-Written Text and its Application An Evaluation Framework for Structured Audio Captions Validated by Controlled Perturbations An Exact Instrument for State Usage in Selective State-Space Models, and the Input-Driven Migration It Reveals Analogical Deep Research Analogical Deep Research: Retrieving and Integrating Historical Analogies for Foresight Analysis Analysing Self-Harm Representations in Language Models: a Cross-Architecture Study AnalyticGeo7K Analytics-Everywhere-Lab Anatomy of Post-Training: Using Interpretability to Characterize Data and Shape the Learning Signal ANCESTRA anchoring effect Andon Labs Andreessen Horowitz Andrej Karpathy Andrew Kelley Andrew Ng Andrew Qu Andrew Wagenmaker Anduril Industries Andy Beam Andy Jassy Andyyyy64 Angelini Pharma AngelSpec Anglophone retrieval bias Angluin's condition anisotropic BAO damping Anjney Midha Anna's Archive Annapurna Labs annotator disagreement anomalyco AnomalyShapeNet ANSI National Accreditation Board AnswerSheet ANSYR ANTAP Anthology Fund Anthony Albanese Anthropic Anthropic Academy Anthropic Advanced AI Framework Anthropic Agent SDK Anthropic Beneficial Deployments Anthropic Claude Code Anthropic Economic Advisory Council Anthropic Economic Futures Programme Anthropic Economic Futures Research Fund Anthropic Economic Index Anthropic Economic Policy Framework Anthropic Fable 5 Anthropic for Startups Anthropic Interviewer Anthropic Korea Anthropic Labs Anthropic Long-Term Benefit Trust Anthropic Mythos Anthropic National Security and Public Sector Advisory Council Anthropic Partner Academy Anthropic Policy Frontier Red Team Anthropic Public Record Anthropic Responsible Scaling Policy Anthropic Safeguards Team Anthropic Terms of Service Anthropic Threat Intelligence Report August 2025 Anthropic Usage Policy Anthropic v. Department of War anthropics/skills anthropomorphism in AI Anti-Drift Rectification Anti-Periodic Positional Encoding: Möbius Boundary Conditions Make In-Context Retrieval Reliable Antigravity Antoine Zambelli Anton Korinek Any-Dimensional Learning by Sampling AnyGroundBench AnyLanguageModel AnyMo Anysphere Anytime-Valid E-Process anytime-valid sequential testing AP2 Apache 2.0 Apache Arrow Apache Ossie Apache Software Foundation Apache Spark Apart Research APE Apertus 70B Apertus-8B-Instruct-2509 APEX-Accounting APEX-Accounting APEX-Agents-AA aphasia API knowledge boundary probing APO: Unsupervised Atomic Policy Optimization for 3D Structure Prediction of Atomic Systems Apollo Apollo Global Management Apollo Research APPA (Agentic Permissions Policy Algebra) Apparent Psychological Profiles of Large Language Models are Largely a Measurement Artifact Appearance Pointers -- Multimodal Region Control of Diffusion Transformers Appia Foundation Apple Apple Foundation Models 3 Apple Foundation Models Framework Apple Health Apple Neural Engine Apple Silicon Applied Intuition apply_patch APPO: Agentic Procedural Policy Optimization Approximate DP APPS Apps SDK AppWorld Apriel-H1 AprielGuard APS-Bench APS-RAG apurvsinghgautam AR military headset Ar9av AraBERT Arabic Arabic Instruction Following Eval (IFEval) Arabic LLMs ArabiGEE AraGen ARC Evals ARC Framework Arc Institute ARC Prize Foundation Arc Virtual Cell Challenge ARC-AGI Archon Arco Education ARDY Are Performance-Optimization Benchmarks Reliably Measuring Coding Agents? Are We Ready For An Agent-Native Memory System? AREA Arena AI Arena Code Arena Leaderboard Arena Search Arena-Hard Arena.ai Code Arena WebDev ArenaHard Argentina Argilla Argilla 2.0 Argonne National Laboratory Argus ariadng ARIS Aristotle API Arithmetic Pedagogy for Language Models Arize AI Arize Phoenix Arm Arm AGI CPU Armin Ronacher ArMIS Army Special Operations Command Arnuv Tandon ArogyaBodha ArogyaSutra ARPA-H ARROW ART (Agent Reinforcement Trainer) Arthur Mensch Artifact Artifacts Artificial Analysis Artificial Analysis Big Bench Audio Artificial Analysis Coding Agent Index Artificial Analysis Conversational Dynamics Artificial Analysis Intelligence Index Artificial Analysis Intelligence Leaderboard Artificial Analysis LLM Performance Leaderboard Artificial Analysis Text to Image Artificial Analysis Text to Image Leaderboard Artificial Bee Colony Artificial Epanorthosis: Why large language models overuse a classical rhetorical figure, and how to mitigate it ArXiv arXiv:2602.05394 ASAM ASAP 2.0 ASAP++ Ascend SuperPOD Ashwyn Sharma Asia-Pacific ASIMOV-2.0 AskChem AskChem-Bench AskDocs ASL-2 ASL-3 ASML ASML Holding NV Aspect-Based Sentiment Analysis ASRD AssemblyAI AssemblyAI Universal Assessing the Operational Impact of Poisoning Attacks over Augmented 3D Point Cloud Public Datasets for Connected and Autonomous Vehicles AssetOpsBench assistant axis Assistants API Assisted Generation ASSISTments Association for Computational Linguistics Associative Recurrent Memory Transformer Astera Institute ASTRA Astral AstrBot AstrBotDevs ASVspoof 5 asynchronous inference Asynchronous Noise Schedule Atari ATDev ATE-Bench Atla Atlas Atlas H&E-TME ATLAS Open Data 13 TeV ATLAS: Active Theory Learning for Automated Science Atlassian AToken Atomic Policy Optimization ATOMIC2020 Atoms of Thought (EEG Microstates paper) Attack Success Rate Attainable Utility Preservation Attention Amnesia in Hybrid LLMs: When CoT Fine-Tuning Breaks Long-Range Recall, and How to Fix It attention entropy Attention Expansion: Enhancing Keyphrase Extraction from Long Documents with Attention-Augmented Contextualized Embeddings attention head circuit Attention Residuals Attention Sorting Attention submodule Attention-based MIL Attractor States Emerge in Multi-Turn LLM Conversations AUC Audex Audio Interaction Model Audio MultiChallenge Audio-Flamingo-3 Audio-LLM Audio-Native Speech Recognition with a Frozen Discrete-Diffusion Language Model AudioCards AudioDER AudioLDM 2 AudioVAE2 Auditing Protocol-Level Shortcuts in Large Audio Language Model Judges for Speech Evaluation AUDITS Augment Augment Code AUPRC AuRA AUROC Australia AI Safety Institute Australia National AI Plan Australian Emergency Department triage notes Australian Government Australian National University Auto Benchmark Audit (ABA) Autodata Autodesk Fusion AutoDex autoDream AutoForest Autoformer AutoGPTQ AutoJourn AUTokenizer AutoLab automated AI research Automated Background Swapping Automated Compliance Mapping in Cloud Security with Domain-Adapted Sentence Transformers Automated Discovery Has No Universally Superior Harness Automated Extraction of Techno-Economic Data from 76,000 Energy System Studies Automated grading of Linux/bash examinations using large language models: a four-level cognitive taxonomy approach automated mechanistic interpretability automated red teaming Automated Reference Verification System Automated reproducibility assessments in the social and behavioral sciences using large language models automated test suite automated theorem proving Automatic Domain Randomization Automatic Post-Editing (APE) Automatic Speech Recognition AutomationBench-AA AutoMem autonomous driving AUTOPILOT-VQA Autoregressive Boltzmann Generators autoregressive transformer autoresearch AutoResearchClaw AutoRound AutoSkillHarm AutoSynthesis avatarin Avenger Labs Aviral Kumar awesome-claude-code awesome-copilot awesome-llm-apps AWQ AWS AWS Bedrock AWS GovCloud AWS Inferentia2 AWS Labs AWS Marketplace AWS Neuron AWS Neuron SDK AWS Trainium AxDafny Axel Backlund Axiom Math Axios Axolotl AXPO Aya Aya Expanse Aya Vision Aya-Expanse-8B-Base Azure Azure AI Studio Azure DevOps Azure Foundry Azure OCR B³D-RWKV BabelJudge BABI-Long BABILong BabyCL Backbone-as-Architect backdoor attack Backdoor Circuit Analysis (Language-Switching) Backdoor Unlearning Generalization: A Path Toward the Removal of Unknown Triggers in LLMs backnotprop backpropagation through time Backstory Bad Grammar Bad Scientist BadWAM BadWAM: When World-Action Models Dream Right but Act Wrong Bag of Words (BoW) BAGEL BAGEL-7B Baidu Baidu Translate Baifeng Shi Baiyu Chen ball balancing ball balancing task Bamba BamiBERT Bandit Bang-v3 BanglaBERT Banque des Territoires Barcelona Supercomputing Center Language Technologies bare metal sandboxes Bargaining Scenarios Dataset Bark Base Baseline Agent Baseline-Log Physical Separation Baseten BashArena basic-memory basicmachines-co Basque Center on Cognition, Brain, and Language Basque language baulab Bayesian decision theory Bayesian Inference and Decision Audits for Public Archives of Frontier AI Evaluations Bayesian Multiobjective Optimization Bayesian Optimal Experimental Design Bayesian Optimization BayesPO BayLing-Duplex BBC News BBQ BBVA BC Protocol BCG X BCR BDD100K Be My Eyes Be My Tutor: On-Policy Co-Distillation for Mutual LLM Improvement via Peer Feedback BEA-Dialogue Beam Search Bebop Bechiri and Lanasri [2026] Becker Friedman Institute for Economics Before You Think: System 0, AI-Mediated Cognition and Cognitive Colonization behavior cloning behavior trees with learning-enabled components behavioral fine-tuning Behavioral Trajectory Tracking Framework behavioral-gradient validator Behavioral-SafetyBench Behavox BEIR Belebele Belief-Space Safety Filter (BeliefSF) BeliefTrack Ben Bernanke Ben Mann Ben Parr Ben Thompson Benchling Benchmark Agent Benchmark Everything Everywhere All at Once Benchmarking Multimodal Large Language Models for Scientific Visualization Literacy benchmaxxing BenCzechMark Bending Spoons Beneficial Deployments Benevity BenevolentAI Bengali Bengaluru BenHalluEval BenHalluScore Benjamin Chang Benjamin Krause Berkeley AI Research (BAIR) Berkeley Artificial Intelligence Research Berkshire Hathaway Berlin Database of Emotional Speech (EMO-DB) BerriAI BERT BERT-base BERT-F1 BERTopic BERTScore BERTurk BERTweet BES (trained models) Bessemer Venture Partners Best-of-N Best-of-N Sampling best@k Better Call GRPO Beyond a Single Judge: Simulating Social Persona Panels for Generative UI Evaluation Beyond Accuracy: Community Perspectives on Machine Translation Beyond Adam: SOAP and Muon for Faster, Label-Efficient Training of Machine Learning Interatomic Potentials Beyond Aggregate Risk: Role-Stratified Conformal Risk Control for LLM Tool Calls Beyond Attack-Success Rate: Action-Graded Severity Scale for Tool-Using AI Agents Beyond Benchmarks: Exposing the Hidden Crisis in Bangla Hate Speech Detection Beyond Fixed Representations: The Vocabulary and Verifier Gaps in Open-Ended AI Beyond Fully Random Masking: Attention-Guided Denoising and Optimization for Diffusion Language Models Beyond Function Calling: Benchmarking Tool-Using Agents under Tool-Environment Unreliability Beyond Global Replanning: Hierarchical Recovery for Cross-Device Agent Systems Beyond Independent Labels: Schwartz-Geometry Decoding for Human Value Detection Beyond Negative-Ridge Endpoints: Mixed-Sign Spectral Regularization via Negative-Shifted Gradient Descent Beyond Runtime Enforcement: Shield Synthesis as Defensibility Analysis for Adversarial Networks Beyond Scale and Generation: Understanding Language Model-based Entity Matching Beyond Sentiment: Structured Information Extraction from Financial News Beyond Success Rate: Cost-Aware Evaluation of Offensive and Defensive Security Agents Beyond Sufficiency: Time Series Explanation with Counterfactual Necessity Beyond Surface Forms: A Comprehensive, Mechanism-Oriented Taxonomy of Indirect Linguistic Encoding for LLM-Based Coded Language Detection Beyond Sycophancy: Structured Resistance and Compliance in LLM Moral Reasoning Beyond task performance: Decoding bioacoustic embeddings with speech features Beyond the Commitment Boundary: Probing Epiphenomenal Chain-of-Thought in Large Reasoning Models Beyond the Hard Budget: Sparsity Regularizers for More Interpretable Top-k Sparse Autoencoders Beyond the Leaderboard: Design Lessons for Trustworthy Multimodal VQA Beyond Third-Person Audits: Situated Interaction Auditing for User-Centered LLM Bias Research Beyond Tokenization: Direct Timestep Embedding and Contrastive Alignment for Time-Series Question Answering Beyond Uniform Tokens: Adaptive Compression for Time Series Language Models BeyondUncertainty BFCL BFCL Multi-Turn BFCL-V3 BFCLv3 BFW schedule hint BGE BGE-M3 BGL bi-manual robotic manipulation Bias Benchmark for Question Answering Bias Leaves a Gradient Trail Biconvex Optimization bidirectional attention Bidirectional Evolutionary Search Bielik Big Bench Big Bench Audio Big-Math BigBird BigCode BigCodeArena BigCodeBench Bigger or Cheaper? Scale and Quantization Effects on Uncertainty Signals in Vision-Language Models Under Image Degradation BigLaw Bench BigScience Bill & Melinda Gates Foundation Bill Dally Bill Gurley Bill McDermott Binarized Neural Network binary quantization BINEVAL Bing binning semiring bioacoustic monitoring bioacoustics BioASQ 14b BioBERT Biographical