Skip to content Get the almanac in your inbox © 2026 Call Me Almanac · Content is AI-generated and may be inaccurate.
Significance High 8–10 Notable 6–7 Minor 4–5
4 Simon Willison'S Weblog · 26h ago · source ↗ Simon Willison published an early alpha release (0.1a0) of llm-mcp-client, a plugin that adds Model Context Protocol client support to his LLM command-line tool. The release extends the LLM ecosystem to interoperate with MCP servers, enabling tool use via the standardized protocol. This is an early-stage but concrete addition to the growing MCP tooling ecosystem.
5 Simon Willison'S Weblog · 26h ago · source ↗ Simon Willison writes about renewed interest in the Model Context Protocol following developments around stateless MCP, and announces two new tools: mcp-explorer and datasette-mcp. The post reflects on how stateless operation changes the practical appeal of MCP for tool integration. This is a practitioner-level signal about MCP adoption patterns and tooling ecosystem growth.
Figma has published an official guide repository for using the Figma MCP (Model Context Protocol) server, written in Python. The repository has accumulated 1,831 stars, indicating meaningful developer interest. This represents Figma's participation in the growing MCP ecosystem for tool integration with AI agents.
AskChem is a new retrieval infrastructure that reframes chemistry literature search around atomic, provenance-carrying claims rather than document rankings, indexing 2.4M claims from 147K papers. Each claim is grounded with a source DOI and verbatim quote, and the system exposes a faceted taxonomy, evidence graph, and living taxonomy for synthesis. The system provides REST, SDK, and MCP access for AI agents, and benchmarking shows that grounding GPT-5.5 in AskChem yields 100% resolvable DOIs versus 88.3% without retrieval. The work is relevant both as a domain-specific RAG infrastructure paper and as an example of MCP-native scientific tooling for AI agents.
5 Simon Willison'S Weblog · 3d ago · source ↗ Simon Willison documents a practical walkthrough for configuring a custom MCP (Model Context Protocol) server with both Claude and ChatGPT. The post covers the concrete steps required to integrate a self-hosted MCP server into two major AI assistant platforms. This is a practitioner-level guide relevant to the growing MCP ecosystem and cross-platform tool-use patterns.
AG-UI (Agent-User Interaction Protocol) is an open-source TypeScript project aimed at standardizing how AI agents integrate with frontend applications. The repository has accumulated 15,007 stars with 48 added today, indicating sustained community interest. It represents an emerging protocol layer for connecting agent backends to user-facing interfaces, complementary to backend-focused protocols like MCP.
jcodemunch-mcp is an open-source MCP server that uses tree-sitter AST parsing to enable precise, symbol-level GitHub code retrieval, claiming 95%+ token cost reduction for code exploration tasks. The tool is compatible with Claude Code, Cursor, and any MCP client. It has accumulated 2,225 GitHub stars with 313B+ tokens reportedly saved across users.
A Python-based MCP server implementation exposing UniFi's suite of applications (Network, Protect, Access, Drive) as tool-callable endpoints. The project enables AI agents to interact with UniFi network infrastructure via the Model Context Protocol. It has 581 GitHub stars with modest daily growth, indicating niche but real adoption.
SurfSense is an open-source Python project positioning itself as a NotebookLM alternative, enabling research across live web sources including Reddit, YouTube, Instagram, TikTok, Google Search, and Maps. It exposes functionality through a platform UI, API, or MCP server. The project has accumulated 15,456 GitHub stars with modest daily momentum (+35 today).
Serena is an open-source Python toolkit that provides an MCP-compatible interface for semantic code retrieval and editing, positioning itself as an IDE layer for AI coding agents. The project has accumulated 26,793 GitHub stars with 85 added today, indicating sustained community traction. It is relevant to the growing ecosystem of MCP-native tooling for agentic software development workflows.
5 Claude Code Release Notes · Jul 21, 2026 · source ↗ Anthropic released Claude Code version 2.1.217 with a notable new cap on concurrently-running subagents (default 20, configurable via CLAUDE_CODE_MAX_CONCURRENT_SUBAGENTS) and a change preventing nested subagent spawning by default. The release also fixes a memory leak in MCP tool output handling, resolves auto-compact failures for Claude Opus 4.8 on Bedrock, and patches several Windows, corporate proxy/TLS, and session isolation bugs. The subagent concurrency and budget enforcement changes are architecturally significant for agentic workflows, preventing unbounded fan-out.
FastMCP, a Python library by PrefectHQ for building Model Context Protocol servers and clients, is trending on GitHub with 26,434 total stars and 77 new stars today. The project positions itself as a high-level, Pythonic interface for MCP development. Its traction signals continued ecosystem growth around the MCP standard for AI tool integration.
SigNoz is an open-source observability platform (30k+ GitHub stars, +343 today) that provides logs, metrics, and traces with APM and distributed tracing capabilities. It has added an MCP server integration and a native AI teammate feature in its cloud offering, positioning it for AI agent observability use cases. The trending activity suggests growing adoption in the AI/ML infrastructure space.
Wigolo is a TypeScript open-source tool providing local-first web search, fetch, crawl, and research capabilities for AI coding agents over the Model Context Protocol (MCP). It requires no API keys and incurs no per-query cost, positioning itself as a zero-cost alternative to cloud-based search APIs. The project gained 192 stars in a single day during its public beta, suggesting meaningful community traction.
A multi-item digest covers five significant AI developments: Apple sued OpenAI alleging trade secret theft via former employees including hardware chief Tang Tan; Meta released Muse Spark 1.1, a multimodal agentic model with 1M-token context and strong tool-use capabilities; OpenAI launched ChatGPT Work, a cloud-based workplace agent competing with Anthropic's Claude Cowork; IBM released CodeAlchemy, a 500B+ token synthetic code dataset with execution traces showing smaller models trained on it outperform those trained on much larger real-code corpora; and OpenAI shut down its Atlas browser in favor of a Chrome extension and desktop integration. These items collectively reflect intensifying competition across agentic products, synthetic data strategies, and legal disputes between major AI players.
mcp-use is a TypeScript framework on GitHub for developing MCP (Model Context Protocol) applications targeting ChatGPT and Claude, as well as MCP servers for AI agents. The project has accumulated over 10,000 stars, indicating meaningful community adoption. It represents a tooling layer in the growing MCP ecosystem for agent-tool integration.
4 Claude Code Release Notes · Jul 15, 2026 · source ↗ Anthropic released Claude Code version 2.1.210, a substantial patch addressing over 30 bugs and adding several improvements to the agentic coding tool. Notable fixes include hardening against indirect prompt injection via subagent-read content, correcting git worktree isolation for subagents, fixing hook callback timeouts that caused unattended sessions to stall, and resolving MCP server teardown during mid-session re-syncs. The release also improves auto-mode permission classification by defaulting to Sonnet 5 for external sessions and adds UX enhancements for the multi-agent dashboard.
Vexa is an open-source Python project providing a meeting transcription API for Google Meet, Microsoft Teams, and Zoom, featuring auto-join bots and real-time WebSocket transcripts. It includes an MCP server integration, making it usable as a tool source for AI agents. The project is available for self-hosting or as a hosted SaaS and has accumulated 2,502 GitHub stars with 72 added today.
memtrace-public is an open-source Python library providing structural memory for AI coding agents via a bi-temporal graph architecture. It is MCP-native and operates without LLM calls, targeting integrations with Cursor, Claude Code, Codex, and VS Code. The project is early-stage with 407 stars and modest daily traction.
4 Claude Code Release Notes · Jul 10, 2026 · source ↗ Anthropic released Claude Code version 2.1.206 with a range of new features and bug fixes. Notable additions include background agent auto-upgrade after Claude Code updates, /login support for Anthropic-operated public gateway endpoints, /doctor checks for trimming CLAUDE.md files, and auto-allow for git push to configured remotes in /commit-push-pr. The release also fixes numerous bugs including MCP server timeout handling, OAuth re-authentication loops, model picker pricing display errors, and a multi-minute Bedrock startup hang on restricted networks.
Mistral has launched version control and system-of-record capabilities for prompts and skills within its Studio platform, targeting enterprise AI governance needs. The feature provides immutable versioning, ownership tracking, audit logs, rollback, and CI/CD integration, allowing non-developer domain experts to iterate on production prompts without engineering bottlenecks. Skills are exposed as MCP servers directly from Studio, connecting governed assets to runtime execution. The release addresses a common enterprise pain point: unmanaged, scattered prompt assets that create compliance and traceability risks.
Meta Superintelligence Labs has released Muse Spark 1.1, a significant upgrade to Muse Spark featuring a 1-million-token context window, strong agentic and computer-use capabilities, and major coding improvements on complex codebases. The model supports multi-agent orchestration, zero-shot generalization to MCP servers and custom tools, and multimodal reasoning including visual-to-code generation and video understanding. Alongside the model release, Meta is launching a public preview of the Meta Model API, giving developers programmatic access for the first time. Safety evaluations were conducted under Meta's Advanced AI Scaling Framework across frontier risk categories.
Notion has published an official Model Context Protocol (MCP) server implementation in TypeScript, enabling AI agents and tools to interact with Notion workspaces via the MCP standard. The repository has accumulated 4,491 stars on GitHub. This represents a major productivity platform adopting MCP as its AI integration layer.
A Python-based Model Context Protocol (MCP) server for Google Analytics has appeared on GitHub trending, published under the googleanalytics organization. The repository has accumulated 2,606 stars with modest daily growth (+14). This represents an official or semi-official Google Analytics integration point for AI agents and tools using the MCP standard.
5 Claude Code Release Notes · Jul 8, 2026 · source ↗ Anthropic released Claude Code version 2.1.203, a substantial patch addressing over 25 bugs primarily affecting background agent sessions, subagent orchestration, and worktree isolation. Key fixes include automatic recovery when daemon session tokens go stale, correct subagent state carry-over when returning to sessions, and PATH/environment variable inheritance issues on Windows. The release also adds MCP roots/list integration for working directories and improves streaming responsiveness.
Simon Willison
Claude
MCP
oraios
OpenTelemetry
MCP
Vexa-ai
any topic Agent and Tool Ecosystem AI Safety Research Alignment and RLHF Enterprise Deployment Patterns Evaluation and Benchmarking Frontier Model Releases Inference Economics Long Context Evolution Multimodal Progress Open Weights Progress Regulatory Developments Training Infrastructure
any topic Agent and Tool Ecosystem AI Safety Research Alignment and RLHF Enterprise Deployment Patterns Evaluation and Benchmarking Frontier Model Releases Inference Economics Long Context Evolution Multimodal Progress Open Weights Progress Regulatory Developments Training Infrastructure
any entity 1-Lipschitz Neural Networks on Hadamard Manifolds 1.58-bit quantization 10 USC 3252 10x Genomics 12-Factor App 1B-scale language models 1M-token context 1X 2-Parameter Logistic IRT Model 2026 FIFA World Cup 2026 State of AI Traffic and Cyberthreat Benchmark Report 2D-RoPE 2WikiMultiHopQA 32B language model 3C3H 3D asset generation 3D Gaussian Splatting 3D Medical Image Classification 3D optical flow 3D point cloud generation 3D tracking 3D-Aware VLMs with Implicit and Explicit Geometries 3D-Fit 3D-RoPE 3E Memory Module 3LM 404 Media 4chan 4D reconstruction 5PILS 6-PACK 6G Radio Access Network 7B language model 8B autoregressive language model A Blueprint for Equilibrium-Based Differentiable Continuous-Variable Thermodynamic Computing A Causal Model of Theory of Mind in Conflict for Artificial Intelligence A Comparative Study of Deep Learning Architectures for Multi-Horizon Behavioural Forecasting for Mobile Health A Computational Audit of Demographic Association Encoding in ClinicalBERT Language Predictions A Constrained Natural-Language Interface for Variational Multi-Physics Finite Element Simulations in FEniCS A Continual Validation, Updating, and Decision-Making Framework for Self-Adaptive Digital Twins via Robust Model Predictive Control A Diffusion Approximation for Temporal-Difference Learning with Linear Features under Markovian Noise A Factorial Study of Synthetic Data Generation for Low-Resource Machine Translation using Grammar Books A History-Aware Visually Grounded Critic for Computer Use Agents A Human-in-the-Loop Corpus for LLM-Based Simplification of Scientific Summaries A Komi-Yazva–Russian Parallel Corpus and Evaluation Protocol for Zero- and Few-Shot LLM Translation A Learning-Rate-Gated Failure of GRPO in a Small Language and Vision-Language Model Web Agent A Methodology for Auditable Trustworthiness Levels in AI Lifecycle Governance A Multi-Agent System for Autonomous, Fine-Tuning-Free Clinical Symptom Detection: Development and Validation Study A Multi-Domain Benchmark for Detecting AI-Generated Text-Rich Images from GPT-Image-2 A Physics-Informed Fourier-Wavelet Transformer for Multiscale Computational Fluid Dynamics Surrogate Modeling A Practical Investigation of Training-free Relaxed Speculative Decoding A Resource for Enthymeme Detection in Controversial Political Discourse A Self-Evolving Agent for Longitudinal Personal Health Management A sleep-like consolidation mechanism for LLMs A Study of Temporal Fusion Strategies for Named Entity Recognition in Historical Texts A Systematic Approach for Selecting Trajectories for Data Augmentation A Taxonomy of Conceptual Alignment in Human-Robot Dialogue A Training-Free Mixture-of-Agents Framework for Multi-Document Summarization using LLMs and Knowledge Graphs A Tree-of-Thoughts Inspired Hybrid Approach for Legal Case Judgement Summarization using LLMs A Unifying Lens on Supervised Fine-Tuning Through Target Distribution Design A24 A2A Protocol A2C A2Z A3C AA-Briefcase AA-Omniscience AAAI Aardvark AARRI-Bench AASIST AB-UPT Abanca ABBEL ABC-Bench Abeba Birhane abhigyanpatwari abjadai Ableton abliteration Abridge Abstraction Gap Accelerate Accelerated Decentralized Stochastic Gradient Descent for Strongly Convex Optimization Accelerating Masked Diffusion Large Language Models: A Survey of Efficient Inference Techniques Accenture Accenture Anthropic Business Group accumulated message effect Accuracy and Satisfaction in Multi-Turn LLM Dialogues for NFR Assessment ACE ACEBench-Agent AceReason-14B Achieving Precise Text-To-Cypher Via Grounded Knowledge Graph Data Generation ACKTR ACL ACL Anthology ACL-Verbatim Acoustic Cue Alignment in Audio Language Models for Speech Emotion Recognition ACROS Act2Answer Action Chunking Transformer ACTION-BED Action-BED: Task-Driven Bayesian Experimental Design with Singly Intractable Objectives Action-Dependent Factorized Baselines Actionable Activation Directions for Detecting and Mitigating Emergent Misalignment Across Language Model Families Activation Atlases activation capping Activation Oracles activation patching Activation Steering active learning Active Listening Active Offline-to-Online Reinforcement Learning ActiveSAM ActiveVision ACTS Ada AdaCodec AdaFlash AdaGrad AdaJEPA AdaJEPA: An Adaptive Latent World Model Adalat AI AdaLook Adam AdamO AdamW AdaPrefix-GRPO ADAPT-GQE ADAPT-VQE adapter fine-tuning Adaptive Clip Policy Optimization Adaptive Data Scheduling Adaptive Depth Sparse Framework Adaptive Gated Feedback Optimization Adaptive Multi-Resolution Procedural Knowledge Compression for Large Language Models Adaptive Multi-Step Lookahead Decoding for Diffusion Language Models Adaptive Parallel Reasoning Adaptive Self-Debiasing adaptive thinking Adaptive Turn-Taking for Real-time Multi-Party Voice Agents adaptive-rank instantiation ADAS AdaSR ADEME ADHD Adithya S K Aditi Krishnapriyan Aditya (co-lead author) Aditya G. Parameswaran Adjacent Contrastive Reasoning ADK Python adk-samples Adobe Adobe Creative Cloud Advanced AI Scaling Framework Advanced Light Source Advanced Matrix Extensions (AMX) Advanced Photon Source Advanced Voice Mode AdvancedMathBench Advancing Omnimodal Embodied Agents from Isolated Skills to Everyday Physical Autonomy AdvBench AdversaBench Adversarial Attacks on Neural Network Policies Adversarial Creation and Detection of AI-Generated Social Bot Content adversarial examples adversarial pragmatics Adversarial Pragmatics for AI Safety Evaluation: A Benchmark for Instruction Conflict, Embedded Commands, and Policy Ambiguity adversarial refinement adversarial robustness adversarial training AdvGRPO Aeneas Aerial Image Dataset (AID) Affinity by Canva AFM 3 AFM 3 Cloud Pro AFM 3 Core Advanced AfriSenti AG-MG Parallel Corpus AG-UI AG-UI Protocol AGC-Bench AGC-Judge Agda AGDO Age of Empires II age prediction Agence France-Presse Agent 365 agent cloud Agent Cognitive Redundancy Ratio Agent Communication Protocol Agent Development Kit agent ecologies Agent Governance Toolkit agent harness Agent JIT Compilation Agent Reasoning Evaluation (ARE) Agent Runtime agent sandboxing agent-device agent-native Agent-Native Immune System Agent-Native Immune System: Architecture, Taxonomy, and Engineering agent-native telemetry Agent-Reach Agent-S agent-skills agent-task efficiency agent-teams-ai agent-to-agent evaluation protocol Agent-to-Agent Protocol (A2A) agent-toolkit-for-aws AgentBeats AgentBeats: Agentifying Agent Assessment for Openness, Standardization, and Reproducibility AgentBoard AgentCL AgentDojo Agentforce agentic AI Agentic AI Foundation Agentic AI Pipelines Agentic AI Systems Agentic Chain-of-Thought Steering Agentic Chain-of-Thought Steering for Efficient and Controllable LLM Reasoning Agentic CLEAR agentic coding Agentic Commerce Protocol Agentic Environment Engineering for Large Language Models: A Survey of Environment Modeling, Synthesis, Evaluation, and Application Agentic generation of verifiable rules for deterministic, self-expanding reaction classification Agentic Permissions Policy Algebra for Taint Confinement in LLM Agents agentic proving Agentic RAG agentic re-optimization framework Agentic Resource Discovery Agentic RL Agentic Self-Instruct Agentic System Monitoring Methodology Agentic Technical Debt agentic vulnerabilities AgenticRL AgenticSTS AgentKit AgentMap agentmemory AgentMob Agentopia Agentopia: Long-Term Life Simulation and Learning in Agent Societies Agents in the Wild: Where Research Meets Deployment Agents-A1 agents-cli Agents-K1 Agents.js AGENTS.md Agents' Last Exam AgentScope AGENTSERVESIM agentskills.io AgentSnare AgentSpec AgentsView AgentWorldBench AGI AGI (Artificial General Intelligence) AGI cognitive framework AGI economy Agility Robotics Agon AGORA agreement attraction AHA-WAM Ahead-of-Time Compilation Ahmad Osman AI Agents AI agents that matter AI alignment AI and Compute AI Andrew AI as Normal Technology AI biosecurity risk assessment AI Clinical Copilot AI Co-Clinician AI control AI Control Roadmap AI Cybersecurity Threat Evaluation Framework AI data sovereignty AI Developer Conference 2025 AI Engineer NYC AI Engineer World's Fair AI Existential Risk AI Exposure Scores: what they measure, what they miss, and what comes next AI for Citizens AI for Game Development AI for Math Initiative AI for Science AI for Science program AI Fund AI Healthcare Chatbots as Information Infrastructure: A Large-Scale Study of User-Reported Breakdowns AI image verification AI inference carbon emissions AI leaderboards AI Liability Directive AI Literacy & Creator Collective AI Mathematical Olympiad AI Overview AI Panic Blog AI Persuasive Framing in Collective Dilemmas AI Registry AI Reproducibility Benchmark AI Risk Management Framework AI Safety Fund AI Safety Level (ASL) AI Safety Level Standards AI Safety via Debate AI scaling laws AI Sheets AI Snake Oil AI systems and the reproduction of (standard) language ideologies in World Englishes AI translation of literary texts is "fine", but readers still prefer human translations AI Verify Foundation AI vs. AI AI Watermark Evidence Fails Forensic Readiness: An Empirical Evaluation AI-accelerated End-to-End Framework for Rapid Professional Upskilling ai-agent-book AI-assisted human evaluation AI-assisted red teaming AI-Assisted Systematization for Evaluating GenAI Systems AI-assisted theorem proving ai-boost/awesome-harness-engineering AI-driven constraint reasoning AI-Driven Life Cycle (AI-DLC) AI-Engineering-Coach ai-hedge-fund AI-MO AI-PAVE-Br AI-Trader AI's Capability in Assisting Scientific Research in Physics, Astrophysics, and Cosmology II: Project Planning and Proposal Evaluation AI@HHMI AI21 Labs AIA Labs AIcrowd Aider AIDev dataset aidlc-workflows Aiera AIEWF AILuminate AIME AIME 2025 AIME 2026 AIME2024 AIME24 AIME25 AIME26 aiming-lab AIMO Interpretability Challenge AIMO Progress Prize AIMS AINews AionUi Air India AIR: Adaptive Interleaved Reasoning with Code in MLLMs Aira Explorer AiraXiv Airbus AirflowAttack AIriskEval-edu Demo AISPA AISPA: User-Centric System Prompt Auditing for Large Language Model Applications aisuite Akshat Bubna Akshay Nathan Albert Gu Alberta AI Academy Alberta Machine Intelligence Institute Alberti ALE-Bench ALER-TI Aletheia Alex Hernandez Alex Kearns-Apuya Alex L. Zhang Alex Lupsasca Alex Rives Alexa+ Alexandr Wang AlexNet ALFRED Alfred Lin ALFWorld Algonauts 2025 Algorithmic and Minimax Complexities in Kernel Bandits Algorithmic Impermeability algorithmic monoculture Algorithmic Monocultures in Hiring Ali Essam Ghareeb Alibaba Alibaba Cloud Alibaba Cloud Model Studio Alibaba DAMO Academy Alibaba Qwen Alibaba Qwen Team ALIGN AlignAtt AlignAtt4LLM ALIGNBEAM alignment auditing alignment faking Alignment Research Center alignment tampering alignment tax Alireza Rezvani All Hands AI All Smoke, No Alarm: Oracle Signals in Agent-Authored Test Code All Tech Is Human all-edge F1 Allam Allegro Allegro Hand Allen Institute Allen Institute for AI AllenAI Allied for Startups ALMANAC ALOHA AlpacaEval 2 Alpamayo R1 Alphabet AlphaEarth Foundations AlphaEvolve AlphaFold AlphaFold2 AlphaFold3 AlphaGenome AlphaGo AlphaOracle AlphaStar AlphaTensor alpic-ai Altana Altera Alternating Token-Weighted Unlearning Altimeter Capital Always-On Evaluation Protocol (AOEP-v0) Always-OnAgents: A Survey of Persistent Memory, State, and Governance in LLM Agents ALX Alyah AM-Parser AMALIA Amar Subramanya AMARIS Amazon Amazon Beauty Amazon Bedrock Amazon Bedrock AgentCore Amazon Graviton Amazon Reviews Dataset Amazon SageMaker Amazon SageMaker Studio Amazon Trainium2 Amazon Web Services Ambient Diffusion Policy AMC12 AMD AMD Instinct AMD Instinct MI300 AMD MI300 AMD MI355X AMEL American Federation of Teachers American Technology Council Ami Vora AMIA Amir Bar amortized variational inference AMP Amrith Setlur AMRS aMUSEd An Agency-Transferring Model-Free Policy Enhancement Technique An Empirical Analysis of Factual Errors in Human-Written Text and its Application An Evaluation Framework for Structured Audio Captions Validated by Controlled Perturbations An Exact Instrument for State Usage in Selective State-Space Models, and the Input-Driven Migration It Reveals Analogical Deep Research Analogical Deep Research: Retrieving and Integrating Historical Analogies for Foresight Analysis Analysing Self-Harm Representations in Language Models: a Cross-Architecture Study AnalyticGeo7K Analytics-Everywhere-Lab Anatomy of Post-Training: Using Interpretability to Characterize Data and Shape the Learning Signal ANCESTRA anchoring effect Andon Labs Andreessen Horowitz Andrej Karpathy Andrew Kelley Andrew Ng Andrew Qu Andrew Wagenmaker Anduril Industries Andy Beam Andy Jassy Andyyyy64 Angelini Pharma AngelSpec Anglophone retrieval bias Angluin's condition anisotropic BAO damping Anjney Midha Anna's Archive Annapurna Labs annotator disagreement anomalyco AnomalyShapeNet ANSI National Accreditation Board AnswerSheet ANSYR ANTAP Anthology Fund Anthony Albanese Anthropic Anthropic Academy Anthropic Advanced AI Framework Anthropic Agent SDK Anthropic Beneficial Deployments Anthropic Claude Code Anthropic Economic Advisory Council Anthropic Economic Futures Programme Anthropic Economic Futures Research Fund Anthropic Economic Index Anthropic Economic Policy Framework Anthropic Fable 5 Anthropic for Startups Anthropic Interviewer Anthropic Korea Anthropic Labs Anthropic Long-Term Benefit Trust Anthropic Mythos Anthropic National Security and Public Sector Advisory Council Anthropic Partner Academy Anthropic Policy Frontier Red Team Anthropic Public Record Anthropic Responsible Scaling Policy Anthropic Safeguards Team Anthropic Terms of Service Anthropic Threat Intelligence Report August 2025 Anthropic Usage Policy Anthropic v. Department of War anthropics/skills anthropomorphism in AI Anti-Drift Rectification Anti-Periodic Positional Encoding: Möbius Boundary Conditions Make In-Context Retrieval Reliable Antigravity Antoine Zambelli Anton Korinek Any-Dimensional Learning by Sampling AnyGroundBench AnyLanguageModel AnyMo Anysphere Anytime-Valid E-Process anytime-valid sequential testing AP2 Apache 2.0 Apache Arrow Apache Ossie Apache Software Foundation Apache Spark Apart Research APE Apertus 70B Apertus-8B-Instruct-2509 APEX-Accounting APEX-Accounting APEX-Agents-AA aphasia API knowledge boundary probing APO: Unsupervised Atomic Policy Optimization for 3D Structure Prediction of Atomic Systems Apollo Apollo Global Management Apollo Research APPA (Agentic Permissions Policy Algebra) Apparent Psychological Profiles of Large Language Models are Largely a Measurement Artifact Appearance Pointers -- Multimodal Region Control of Diffusion Transformers Appia Foundation Apple Apple Foundation Models 3 Apple Foundation Models Framework Apple Health Apple Neural Engine Apple Silicon Applied Intuition apply_patch APPO: Agentic Procedural Policy Optimization Approximate DP APPS Apps SDK AppWorld Apriel-H1 AprielGuard APS-Bench APS-RAG apurvsinghgautam AR military headset Ar9av AraBERT Arabic Arabic Instruction Following Eval (IFEval) Arabic LLMs ArabiGEE AraGen ARC Evals ARC Framework Arc Institute ARC Prize Foundation Arc Virtual Cell Challenge ARC-AGI Archon Arco Education ARDY Are Performance-Optimization Benchmarks Reliably Measuring Coding Agents? Are We Ready For An Agent-Native Memory System? AREA Arena AI Arena Code Arena Leaderboard Arena Search Arena-Hard Arena.ai Code Arena WebDev ArenaHard Argentina Argilla Argilla 2.0 Argonne National Laboratory Argus ariadng ARIS Aristotle API Arithmetic Pedagogy for Language Models Arize AI Arize Phoenix Arm Arm AGI CPU Armin Ronacher ArMIS Army Special Operations Command Arnuv Tandon ArogyaBodha ArogyaSutra ARPA-H ARROW ART (Agent Reinforcement Trainer) Arthur Mensch Artifact Artifacts Artificial Analysis Artificial Analysis Big Bench Audio Artificial Analysis Coding Agent Index Artificial Analysis Conversational Dynamics Artificial Analysis Intelligence Index Artificial Analysis Intelligence Leaderboard Artificial Analysis LLM Performance Leaderboard Artificial Analysis Text to Image Artificial Analysis Text to Image Leaderboard Artificial Bee Colony Artificial Epanorthosis: Why large language models overuse a classical rhetorical figure, and how to mitigate it ArXiv arXiv:2602.05394 ASAM ASAP 2.0 ASAP++ Ascend SuperPOD Ashwyn Sharma Asia-Pacific ASIMOV-2.0 AskChem AskChem-Bench AskDocs ASL-2 ASL-3 ASML ASML Holding NV Aspect-Based Sentiment Analysis ASRD AssemblyAI AssemblyAI Universal Assessing the Operational Impact of Poisoning Attacks over Augmented 3D Point Cloud Public Datasets for Connected and Autonomous Vehicles AssetOpsBench assistant axis Assistants API Assisted Generation ASSISTments Association for Computational Linguistics Associative Recurrent Memory Transformer Astera Institute ASTRA Astral AstrBot AstrBotDevs ASVspoof 5 asynchronous inference Asynchronous Noise Schedule Atari ATDev ATE-Bench Atla Atlas Atlas H&E-TME ATLAS Open Data 13 TeV ATLAS: Active Theory Learning for Automated Science Atlassian AToken Atomic Policy Optimization ATOMIC2020 Atoms of Thought (EEG Microstates paper) Attack Success Rate Attainable Utility Preservation Attention Amnesia in Hybrid LLMs: When CoT Fine-Tuning Breaks Long-Range Recall, and How to Fix It attention entropy Attention Expansion: Enhancing Keyphrase Extraction from Long Documents with Attention-Augmented Contextualized Embeddings attention head circuit Attention Residuals Attention Sorting Attention submodule Attention-based MIL Attractor States Emerge in Multi-Turn LLM Conversations AUC Audex Audio Interaction Model Audio MultiChallenge Audio-Flamingo-3 Audio-LLM Audio-Native Speech Recognition with a Frozen Discrete-Diffusion Language Model AudioCards AudioDER AudioLDM 2 AudioVAE2 Auditing Protocol-Level Shortcuts in Large Audio Language Model Judges for Speech Evaluation AUDITS Augment Augment Code AUPRC AuRA AUROC Australia AI Safety Institute Australia National AI Plan Australian Emergency Department triage notes Australian Government Australian National University Auto Benchmark Audit (ABA) Autodata Autodesk Fusion AutoDex autoDream AutoForest Autoformer AutoGPTQ AutoJourn AUTokenizer AutoLab automated AI research Automated Background Swapping Automated Compliance Mapping in Cloud Security with Domain-Adapted Sentence Transformers Automated Discovery Has No Universally Superior Harness Automated Extraction of Techno-Economic Data from 76,000 Energy System Studies Automated grading of Linux/bash examinations using large language models: a four-level cognitive taxonomy approach automated mechanistic interpretability automated red teaming Automated Reference Verification System Automated reproducibility assessments in the social and behavioral sciences using large language models automated test suite automated theorem proving Automatic Domain Randomization Automatic Post-Editing (APE) Automatic Speech Recognition AutomationBench-AA AutoMem autonomous driving AUTOPILOT-VQA Autoregressive Boltzmann Generators autoregressive transformer autoresearch AutoResearchClaw AutoRound AutoSkillHarm AutoSynthesis avatarin Avenger Labs Aviral Kumar awesome-claude-code awesome-copilot awesome-llm-apps AWQ AWS AWS Bedrock AWS GovCloud AWS Inferentia2 AWS Labs AWS Marketplace AWS Neuron AWS Neuron SDK AWS Trainium AxDafny Axel Backlund Axiom Math Axios Axolotl AXPO Aya Aya Expanse Aya Vision Aya-Expanse-8B-Base Azure Azure AI Studio Azure DevOps Azure Foundry Azure OCR B³D-RWKV BabelJudge BABI-Long BABILong BabyCL Backbone-as-Architect backdoor attack Backdoor Circuit Analysis (Language-Switching) Backdoor Unlearning Generalization: A Path Toward the Removal of Unknown Triggers in LLMs backnotprop backpropagation through time Backstory Bad Grammar Bad Scientist BadWAM BadWAM: When World-Action Models Dream Right but Act Wrong Bag of Words (BoW) BAGEL BAGEL-7B Baidu Baidu Translate Baifeng Shi Baiyu Chen ball balancing ball balancing task Bamba BamiBERT Bandit Bang-v3 BanglaBERT Banque des Territoires Barcelona Supercomputing Center Language Technologies bare metal sandboxes Bargaining Scenarios Dataset Bark Base Baseline Agent Baseline-Log Physical Separation Baseten BashArena basic-memory basicmachines-co Basque Center on Cognition, Brain, and Language Basque language baulab Bayesian decision theory Bayesian Inference and Decision Audits for Public Archives of Frontier AI Evaluations Bayesian Multiobjective Optimization Bayesian Optimal Experimental Design Bayesian Optimization BayesPO BayLing-Duplex BBC News BBQ BBVA BC Protocol BCG X BCR BDD100K Be My Eyes Be My Tutor: On-Policy Co-Distillation for Mutual LLM Improvement via Peer Feedback BEA-Dialogue Beam Search Bebop Bechiri and Lanasri [2026] Becker Friedman Institute for Economics Before You Think: System 0, AI-Mediated Cognition and Cognitive Colonization behavior cloning behavior trees with learning-enabled components behavioral fine-tuning Behavioral Trajectory Tracking Framework behavioral-gradient validator Behavioral-SafetyBench Behavox BEIR Belebele Belief-Space Safety Filter (BeliefSF) BeliefTrack Ben Bernanke Ben Mann Ben Parr Ben Thompson Benchling Benchmark Agent Benchmark Everything Everywhere All at Once Benchmarking Multimodal Large Language Models for Scientific Visualization Literacy benchmaxxing BenCzechMark Bending Spoons Beneficial Deployments Benevity BenevolentAI Bengali Bengaluru BenHalluEval BenHalluScore Benjamin Chang Benjamin Krause Berkeley AI Research (BAIR) Berkeley Artificial Intelligence Research Berkshire Hathaway Berlin Database of Emotional Speech (EMO-DB) BerriAI BERT BERT-base BERT-F1 BERTopic BERTScore BERTurk BERTweet BES (trained models) Bessemer Venture Partners Best-of-N Best-of-N Sampling best@k Better Call GRPO Beyond a Single Judge: Simulating Social Persona Panels for Generative UI Evaluation Beyond Accuracy: Community Perspectives on Machine Translation Beyond Adam: SOAP and Muon for Faster, Label-Efficient Training of Machine Learning Interatomic Potentials Beyond Aggregate Risk: Role-Stratified Conformal Risk Control for LLM Tool Calls Beyond Attack-Success Rate: Action-Graded Severity Scale for Tool-Using AI Agents Beyond Benchmarks: Exposing the Hidden Crisis in Bangla Hate Speech Detection Beyond Fixed Representations: The Vocabulary and Verifier Gaps in Open-Ended AI Beyond Fully Random Masking: Attention-Guided Denoising and Optimization for Diffusion Language Models Beyond Function Calling: Benchmarking Tool-Using Agents under Tool-Environment Unreliability Beyond Global Replanning: Hierarchical Recovery for Cross-Device Agent Systems Beyond Independent Labels: Schwartz-Geometry Decoding for Human Value Detection Beyond Negative-Ridge Endpoints: Mixed-Sign Spectral Regularization via Negative-Shifted Gradient Descent Beyond Runtime Enforcement: Shield Synthesis as Defensibility Analysis for Adversarial Networks Beyond Scale and Generation: Understanding Language Model-based Entity Matching Beyond Sentiment: Structured Information Extraction from Financial News Beyond Success Rate: Cost-Aware Evaluation of Offensive and Defensive Security Agents Beyond Sufficiency: Time Series Explanation with Counterfactual Necessity Beyond Surface Forms: A Comprehensive, Mechanism-Oriented Taxonomy of Indirect Linguistic Encoding for LLM-Based Coded Language Detection Beyond Sycophancy: Structured Resistance and Compliance in LLM Moral Reasoning Beyond task performance: Decoding bioacoustic embeddings with speech features Beyond the Commitment Boundary: Probing Epiphenomenal Chain-of-Thought in Large Reasoning Models Beyond the Hard Budget: Sparsity Regularizers for More Interpretable Top-k Sparse Autoencoders Beyond the Leaderboard: Design Lessons for Trustworthy Multimodal VQA Beyond Third-Person Audits: Situated Interaction Auditing for User-Centered LLM Bias Research Beyond Tokenization: Direct Timestep Embedding and Contrastive Alignment for Time-Series Question Answering Beyond Uniform Tokens: Adaptive Compression for Time Series Language Models BeyondUncertainty BFCL BFCL Multi-Turn BFCL-V3 BFCLv3 BFW schedule hint BGE BGE-M3 BGL bi-manual robotic manipulation Bias Benchmark for Question Answering Bias Leaves a Gradient Trail Biconvex Optimization bidirectional attention Bidirectional Evolutionary Search Bielik Big Bench Big Bench Audio Big-Math BigBird BigCode BigCodeArena BigCodeBench Bigger or Cheaper? Scale and Quantization Effects on Uncertainty Signals in Vision-Language Models Under Image Degradation BigLaw Bench BigScience Bill & Melinda Gates Foundation Bill Dally Bill Gurley Bill McDermott Binarized Neural Network binary quantization BINEVAL Bing binning semiring bioacoustic monitoring bioacoustics BioASQ 14b BioBERT Biographical
any entity 1-Lipschitz Neural Networks on Hadamard Manifolds 1.58-bit quantization 10 USC 3252 10x Genomics 12-Factor App 1B-scale language models 1M-token context 1X 2-Parameter Logistic IRT Model 2026 FIFA World Cup 2026 State of AI Traffic and Cyberthreat Benchmark Report 2D-RoPE 2WikiMultiHopQA 32B language model 3C3H 3D asset generation 3D Gaussian Splatting 3D Medical Image Classification 3D optical flow 3D point cloud generation 3D tracking 3D-Aware VLMs with Implicit and Explicit Geometries 3D-Fit 3D-RoPE 3E Memory Module 3LM 404 Media 4chan 4D reconstruction 5PILS 6-PACK 6G Radio Access Network 7B language model 8B autoregressive language model A Blueprint for Equilibrium-Based Differentiable Continuous-Variable Thermodynamic Computing A Causal Model of Theory of Mind in Conflict for Artificial Intelligence A Comparative Study of Deep Learning Architectures for Multi-Horizon Behavioural Forecasting for Mobile Health A Computational Audit of Demographic Association Encoding in ClinicalBERT Language Predictions A Constrained Natural-Language Interface for Variational Multi-Physics Finite Element Simulations in FEniCS A Continual Validation, Updating, and Decision-Making Framework for Self-Adaptive Digital Twins via Robust Model Predictive Control A Diffusion Approximation for Temporal-Difference Learning with Linear Features under Markovian Noise A Factorial Study of Synthetic Data Generation for Low-Resource Machine Translation using Grammar Books A History-Aware Visually Grounded Critic for Computer Use Agents A Human-in-the-Loop Corpus for LLM-Based Simplification of Scientific Summaries A Komi-Yazva–Russian Parallel Corpus and Evaluation Protocol for Zero- and Few-Shot LLM Translation A Learning-Rate-Gated Failure of GRPO in a Small Language and Vision-Language Model Web Agent A Methodology for Auditable Trustworthiness Levels in AI Lifecycle Governance A Multi-Agent System for Autonomous, Fine-Tuning-Free Clinical Symptom Detection: Development and Validation Study A Multi-Domain Benchmark for Detecting AI-Generated Text-Rich Images from GPT-Image-2 A Physics-Informed Fourier-Wavelet Transformer for Multiscale Computational Fluid Dynamics Surrogate Modeling A Practical Investigation of Training-free Relaxed Speculative Decoding A Resource for Enthymeme Detection in Controversial Political Discourse A Self-Evolving Agent for Longitudinal Personal Health Management A sleep-like consolidation mechanism for LLMs A Study of Temporal Fusion Strategies for Named Entity Recognition in Historical Texts A Systematic Approach for Selecting Trajectories for Data Augmentation A Taxonomy of Conceptual Alignment in Human-Robot Dialogue A Training-Free Mixture-of-Agents Framework for Multi-Document Summarization using LLMs and Knowledge Graphs A Tree-of-Thoughts Inspired Hybrid Approach for Legal Case Judgement Summarization using LLMs A Unifying Lens on Supervised Fine-Tuning Through Target Distribution Design A24 A2A Protocol A2C A2Z A3C AA-Briefcase AA-Omniscience AAAI Aardvark AARRI-Bench AASIST AB-UPT Abanca ABBEL ABC-Bench Abeba Birhane abhigyanpatwari abjadai Ableton abliteration Abridge Abstraction Gap Accelerate Accelerated Decentralized Stochastic Gradient Descent for Strongly Convex Optimization Accelerating Masked Diffusion Large Language Models: A Survey of Efficient Inference Techniques Accenture Accenture Anthropic Business Group accumulated message effect Accuracy and Satisfaction in Multi-Turn LLM Dialogues for NFR Assessment ACE ACEBench-Agent AceReason-14B Achieving Precise Text-To-Cypher Via Grounded Knowledge Graph Data Generation ACKTR ACL ACL Anthology ACL-Verbatim Acoustic Cue Alignment in Audio Language Models for Speech Emotion Recognition ACROS Act2Answer Action Chunking Transformer ACTION-BED Action-BED: Task-Driven Bayesian Experimental Design with Singly Intractable Objectives Action-Dependent Factorized Baselines Actionable Activation Directions for Detecting and Mitigating Emergent Misalignment Across Language Model Families Activation Atlases activation capping Activation Oracles activation patching Activation Steering active learning Active Listening Active Offline-to-Online Reinforcement Learning ActiveSAM ActiveVision ACTS Ada AdaCodec AdaFlash AdaGrad AdaJEPA AdaJEPA: An Adaptive Latent World Model Adalat AI AdaLook Adam AdamO AdamW AdaPrefix-GRPO ADAPT-GQE ADAPT-VQE adapter fine-tuning Adaptive Clip Policy Optimization Adaptive Data Scheduling Adaptive Depth Sparse Framework Adaptive Gated Feedback Optimization Adaptive Multi-Resolution Procedural Knowledge Compression for Large Language Models Adaptive Multi-Step Lookahead Decoding for Diffusion Language Models Adaptive Parallel Reasoning Adaptive Self-Debiasing adaptive thinking Adaptive Turn-Taking for Real-time Multi-Party Voice Agents adaptive-rank instantiation ADAS AdaSR ADEME ADHD Adithya S K Aditi Krishnapriyan Aditya (co-lead author) Aditya G. Parameswaran Adjacent Contrastive Reasoning ADK Python adk-samples Adobe Adobe Creative Cloud Advanced AI Scaling Framework Advanced Light Source Advanced Matrix Extensions (AMX) Advanced Photon Source Advanced Voice Mode AdvancedMathBench Advancing Omnimodal Embodied Agents from Isolated Skills to Everyday Physical Autonomy AdvBench AdversaBench Adversarial Attacks on Neural Network Policies Adversarial Creation and Detection of AI-Generated Social Bot Content adversarial examples adversarial pragmatics Adversarial Pragmatics for AI Safety Evaluation: A Benchmark for Instruction Conflict, Embedded Commands, and Policy Ambiguity adversarial refinement adversarial robustness adversarial training AdvGRPO Aeneas Aerial Image Dataset (AID) Affinity by Canva AFM 3 AFM 3 Cloud Pro AFM 3 Core Advanced AfriSenti AG-MG Parallel Corpus AG-UI AG-UI Protocol AGC-Bench AGC-Judge Agda AGDO Age of Empires II age prediction Agence France-Presse Agent 365 agent cloud Agent Cognitive Redundancy Ratio Agent Communication Protocol Agent Development Kit agent ecologies Agent Governance Toolkit agent harness Agent JIT Compilation Agent Reasoning Evaluation (ARE) Agent Runtime agent sandboxing agent-device agent-native Agent-Native Immune System Agent-Native Immune System: Architecture, Taxonomy, and Engineering agent-native telemetry Agent-Reach Agent-S agent-skills agent-task efficiency agent-teams-ai agent-to-agent evaluation protocol Agent-to-Agent Protocol (A2A) agent-toolkit-for-aws AgentBeats AgentBeats: Agentifying Agent Assessment for Openness, Standardization, and Reproducibility AgentBoard AgentCL AgentDojo Agentforce agentic AI Agentic AI Foundation Agentic AI Pipelines Agentic AI Systems Agentic Chain-of-Thought Steering Agentic Chain-of-Thought Steering for Efficient and Controllable LLM Reasoning Agentic CLEAR agentic coding Agentic Commerce Protocol Agentic Environment Engineering for Large Language Models: A Survey of Environment Modeling, Synthesis, Evaluation, and Application Agentic generation of verifiable rules for deterministic, self-expanding reaction classification Agentic Permissions Policy Algebra for Taint Confinement in LLM Agents agentic proving Agentic RAG agentic re-optimization framework Agentic Resource Discovery Agentic RL Agentic Self-Instruct Agentic System Monitoring Methodology Agentic Technical Debt agentic vulnerabilities AgenticRL AgenticSTS AgentKit AgentMap agentmemory AgentMob Agentopia Agentopia: Long-Term Life Simulation and Learning in Agent Societies Agents in the Wild: Where Research Meets Deployment Agents-A1 agents-cli Agents-K1 Agents.js AGENTS.md Agents' Last Exam AgentScope AGENTSERVESIM agentskills.io AgentSnare AgentSpec AgentsView AgentWorldBench AGI AGI (Artificial General Intelligence) AGI cognitive framework AGI economy Agility Robotics Agon AGORA agreement attraction AHA-WAM Ahead-of-Time Compilation Ahmad Osman AI Agents AI agents that matter AI alignment AI and Compute AI Andrew AI as Normal Technology AI biosecurity risk assessment AI Clinical Copilot AI Co-Clinician AI control AI Control Roadmap AI Cybersecurity Threat Evaluation Framework AI data sovereignty AI Developer Conference 2025 AI Engineer NYC AI Engineer World's Fair AI Existential Risk AI Exposure Scores: what they measure, what they miss, and what comes next AI for Citizens AI for Game Development AI for Math Initiative AI for Science AI for Science program AI Fund AI Healthcare Chatbots as Information Infrastructure: A Large-Scale Study of User-Reported Breakdowns AI image verification AI inference carbon emissions AI leaderboards AI Liability Directive AI Literacy & Creator Collective AI Mathematical Olympiad AI Overview AI Panic Blog AI Persuasive Framing in Collective Dilemmas AI Registry AI Reproducibility Benchmark AI Risk Management Framework AI Safety Fund AI Safety Level (ASL) AI Safety Level Standards AI Safety via Debate AI scaling laws AI Sheets AI Snake Oil AI systems and the reproduction of (standard) language ideologies in World Englishes AI translation of literary texts is "fine", but readers still prefer human translations AI Verify Foundation AI vs. AI AI Watermark Evidence Fails Forensic Readiness: An Empirical Evaluation AI-accelerated End-to-End Framework for Rapid Professional Upskilling ai-agent-book AI-assisted human evaluation AI-assisted red teaming AI-Assisted Systematization for Evaluating GenAI Systems AI-assisted theorem proving ai-boost/awesome-harness-engineering AI-driven constraint reasoning AI-Driven Life Cycle (AI-DLC) AI-Engineering-Coach ai-hedge-fund AI-MO AI-PAVE-Br AI-Trader AI's Capability in Assisting Scientific Research in Physics, Astrophysics, and Cosmology II: Project Planning and Proposal Evaluation AI@HHMI AI21 Labs AIA Labs AIcrowd Aider AIDev dataset aidlc-workflows Aiera AIEWF AILuminate AIME AIME 2025 AIME 2026 AIME2024 AIME24 AIME25 AIME26 aiming-lab AIMO Interpretability Challenge AIMO Progress Prize AIMS AINews AionUi Air India AIR: Adaptive Interleaved Reasoning with Code in MLLMs Aira Explorer AiraXiv Airbus AirflowAttack AIriskEval-edu Demo AISPA AISPA: User-Centric System Prompt Auditing for Large Language Model Applications aisuite Akshat Bubna Akshay Nathan Albert Gu Alberta AI Academy Alberta Machine Intelligence Institute Alberti ALE-Bench ALER-TI Aletheia Alex Hernandez Alex Kearns-Apuya Alex L. Zhang Alex Lupsasca Alex Rives Alexa+ Alexandr Wang AlexNet ALFRED Alfred Lin ALFWorld Algonauts 2025 Algorithmic and Minimax Complexities in Kernel Bandits Algorithmic Impermeability algorithmic monoculture Algorithmic Monocultures in Hiring Ali Essam Ghareeb Alibaba Alibaba Cloud Alibaba Cloud Model Studio Alibaba DAMO Academy Alibaba Qwen Alibaba Qwen Team ALIGN AlignAtt AlignAtt4LLM ALIGNBEAM alignment auditing alignment faking Alignment Research Center alignment tampering alignment tax Alireza Rezvani All Hands AI All Smoke, No Alarm: Oracle Signals in Agent-Authored Test Code All Tech Is Human all-edge F1 Allam Allegro Allegro Hand Allen Institute Allen Institute for AI AllenAI Allied for Startups ALMANAC ALOHA AlpacaEval 2 Alpamayo R1 Alphabet AlphaEarth Foundations AlphaEvolve AlphaFold AlphaFold2 AlphaFold3 AlphaGenome AlphaGo AlphaOracle AlphaStar AlphaTensor alpic-ai Altana Altera Alternating Token-Weighted Unlearning Altimeter Capital Always-On Evaluation Protocol (AOEP-v0) Always-OnAgents: A Survey of Persistent Memory, State, and Governance in LLM Agents ALX Alyah AM-Parser AMALIA Amar Subramanya AMARIS Amazon Amazon Beauty Amazon Bedrock Amazon Bedrock AgentCore Amazon Graviton Amazon Reviews Dataset Amazon SageMaker Amazon SageMaker Studio Amazon Trainium2 Amazon Web Services Ambient Diffusion Policy AMC12 AMD AMD Instinct AMD Instinct MI300 AMD MI300 AMD MI355X AMEL American Federation of Teachers American Technology Council Ami Vora AMIA Amir Bar amortized variational inference AMP Amrith Setlur AMRS aMUSEd An Agency-Transferring Model-Free Policy Enhancement Technique An Empirical Analysis of Factual Errors in Human-Written Text and its Application An Evaluation Framework for Structured Audio Captions Validated by Controlled Perturbations An Exact Instrument for State Usage in Selective State-Space Models, and the Input-Driven Migration It Reveals Analogical Deep Research Analogical Deep Research: Retrieving and Integrating Historical Analogies for Foresight Analysis Analysing Self-Harm Representations in Language Models: a Cross-Architecture Study AnalyticGeo7K Analytics-Everywhere-Lab Anatomy of Post-Training: Using Interpretability to Characterize Data and Shape the Learning Signal ANCESTRA anchoring effect Andon Labs Andreessen Horowitz Andrej Karpathy Andrew Kelley Andrew Ng Andrew Qu Andrew Wagenmaker Anduril Industries Andy Beam Andy Jassy Andyyyy64 Angelini Pharma AngelSpec Anglophone retrieval bias Angluin's condition anisotropic BAO damping Anjney Midha Anna's Archive Annapurna Labs annotator disagreement anomalyco AnomalyShapeNet ANSI National Accreditation Board AnswerSheet ANSYR ANTAP Anthology Fund Anthony Albanese Anthropic Anthropic Academy Anthropic Advanced AI Framework Anthropic Agent SDK Anthropic Beneficial Deployments Anthropic Claude Code Anthropic Economic Advisory Council Anthropic Economic Futures Programme Anthropic Economic Futures Research Fund Anthropic Economic Index Anthropic Economic Policy Framework Anthropic Fable 5 Anthropic for Startups Anthropic Interviewer Anthropic Korea Anthropic Labs Anthropic Long-Term Benefit Trust Anthropic Mythos Anthropic National Security and Public Sector Advisory Council Anthropic Partner Academy Anthropic Policy Frontier Red Team Anthropic Public Record Anthropic Responsible Scaling Policy Anthropic Safeguards Team Anthropic Terms of Service Anthropic Threat Intelligence Report August 2025 Anthropic Usage Policy Anthropic v. Department of War anthropics/skills anthropomorphism in AI Anti-Drift Rectification Anti-Periodic Positional Encoding: Möbius Boundary Conditions Make In-Context Retrieval Reliable Antigravity Antoine Zambelli Anton Korinek Any-Dimensional Learning by Sampling AnyGroundBench AnyLanguageModel AnyMo Anysphere Anytime-Valid E-Process anytime-valid sequential testing AP2 Apache 2.0 Apache Arrow Apache Ossie Apache Software Foundation Apache Spark Apart Research APE Apertus 70B Apertus-8B-Instruct-2509 APEX-Accounting APEX-Accounting APEX-Agents-AA aphasia API knowledge boundary probing APO: Unsupervised Atomic Policy Optimization for 3D Structure Prediction of Atomic Systems Apollo Apollo Global Management Apollo Research APPA (Agentic Permissions Policy Algebra) Apparent Psychological Profiles of Large Language Models are Largely a Measurement Artifact Appearance Pointers -- Multimodal Region Control of Diffusion Transformers Appia Foundation Apple Apple Foundation Models 3 Apple Foundation Models Framework Apple Health Apple Neural Engine Apple Silicon Applied Intuition apply_patch APPO: Agentic Procedural Policy Optimization Approximate DP APPS Apps SDK AppWorld Apriel-H1 AprielGuard APS-Bench APS-RAG apurvsinghgautam AR military headset Ar9av AraBERT Arabic Arabic Instruction Following Eval (IFEval) Arabic LLMs ArabiGEE AraGen ARC Evals ARC Framework Arc Institute ARC Prize Foundation Arc Virtual Cell Challenge ARC-AGI Archon Arco Education ARDY Are Performance-Optimization Benchmarks Reliably Measuring Coding Agents? Are We Ready For An Agent-Native Memory System? AREA Arena AI Arena Code Arena Leaderboard Arena Search Arena-Hard Arena.ai Code Arena WebDev ArenaHard Argentina Argilla Argilla 2.0 Argonne National Laboratory Argus ariadng ARIS Aristotle API Arithmetic Pedagogy for Language Models Arize AI Arize Phoenix Arm Arm AGI CPU Armin Ronacher ArMIS Army Special Operations Command Arnuv Tandon ArogyaBodha ArogyaSutra ARPA-H ARROW ART (Agent Reinforcement Trainer) Arthur Mensch Artifact Artifacts Artificial Analysis Artificial Analysis Big Bench Audio Artificial Analysis Coding Agent Index Artificial Analysis Conversational Dynamics Artificial Analysis Intelligence Index Artificial Analysis Intelligence Leaderboard Artificial Analysis LLM Performance Leaderboard Artificial Analysis Text to Image Artificial Analysis Text to Image Leaderboard Artificial Bee Colony Artificial Epanorthosis: Why large language models overuse a classical rhetorical figure, and how to mitigate it ArXiv arXiv:2602.05394 ASAM ASAP 2.0 ASAP++ Ascend SuperPOD Ashwyn Sharma Asia-Pacific ASIMOV-2.0 AskChem AskChem-Bench AskDocs ASL-2 ASL-3 ASML ASML Holding NV Aspect-Based Sentiment Analysis ASRD AssemblyAI AssemblyAI Universal Assessing the Operational Impact of Poisoning Attacks over Augmented 3D Point Cloud Public Datasets for Connected and Autonomous Vehicles AssetOpsBench assistant axis Assistants API Assisted Generation ASSISTments Association for Computational Linguistics Associative Recurrent Memory Transformer Astera Institute ASTRA Astral AstrBot AstrBotDevs ASVspoof 5 asynchronous inference Asynchronous Noise Schedule Atari ATDev ATE-Bench Atla Atlas Atlas H&E-TME ATLAS Open Data 13 TeV ATLAS: Active Theory Learning for Automated Science Atlassian AToken Atomic Policy Optimization ATOMIC2020 Atoms of Thought (EEG Microstates paper) Attack Success Rate Attainable Utility Preservation Attention Amnesia in Hybrid LLMs: When CoT Fine-Tuning Breaks Long-Range Recall, and How to Fix It attention entropy Attention Expansion: Enhancing Keyphrase Extraction from Long Documents with Attention-Augmented Contextualized Embeddings attention head circuit Attention Residuals Attention Sorting Attention submodule Attention-based MIL Attractor States Emerge in Multi-Turn LLM Conversations AUC Audex Audio Interaction Model Audio MultiChallenge Audio-Flamingo-3 Audio-LLM Audio-Native Speech Recognition with a Frozen Discrete-Diffusion Language Model AudioCards AudioDER AudioLDM 2 AudioVAE2 Auditing Protocol-Level Shortcuts in Large Audio Language Model Judges for Speech Evaluation AUDITS Augment Augment Code AUPRC AuRA AUROC Australia AI Safety Institute Australia National AI Plan Australian Emergency Department triage notes Australian Government Australian National University Auto Benchmark Audit (ABA) Autodata Autodesk Fusion AutoDex autoDream AutoForest Autoformer AutoGPTQ AutoJourn AUTokenizer AutoLab automated AI research Automated Background Swapping Automated Compliance Mapping in Cloud Security with Domain-Adapted Sentence Transformers Automated Discovery Has No Universally Superior Harness Automated Extraction of Techno-Economic Data from 76,000 Energy System Studies Automated grading of Linux/bash examinations using large language models: a four-level cognitive taxonomy approach automated mechanistic interpretability automated red teaming Automated Reference Verification System Automated reproducibility assessments in the social and behavioral sciences using large language models automated test suite automated theorem proving Automatic Domain Randomization Automatic Post-Editing (APE) Automatic Speech Recognition AutomationBench-AA AutoMem autonomous driving AUTOPILOT-VQA Autoregressive Boltzmann Generators autoregressive transformer autoresearch AutoResearchClaw AutoRound AutoSkillHarm AutoSynthesis avatarin Avenger Labs Aviral Kumar awesome-claude-code awesome-copilot awesome-llm-apps AWQ AWS AWS Bedrock AWS GovCloud AWS Inferentia2 AWS Labs AWS Marketplace AWS Neuron AWS Neuron SDK AWS Trainium AxDafny Axel Backlund Axiom Math Axios Axolotl AXPO Aya Aya Expanse Aya Vision Aya-Expanse-8B-Base Azure Azure AI Studio Azure DevOps Azure Foundry Azure OCR B³D-RWKV BabelJudge BABI-Long BABILong BabyCL Backbone-as-Architect backdoor attack Backdoor Circuit Analysis (Language-Switching) Backdoor Unlearning Generalization: A Path Toward the Removal of Unknown Triggers in LLMs backnotprop backpropagation through time Backstory Bad Grammar Bad Scientist BadWAM BadWAM: When World-Action Models Dream Right but Act Wrong Bag of Words (BoW) BAGEL BAGEL-7B Baidu Baidu Translate Baifeng Shi Baiyu Chen ball balancing ball balancing task Bamba BamiBERT Bandit Bang-v3 BanglaBERT Banque des Territoires Barcelona Supercomputing Center Language Technologies bare metal sandboxes Bargaining Scenarios Dataset Bark Base Baseline Agent Baseline-Log Physical Separation Baseten BashArena basic-memory basicmachines-co Basque Center on Cognition, Brain, and Language Basque language baulab Bayesian decision theory Bayesian Inference and Decision Audits for Public Archives of Frontier AI Evaluations Bayesian Multiobjective Optimization Bayesian Optimal Experimental Design Bayesian Optimization BayesPO BayLing-Duplex BBC News BBQ BBVA BC Protocol BCG X BCR BDD100K Be My Eyes Be My Tutor: On-Policy Co-Distillation for Mutual LLM Improvement via Peer Feedback BEA-Dialogue Beam Search Bebop Bechiri and Lanasri [2026] Becker Friedman Institute for Economics Before You Think: System 0, AI-Mediated Cognition and Cognitive Colonization behavior cloning behavior trees with learning-enabled components behavioral fine-tuning Behavioral Trajectory Tracking Framework behavioral-gradient validator Behavioral-SafetyBench Behavox BEIR Belebele Belief-Space Safety Filter (BeliefSF) BeliefTrack Ben Bernanke Ben Mann Ben Parr Ben Thompson Benchling Benchmark Agent Benchmark Everything Everywhere All at Once Benchmarking Multimodal Large Language Models for Scientific Visualization Literacy benchmaxxing BenCzechMark Bending Spoons Beneficial Deployments Benevity BenevolentAI Bengali Bengaluru BenHalluEval BenHalluScore Benjamin Chang Benjamin Krause Berkeley AI Research (BAIR) Berkeley Artificial Intelligence Research Berkshire Hathaway Berlin Database of Emotional Speech (EMO-DB) BerriAI BERT BERT-base BERT-F1 BERTopic BERTScore BERTurk BERTweet BES (trained models) Bessemer Venture Partners Best-of-N Best-of-N Sampling best@k Better Call GRPO Beyond a Single Judge: Simulating Social Persona Panels for Generative UI Evaluation Beyond Accuracy: Community Perspectives on Machine Translation Beyond Adam: SOAP and Muon for Faster, Label-Efficient Training of Machine Learning Interatomic Potentials Beyond Aggregate Risk: Role-Stratified Conformal Risk Control for LLM Tool Calls Beyond Attack-Success Rate: Action-Graded Severity Scale for Tool-Using AI Agents Beyond Benchmarks: Exposing the Hidden Crisis in Bangla Hate Speech Detection Beyond Fixed Representations: The Vocabulary and Verifier Gaps in Open-Ended AI Beyond Fully Random Masking: Attention-Guided Denoising and Optimization for Diffusion Language Models Beyond Function Calling: Benchmarking Tool-Using Agents under Tool-Environment Unreliability Beyond Global Replanning: Hierarchical Recovery for Cross-Device Agent Systems Beyond Independent Labels: Schwartz-Geometry Decoding for Human Value Detection Beyond Negative-Ridge Endpoints: Mixed-Sign Spectral Regularization via Negative-Shifted Gradient Descent Beyond Runtime Enforcement: Shield Synthesis as Defensibility Analysis for Adversarial Networks Beyond Scale and Generation: Understanding Language Model-based Entity Matching Beyond Sentiment: Structured Information Extraction from Financial News Beyond Success Rate: Cost-Aware Evaluation of Offensive and Defensive Security Agents Beyond Sufficiency: Time Series Explanation with Counterfactual Necessity Beyond Surface Forms: A Comprehensive, Mechanism-Oriented Taxonomy of Indirect Linguistic Encoding for LLM-Based Coded Language Detection Beyond Sycophancy: Structured Resistance and Compliance in LLM Moral Reasoning Beyond task performance: Decoding bioacoustic embeddings with speech features Beyond the Commitment Boundary: Probing Epiphenomenal Chain-of-Thought in Large Reasoning Models Beyond the Hard Budget: Sparsity Regularizers for More Interpretable Top-k Sparse Autoencoders Beyond the Leaderboard: Design Lessons for Trustworthy Multimodal VQA Beyond Third-Person Audits: Situated Interaction Auditing for User-Centered LLM Bias Research Beyond Tokenization: Direct Timestep Embedding and Contrastive Alignment for Time-Series Question Answering Beyond Uniform Tokens: Adaptive Compression for Time Series Language Models BeyondUncertainty BFCL BFCL Multi-Turn BFCL-V3 BFCLv3 BFW schedule hint BGE BGE-M3 BGL bi-manual robotic manipulation Bias Benchmark for Question Answering Bias Leaves a Gradient Trail Biconvex Optimization bidirectional attention Bidirectional Evolutionary Search Bielik Big Bench Big Bench Audio Big-Math BigBird BigCode BigCodeArena BigCodeBench Bigger or Cheaper? Scale and Quantization Effects on Uncertainty Signals in Vision-Language Models Under Image Degradation BigLaw Bench BigScience Bill & Melinda Gates Foundation Bill Dally Bill Gurley Bill McDermott Binarized Neural Network binary quantization BINEVAL Bing binning semiring bioacoustic monitoring bioacoustics BioASQ 14b BioBERT Biographical