4arXiv cs.CL (Computation and Language)·6d ago

Task decomposition framework for reducing inferential load in structured annotation

A new arXiv preprint proposes decomposing complex structured annotation tasks into sub-tasks to reduce aggregate inferential load across heterogeneous annotator pools that mix human experts and models. The authors introduce a formal model of inferential load grounded in centering theory, using 'centers' (salient anchor entities) to constrain output space complexity. They provide decomposition guidelines and a budget-aware sub-task allocation procedure, with cost-efficiency gains demonstrated from prior work.

Evaluation and Benchmarking centering theory Task Decomposition for Efficient Annotation

Related guides (1)

Evaluation and BenchmarkingTopic guide

Evaluation and Benchmarking: How We Measure AI — and Why It Keeps Getting Harder

Read asBeginner In-depth

Related events (8)

6arXiv · cs.CL·7d ago·source ↗

SelfCompact: Model-driven adaptive context compaction for long agent traces

Researchers propose SelfCompact, a scaffold that lets language models decide when and how to compact their own accumulated context during long agentic runs, rather than relying on fixed token-threshold triggers. The system pairs a compaction tool with a lightweight rubric specifying when to invoke or suppress compaction based on trajectory structure (e.g., sub-task completion vs. mid-derivation). Evaluated across six benchmarks and seven models, SelfCompact matches or exceeds fixed-interval summarization while reducing per-question token cost by 30-70%, with gains of up to 18.1 points on math tasks and 5-9 points on agentic search. The work identifies a 'meta-cognitive gap' in unprompted models and shows it can be closed via scaffolding without fine-tuning.

Long Context Evolution Inference Economics SelfCompact Self-Compacting Language Model Agents +1 more

5arXiv · cs.CL·1mo ago·source ↗

Cross-Annotator Preference Optimization (CAPO) for Learning Annotator-Specific Explanation Behavior

This paper investigates whether LLMs can learn and reproduce individual annotator-specific reasoning patterns, not just label choices, using two sentence-pair tasks (NLI and paraphrase judgment) with four annotators each. The authors find that annotator-specific patterns are weak at the single-annotation level but detectable after aggregation, and propose CAPO—a preference optimization method that contrasts a target annotator's response against other valid but less target-specific annotations. CAPO outperforms prompting and supervised fine-tuning baselines in capturing annotator-specific label-explanation behavior. The work suggests a path toward scalable annotation pipelines grounded in annotator histories rather than labels alone.

Evaluation and Benchmarking Alignment and RLHF Cross-Annotator Preference Optimization (CAPO)Human Label Variation (HLV)Natural Language Inference +2 more

5arXiv · cs.LG·11d ago·source ↗

Multi-Task Bayesian In-Context Learning for Amortized Hierarchical Inference

A new arXiv preprint introduces a multi-task in-context learning framework for amortized hierarchical Bayesian predictive inference, representing prior information as a prefix of in-context datasets fed to a transformer. The model learns to adapt predictions across families of priors, addressing the brittleness of prior-data fitted models under distribution shift. On evaluations including out-of-meta-distribution priors and high-dimensional latent structures, the method matches oracle Bayesian predictors while being orders of magnitude faster, with a real-world spatiotemporal temperature prediction demonstration.

Evaluation and Benchmarking Multi-Task Bayesian In-Context Learning Prior-Data Fitted Networks

6arXiv · cs.LG·18d ago·source ↗

Task exchangeability framework enables statistically valid inference from synthetic data

A new arXiv preprint proposes a statistical framework for using synthetic data in scientific research with provable validity guarantees, centered on a condition called 'task exchangeability.' The framework requires identifying historical tasks with real data that are exchangeable with the current task of interest, enabling valid inference even when synthetic data is biased or misspecified. The authors demonstrate the approach on LLM-generated 'silicon samples' for public opinion surveys and LLM-as-a-judge AI evaluation settings. This addresses a foundational concern about the reliability of synthetic data pipelines increasingly used across AI evaluation and scientific research.

Evaluation and Benchmarking AI Safety Research Valid Inference with Synthetic Data via Task Exchangeability task exchangeability

7arXiv · cs.CL·18d ago·source ↗

Research identifies 'commitment boundary' in chain-of-thought reasoning, enabling 55% CoT length reduction

A new arXiv preprint introduces the concept of a 'commitment boundary' in chain-of-thought reasoning — a sharp transition point where a model's answer stabilizes, after which subsequent reasoning steps are 'epiphenomenal' and causally inert. The authors use early-exit probing and attention probes to detect this boundary, finding it can be linearly decoded from intermediate steps and generalizes across tasks. Exploiting this signal to exit reasoning blocks at the commitment boundary reduces CoT length by up to 55% on average with negligible performance loss, with direct implications for inference efficiency in large reasoning models.

Frontier Model Releases Evaluation and Benchmarking Chain-of-Thought Reasoning Beyond the Commitment Boundary: Probing Epiphenomenal Chain-of-Thought in Large Reasoning Models +1 more

4arXiv · cs.AI·1mo ago·source ↗

Structure-Aware Code Change Labeling with LLMs via Two-Stage Taxonomy Pipeline

This paper presents a systematic study of using LLMs for taxonomy-based labeling of code diff hunks, going beyond summarization to assign structured labels capturing semantic attributes like renames, moves, and logic modifications. The authors introduce a two-stage pipeline combining diff-hunk labeling with structural refinement, using few-shot prompting to remain language-agnostic. Evaluated across four LLMs on a curated benchmark of natural and synthetic patches, the best configuration achieves 84% recall and 81% precision. Results suggest LLM-based structured labeling can complement static analysis tools in code review workflows.

Enterprise Deployment Patterns Agent and Tool Ecosystem few-shot prompting code review automation diff hunk taxonomy benchmark +1 more

5Openai Blog·1mo ago·source ↗

Summarizing Books with Human Feedback

OpenAI published research on using human feedback to train models to summarize entire books, addressing the challenge of scaling human oversight to tasks that are difficult for humans to evaluate directly. The work explores recursive task decomposition, where models summarize smaller chunks and then summarize those summaries, with humans providing feedback at each level. This represents an early concrete application of scalable oversight techniques to long-document understanding.

Long Context Evolution AI Safety Research Recursive Summarization Reinforcement Learning from Human Feedback OpenAI +2 more

5arXiv · cs.CL·19d ago·source ↗

Doc-to-Atom: Compositional parametric memory via semantically typed micro-LoRA adapters

Doc-to-Atom (Doc2Atom) proposes a framework that decomposes documents into semantically typed knowledge atoms, each compiled into an independent micro-LoRA adapter with a retrieval key. At inference, a lightweight query router assembles only relevant atoms into a query-specific adapter injected into a frozen base model, addressing the irrelevant-query interference and scalability problems of monolithic adapter approaches like Doc-to-LoRA. The system is trained end-to-end via multi-objective distillation and outperforms Doc-to-LoRA baselines on six QA benchmarks while reducing memory cost.

Long Context Evolution Inference Economics Doc-to-LoRA Doc-to-Atom LoRA