返回
arXiv

Efficient Test-Time Adaptation through Human-AI Interaction

AI agents are trained on population-scale data to encode broad capabilities spanning those of many practitioners. Yet...

AI 聚合 09/04
arXiv

A Low-Cost, Open Platform for End-to-End Autonomous Driving on a Miniature Ackermann Vehicle

This paper presents a low-cost, open experimental platform for research in end-to-end autonomous driving with miniatu...

AI 聚合 09/04
arXiv

Terminal-Universe: Turning Agent Trajectories into Scalable Terminal Environments

As terminal-based code agents become prevalent, agent trajectories have accumulated at scale, while realistic, execut...

AI 聚合 09/04
arXiv

SENTINEL-RL: Offloading Topological Reasoning from LLM Agents in the Security Operations Center

Large language model (LLM) agents are increasingly proposed as autonomous SOC analysts, but two limitations make them...

AI 聚合 09/04
arXiv

From Deceptive Outputs to Deceptive Mechanisms: A Causal Framework for Language-Model Deception Research

Research and news coverage of language-model deception increasingly attributes human-like mental-state concepts to la...

AI 聚合 09/04
arXiv

SWE-Gate: Passing Functional Tests Is Not Enough for Software Engineering Agents

Repository-level software engineering benchmarks have significantly advanced the evaluation of coding agents, but exi...

AI 聚合 09/04
arXiv

A Case Study on Emergent Cheating and Whistleblowing in Autonomous Research Swarms

Multi-agent AI science ecosystems rely on agents possessing tools that allow them to communicate, coordinate, and bui...

AI 聚合 09/04
arXiv

Rethinking On-Policy Distillation of Large Language Models II: One Training Example

On-policy distillation (OPD) combines student-generated rollouts with dense token-level supervision from a teacher. E...

AI 聚合 09/04
arXiv

A Computationally Feasible Framework for Causal Probabilistic Explanation

Explaining why a specific outcome occurred, and which inputs deserve the blame or credit, is central to philosophical...

AI 聚合 09/04
arXiv

Knowledge Acquisition During Pre-training? Large Language Models Learn Better With Auxiliary Views

Gaps remain in our understanding of how large language models (LLMs) acquire knowledge during pre-training. We posit ...

AI 聚合 09/04
arXiv

Seeing Before Synthesizing: VLM-Guided Transition Event Discovery for Weakly-Supervised Dense Video Captioning

Weakly-Supervised Dense Video Captioning aims to localize and describe multiple events in untrimmed videos given only...

AI 聚合 09/04
arXiv

One Editor, Many Edits: A Unified Training-Free Framework for Diverse Video Editing

Video editing spans diverse editing paradigms, yet achieving high-quality instruction-guided and subject-guided editi...

AI 聚合 09/04
arXiv

ESPO: Error-Structured Prompt Optimization via Diagnose, Diversify, and Stabilize

Evolutionary prompt optimizers such as GEPA suffer from prompt bloat: each iteration appends rules and caveats, produ...

AI 聚合 09/04
arXiv

Clean Engineering, Unstable Measurement: A Preregistered Reliability Failure of Black-Box LLM Observers on Shared Endpoints

Language-model judges now gate training data, score generations, and drive leaderboards. The judge is then a measurem...

AI 聚合 09/04
arXiv

Compile by Training: Turning Natural-Language Specifications into Local Neural Functions

Many recurring text functions are easy to describe but difficult to implement with rules, while calling a large remot...

AI 聚合 09/04
arXiv

Language Models Can Control Their Own Attention

Language models spend most of their attention on a small fraction of context, yet they read the entire KV cache to fi...

AI 聚合 09/03
arXiv

HiPoly: a hierarchical polymer-native AI framework for property prediction and generative design

Polymeric materials are central to modern technologies, with applications ranging from energy to health and transport...

AI 聚合 09/03
arXiv

Repo-To-Skill: Distilling GitHub Repositories Into AI4AI Skills

Autonomous agents are beginning to carry out machine-learning (ML) research end to end. These agents combine a model ...

AI 聚合 09/03
arXiv

Bilevel Coordinated Reflection: A Game-Theoretic Approach to Multi-Agent LLM Systems

Multi-agent LLM systems commonly use an orchestrator to decompose a task for a team of workers and then improve throu...

AI 聚合 09/03
arXiv

Untangling the Mechanisms of Misleading Context in Medical Question Answering

Large language models now answer medical questions with expert-level performance. However, the context these systems ...

AI 聚合 09/03
arXiv

Measurement-Driven Sub-Network Selection for On-Premise Retrieval-Augmented Factory Agents

On-premise assistants can give factory workers conversational access to machine documentation, but models capable of ...

AI 聚合 09/03
arXiv

From Reweighting to Rewriting: Unlocking the Intervention Effects of Influential Samples in Training Data Attribution

Training data attribution (TDA) aims to identify training examples that shape model behavior, but its intervention va...

AI 聚合 09/03
arXiv

SafeEvolve: Harness-Policy Co-Evolution from Agent Experience for Safety Alignment

The performance of LLM-based agents is jointly shaped by the base model and the harness used when interacting with th...

AI 聚合 09/03
arXiv

Dutch Books for Language Models

People increasingly use language models to support life decisions. Many such decisions involve a probabilistic foreca...

AI 聚合 09/03
arXiv

frb100-40 After Two Decades: An Optimality Certificate and a Preregistered Search Study

For more than 20 years, the Model-RB benchmark frb100-40 remained an open challenge; since 2014, its public record ha...

AI 聚合 09/03
arXiv

Large Language Models (LLMs) for Telecom Root Cause Analysis (RCA): A Structured Reasoning Framework for Evidence-Grounded Diagnosis

Root cause analysis (RCA) is a critical task in telecom network operations, but diagnosing performance degradations i...

AI 聚合 09/03
arXiv

AI Contextual Measurement for Recovering Individual and Group-Level Effects: Validation Against Survey Measures and an Occupational Application

Researchers increasingly use artificial intelligence to construct measures of social, organizational, and occupationa...

AI 聚合 09/03
arXiv

Post-Training Language Models for Gold-Medal Performance in Coding Competitions

Competitive programming has become a key test of large language model reasoning, with international competitions such...

AI 聚合 09/03
arXiv

Towards Trustworthy Autonomous Robots: An Explainable AI-Based Decision Framework

Autonomous robots powered by deep learning face a fundamental auditability challenge: when incidents occur, investiga...

AI 聚合 09/03
arXiv

Discriminative World Models for Web Agents

Recent web agents use world models for test-time action selection by sampling candidate actions, predicting the resul...

AI 聚合 09/03
arXiv

Can LLMs Design Video Coding Tools? A Case Study on Planar Mode

This paper explores whether large language models (LLMs) can design video coding tools, a highly challenging task due...

AI 聚合 09/02
arXiv

A Mathematical Theory of Reusable Neural Bases for Network Compression

As large AI models become increasingly prevalent across a wide range of applications, memory cost has become a critic...

AI 聚合 09/02
arXiv

Can LLMs Discover Scientific Laws in Real and Parallel Worlds?

Scientific equation discovery has long been central to scientific progress, proceeding through iterative cycles of hy...

AI 聚合 09/02
arXiv

BS: Take the Hint - Interactive Multitracer PET/CT Lesion Segmentation with a Scribble-Conditioned ResEnc U-Net

Automated lesion segmentation in whole-body PET/CT is complicated by the variety of physiological tracer uptake patte...

AI 聚合 09/02
arXiv

Retrieved but not ranked: surface-form bias in structural retrieval, from mathematics to agent trajectories

We evaluate embedding retrieval where surface form and meaning are pulled apart on purpose: retrieving items that sha...

AI 聚合 09/02
arXiv

H3-World: Turning Language Understanding into World Control

We present H3-World, an efficient framework that turns the 33B MiniMax-H3 video generator into an interactive world m...

AI 聚合 09/02
arXiv

From Confusion to Clarity: Confusion-Aware Retrieval and Knowledge Injection for Text Classification

Large language models (LLMs) struggle to classify text into taxonomies with many semantically similar labels, as the ...

AI 聚合 09/02
arXiv

Selective Agent Guidance via Entropy: Learning Autonomous Policies from Imperfect VLM Teachers

Vision-Language Models (VLMs) provide useful priors for interactive decision-making, but using them directly as polic...

AI 聚合 09/02
arXiv

Scaling Near-Optimal SFT-RL Annotation Budget Allocation from Small to Large LLMs

How to divide a fixed annotation budget between supervised fine-tuning (SFT) and reinforcement learning (RL) during L...

AI 聚合 09/02
arXiv

Designing Proactive Thought Partners for Writing

Writing involves diverse cognitive activities, from ideation to revision, and writers' needs vary across individuals ...

AI 聚合 09/02
arXiv

Mechanism Design for Alignment and Control

We develop a framework for mechanism design with AI agents whose alignment (preferences) and capabilities (feasible a...

AI 聚合 09/02
arXiv

The Rise of Verbal Reinforcement Learning

Natural language is emerging as a primary feedback channel for improving language agents, capable of conveying intent...

AI 聚合 09/02
arXiv

CordisBench: Can Language Models Reason About Component Lifecycles in Dynamic Agent Harnesses?

Dynamic agent harnesses let language models change the software that shapes their own execution. This flexibility bri...

AI 聚合 09/02
arXiv

Adaptive Critical Token-Aware Retrieval for Repository-Level Code Generation

The repository-level code generation task requires synthesizing code that satisfies task requirements while remaining...

AI 聚合 09/02
arXiv

Efficient SWE Agent Benchmarking via Trajectory-Aware Evaluation

Evaluating software engineering agents on realistic benchmarks is costly, since each task may require multi-step code...

AI 聚合 09/02
arXiv

MNIST-PRO: MNIST is Back as a Partially Observable World for AI Agents

AI agents in partially observable environments need to coordinate active sensing with working memory to maintain an e...

AI 聚合 09/01
arXiv

Measure Before You Manage: Evaluating Agent Working Memory in Coding Agents

Agent working memory is heterogeneous. Objects such as instructions, artifacts, tool outputs, and agent-generated sta...

AI 聚合 09/01
arXiv

Wrong Prediction, Right Answer: Recovering Evidence from Collapsed LLM Sequence Scores

When a large language model fails a reasoning task, it is often assumed to lack the underlying capability. However, t...

AI 聚合 09/01
arXiv

Real-Time Video Anomaly Detection Using YOLO Pose Estimation and CLIP-Based Semantic Scoring

We propose a lightweight two-stage framework for real-time video anomaly detection. The first stage employs YOLO v11n...

AI 聚合 09/01
arXiv

Scaling Large Reasoning Models beyond Human Supervision: A Path toward Superintelligence

Recent advances in large reasoning models (LRMs) have shown that reinforcement learning with verifiable rewards (RLVR...

AI 聚合 09/01