Efficient Test-Time Adaptation through Human-AI Interaction
AI agents are trained on population-scale data to encode broad capabilities spanning those of many practitioners. Yet...
AI agents are trained on population-scale data to encode broad capabilities spanning those of many practitioners. Yet...
This paper presents a low-cost, open experimental platform for research in end-to-end autonomous driving with miniatu...
As terminal-based code agents become prevalent, agent trajectories have accumulated at scale, while realistic, execut...
Large language model (LLM) agents are increasingly proposed as autonomous SOC analysts, but two limitations make them...
Research and news coverage of language-model deception increasingly attributes human-like mental-state concepts to la...
Repository-level software engineering benchmarks have significantly advanced the evaluation of coding agents, but exi...
Multi-agent AI science ecosystems rely on agents possessing tools that allow them to communicate, coordinate, and bui...
On-policy distillation (OPD) combines student-generated rollouts with dense token-level supervision from a teacher. E...
Explaining why a specific outcome occurred, and which inputs deserve the blame or credit, is central to philosophical...
Gaps remain in our understanding of how large language models (LLMs) acquire knowledge during pre-training. We posit ...
Weakly-Supervised Dense Video Captioning aims to localize and describe multiple events in untrimmed videos given only...
Video editing spans diverse editing paradigms, yet achieving high-quality instruction-guided and subject-guided editi...
Evolutionary prompt optimizers such as GEPA suffer from prompt bloat: each iteration appends rules and caveats, produ...
Language-model judges now gate training data, score generations, and drive leaderboards. The judge is then a measurem...
Many recurring text functions are easy to describe but difficult to implement with rules, while calling a large remot...
Language models spend most of their attention on a small fraction of context, yet they read the entire KV cache to fi...
Polymeric materials are central to modern technologies, with applications ranging from energy to health and transport...
Autonomous agents are beginning to carry out machine-learning (ML) research end to end. These agents combine a model ...
Multi-agent LLM systems commonly use an orchestrator to decompose a task for a team of workers and then improve throu...
Large language models now answer medical questions with expert-level performance. However, the context these systems ...
On-premise assistants can give factory workers conversational access to machine documentation, but models capable of ...
Training data attribution (TDA) aims to identify training examples that shape model behavior, but its intervention va...
The performance of LLM-based agents is jointly shaped by the base model and the harness used when interacting with th...
People increasingly use language models to support life decisions. Many such decisions involve a probabilistic foreca...
For more than 20 years, the Model-RB benchmark frb100-40 remained an open challenge; since 2014, its public record ha...
Root cause analysis (RCA) is a critical task in telecom network operations, but diagnosing performance degradations i...
Researchers increasingly use artificial intelligence to construct measures of social, organizational, and occupationa...
Competitive programming has become a key test of large language model reasoning, with international competitions such...
Autonomous robots powered by deep learning face a fundamental auditability challenge: when incidents occur, investiga...
Recent web agents use world models for test-time action selection by sampling candidate actions, predicting the resul...
This paper explores whether large language models (LLMs) can design video coding tools, a highly challenging task due...
As large AI models become increasingly prevalent across a wide range of applications, memory cost has become a critic...
Scientific equation discovery has long been central to scientific progress, proceeding through iterative cycles of hy...
Automated lesion segmentation in whole-body PET/CT is complicated by the variety of physiological tracer uptake patte...
We evaluate embedding retrieval where surface form and meaning are pulled apart on purpose: retrieving items that sha...
We present H3-World, an efficient framework that turns the 33B MiniMax-H3 video generator into an interactive world m...
Large language models (LLMs) struggle to classify text into taxonomies with many semantically similar labels, as the ...
Vision-Language Models (VLMs) provide useful priors for interactive decision-making, but using them directly as polic...
How to divide a fixed annotation budget between supervised fine-tuning (SFT) and reinforcement learning (RL) during L...
Writing involves diverse cognitive activities, from ideation to revision, and writers' needs vary across individuals ...
We develop a framework for mechanism design with AI agents whose alignment (preferences) and capabilities (feasible a...
Natural language is emerging as a primary feedback channel for improving language agents, capable of conveying intent...
Dynamic agent harnesses let language models change the software that shapes their own execution. This flexibility bri...
The repository-level code generation task requires synthesizing code that satisfies task requirements while remaining...
Evaluating software engineering agents on realistic benchmarks is costly, since each task may require multi-step code...
AI agents in partially observable environments need to coordinate active sensing with working memory to maintain an e...
Agent working memory is heterogeneous. Objects such as instructions, artifacts, tool outputs, and agent-generated sta...
When a large language model fails a reasoning task, it is often assumed to lack the underlying capability. However, t...
We propose a lightweight two-stage framework for real-time video anomaly detection. The first stage employs YOLO v11n...
Recent advances in large reasoning models (LRMs) have shown that reinforcement learning with verifiable rewards (RLVR...