Terminal-Universe: Turning Agent Trajectories into Scalable Terminal Environments
As terminal-based code agents become prevalent, agent trajectories have accumulated at scale, while realistic, execut...
As terminal-based code agents become prevalent, agent trajectories have accumulated at scale, while realistic, execut...
Evaluating physical reasoning in video models is difficult because absolute motion measurements depend on frame rate,...
Scaling interactive and verifiable environments is critical for training terminal agents. As frontier models become m...
While diffusion base models such as GPT-Image-2 and Nano-Banana exhibit remarkable visual expressiveness, their end-t...
MLLM-based embedding models remain limited in compositional retrieval, often failing to distinguish scenes containing...
On-policy distillation (OPD) combines student-generated rollouts with dense token-level supervision from a teacher. E...
Large language models achieve superior performance on tasks that require extended reasoning, but long chains of thoug...
Camera-conditioned world models generate interactive videos in which commanded actions should induce the expected sce...
We present FlashRender, a few-step generative rendering framework that retakes a source video along a target camera t...
Joint audio-video generation models have made substantial progress in visual quality and audio-visual synchronization...
Streaming video understanding requires multimodal large language models (MLLMs) to process continuous visual inputs a...
We propose Puffin-World, a unified multimodal architecture that integrates physical understanding, spatial simulation...
Large language models offer broad capabilities, but adapting them to evolving domains, tools, and requirements often ...
Personalized assistants should not only comply with user requests but also assess whether those requests are appropri...
We study the dynamics of Stochastic Gradient Descent (SGD), which is known to steer deep neural networks toward invar...
Many recurring text functions are easy to describe but difficult to implement with rules, while calling a large remot...
Post-hoc calibration corrects reported confidence, yet a multiclass calibrator can also change the associated top-1 p...
Online 3D reconstruction models perform poorly on long videos. This happens because regressing poses relative to a fi...
Understanding agent behavior requires methods that scale to thousands of trajectories and surface new patterns in lon...
We present FoldingAgent, an agentic framework for inferring explicit parametric folding programs directly from origam...
Multimodal models often build on architectures designed for generative vision-language modeling, typically combining ...
Spatial return models take the interaction matrix as given and leave feedback uninterpreted. We construct a bandwidth...
Portfolio risk assessment ordinarily relies on reliable estimates of cross-asset return covariances, which are diffic...
Model compression techniques such as pruning and quantization facilitate the efficient deployment and acceleration of...
Rubric-based reinforcement learning extends RL beyond tasks with exact answers or rule-based verifiers by scoring res...
Agent performance depends jointly on the model parameters and the executable harness code that manages context and co...
A language model's prediction of its next token develops across layers, and lens methods track this process by decodi...
Conversational Recommender Systems (CRS) typically require domain-specific dialogue data, which is costly, scarce, an...
Reinforcement learning with verifiable rewards (RLVR) has emerged as a powerful paradigm for large language model (LL...
Understanding animal motion is fundamental to modeling animal behavior and biomechanics, yet progress in this area la...
Competitive programming has become a key test of large language model reasoning, with international competitions such...
Text-rich visual inputs require models that can read, retrieve, and compress language directly in pixel space, yet ex...
Language models spend most of their attention on a small fraction of context, yet they read the entire KV cache to fi...
Evaluating LLM agents is essential for guiding their development, yet it has grown prohibitively expensive: a single ...
Autonomous agents are beginning to carry out machine-learning (ML) research end to end. These agents combine a model ...
Large language models (LLMs) increasingly interact with external environments and accumulate substantial behavioral e...
Many important forms of human learning begin with a vague goal, such as "become a better physicist" or "improve at re...
As agents move from research prototypes to deployed tools, their capability increasingly depends on model-external ex...
Historical newspapers are an abundant record of public life, but their dense, irregular and sometimes noisy layouts m...
The attention prefilling phase of long-context LLM inference scales quadratically, making self-attention a severe com...
An image may be worth a thousand words, but most captioning models describe it in only a few. Modern vision-language ...
Traditional speaker-attributed ASR systems treated ASR and speaker diarization as two separate tasks. Recently, end-t...
Mobile AI acts as a visual oracle, empowering users to snap a picture of something and ask for information. Snap-and-...
Compact token sequences are essential for efficient 3D generation. However, existing 3D tokenizers typically organize...
Large language models (LLMs) trained only on text and code can sometimes generate programs that draw recognizable ima...
Knowledge-Based Visual Question Answering (KB-VQA) relies on retrieving external information to answer queries involv...
LLMs are increasingly deployed as orchestrators that coordinate specialized subagents to solve complex tasks through ...
Professional agent tasks often depend on conventions that are absent from public corpora, yet benchmarks rarely contr...
Data-residency constraints force enterprises to self-host LLMs, but continuous adoption of newer models without decom...
Multimodal geometry reasoning requires VLMs to extract precise visual relations and preserve them through multi-step ...