返回
HuggingFace

Terminal-Universe: Turning Agent Trajectories into Scalable Terminal Environments

As terminal-based code agents become prevalent, agent trajectories have accumulated at scale, while realistic, execut...

AI 聚合 09/04
HuggingFace

Principia: Relational Physics Tests for Video Models

Evaluating physical reasoning in video models is difficult because absolute motion measurements depend on frame rate,...

AI 聚合 09/04
HuggingFace

Environment Evolution for Terminal Agents

Scaling interactive and verifiable environments is critical for training terminal agents. As frontier models become m...

AI 聚合 09/04
HuggingFace

Editable Visual Design

While diffusion base models such as GPT-Image-2 and Nano-Banana exhibit remarkable visual expressiveness, their end-t...

AI 聚合 09/04
HuggingFace

CORE: Improving Compositional Reasoning in MLLM Embedding via Reranker Distillation

MLLM-based embedding models remain limited in compositional retrieval, often failing to distinguish scenes containing...

AI 聚合 09/04
HuggingFace

Rethinking On-Policy Distillation of Large Language Models II: One Training Example

On-policy distillation (OPD) combines student-generated rollouts with dense token-level supervision from a teacher. E...

AI 聚合 09/04
HuggingFace

Random Attention: Rethinking KV Cache Eviction for Efficient Reasoning

Large language models achieve superior performance on tasks that require extended reasoning, but long chains of thoug...

AI 聚合 09/04
HuggingFace

WorldReward: Reward Modeling for Camera-Conditioned World Models

Camera-conditioned world models generate interactive videos in which commanded actions should induce the expected sce...

AI 聚合 09/04
HuggingFace

FlashRender: Few-Step Generative Rendering via Camera-Controlled Video MeanFlow

We present FlashRender, a few-step generative rendering framework that retakes a source video along a target camera t...

AI 聚合 09/04
HuggingFace

The Missing Temporal Link: Temporal Context Routing for Script-Driven Audio-Video Generation

Joint audio-video generation models have made substantial progress in visual quality and audio-visual synchronization...

AI 聚合 09/04
HuggingFace

Beyond Retrieval: Progressive Latent Memory Evolution for Streaming Video Understanding

Streaming video understanding requires multimodal large language models (MLLMs) to process continuous visual inputs a...

AI 聚合 09/04
HuggingFace

Puffin-World: Scaling a Unified Multimodal Model with Native 3D World States

We propose Puffin-World, a unified multimodal architecture that integrates physical understanding, spatial simulation...

AI 聚合 09/04
HuggingFace

Knowing When Not to Reuse: Conditional Experience Transfer in Autonomous LLM Post-Training

Large language models offer broad capabilities, but adapting them to evolving domains, tools, and requirements often ...

AI 聚合 09/04
HuggingFace

PACE: Towards Surfacing Hidden Conflicts in User Requests

Personalized assistants should not only comply with user requests but also assess whether those requests are appropri...

AI 聚合 09/04
HuggingFace

Percolation Dynamics in Optimization : Variance Cascades and Discrete Scale Invariance

We study the dynamics of Stochastic Gradient Descent (SGD), which is known to steer deep neural networks toward invar...

AI 聚合 09/04
HuggingFace

Compile by Training: Turning Natural-Language Specifications into Local Neural Functions

Many recurring text functions are easy to describe but difficult to implement with rules, while calling a large remot...

AI 聚合 09/04
HuggingFace

Let Confidence Change, Not the Prediction: Prediction-Preserving Repair for Post-hoc Calibration

Post-hoc calibration corrects reported confidence, yet a multiclass calibrator can also change the associated top-1 p...

AI 聚合 09/04
HuggingFace

Scal3R: Learning Efficient Multi-Relative Pose Query for Scalable Online 3D Reconstruction

Online 3D reconstruction models perform poorly on long videos. This happens because regressing poses relative to a fi...

AI 聚合 09/04
HuggingFace

Using Grounded Theory for Agent Behavior Analysis at Scale

Understanding agent behavior requires methods that scale to thousands of trajectories and surface new patterns in lon...

AI 聚合 09/04
HuggingFace

FoldingAgent: Inferring Parametric Origami Procedures from Demonstration Videos

We present FoldingAgent, an agentic framework for inferring explicit parametric folding programs directly from origam...

AI 聚合 09/04
HuggingFace

NeoMME: A Single-Tower Multimodal-Native Multilingual Foundation Encoder for Efficient Fine-Tuning and Inference

Multimodal models often build on architectures designed for generative vision-language modeling, typically combining ...

AI 聚合 09/04
HuggingFace

Wasserstein-Barycentric Interaction Fields for Spatial Factor Models: Evidence from Language-Model Representations

Spatial return models take the interaction matrix as given and leave feedback uninterpreted. We construct a bandwidth...

AI 聚合 09/04
HuggingFace

Portfolio Risk Bounds without Cross-Asset Return Covariances: Distributional Fields from Language-Model Representations

Portfolio risk assessment ordinarily relies on reliable estimates of cross-asset return covariances, which are diffic...

AI 聚合 09/04
HuggingFace

Debias-SparseGPT: Bias-Aware Pruning for Large Language Models

Model compression techniques such as pruning and quantization facilitate the efficient deployment and acceleration of...

AI 聚合 09/04
HuggingFace

Small Language Models as Judges for Rubric-Based Reinforcement Learning

Rubric-based reinforcement learning extends RL beyond tasks with exact answers or rule-based verifiers by scoring res...

AI 聚合 09/04
HuggingFace

WHALE: A Simple Recipe for Joint Harness-Weight Optimization

Agent performance depends jointly on the model parameters and the executable harness code that manages context and co...

AI 聚合 09/04
HuggingFace

Sparse Readout Prism: Explaining Logit-Lens Scores in Features Instead of Tokens

A language model's prediction of its next token develops across layers, and lens methods track this process by decodi...

AI 聚合 09/04
HuggingFace

An Empirical Study on Zero-Data Bootstrapping for Conversational Recommender Systems

Conversational Recommender Systems (CRS) typically require domain-specific dialogue data, which is costly, scarce, an...

AI 聚合 09/04
HuggingFace

Cliff: Learning Process Rewards from the First Mistake

Reinforcement learning with verifiable rewards (RLVR) has emerged as a powerful paradigm for large language model (LL...

AI 聚合 09/03
HuggingFace

Kirin: Animal Motion Generation from In-the-Wild Video

Understanding animal motion is fundamental to modeling animal behavior and biomechanics, yet progress in this area la...

AI 聚合 09/03
HuggingFace

Post-Training Language Models for Gold-Medal Performance in Coding Competitions

Competitive programming has become a key test of large language model reasoning, with international competitions such...

AI 聚合 09/03
HuggingFace

On the Design Fundamentals of Pixel Text Representation Learning

Text-rich visual inputs require models that can read, retrieve, and compress language directly in pixel space, yet ex...

AI 聚合 09/03
HuggingFace

Language Models Can Control Their Own Attention

Language models spend most of their attention on a small fraction of context, yet they read the entire KV cache to fi...

AI 聚合 09/03
HuggingFace

EarlyEval: Cheaper Agent Evaluation via Early Outcome Prediction

Evaluating LLM agents is essential for guiding their development, yet it has grown prohibitively expensive: a single ...

AI 聚合 09/03
HuggingFace

Repo-To-Skill: Distilling GitHub Repositories Into AI4AI Skills

Autonomous agents are beginning to carry out machine-learning (ML) research end to end. These agents combine a model ...

AI 聚合 09/03
HuggingFace

S3Gym: Can LLMs Turn Self-Testing and Self-Judging into Self-Improvement?

Large language models (LLMs) increasingly interact with external environments and accumulate substantial behavioral e...

AI 聚合 09/03
HuggingFace

Aspire: Can Models Self-Evolve from Vague Goals?

Many important forms of human learning begin with a vague goal, such as "become a better physicist" or "improve at re...

AI 聚合 09/03
HuggingFace

HarnessDev: Can LLMs Create and Evolve Their Own Agent Harness?

As agents move from research prototypes to deployed tools, their capability increasingly depends on model-external ex...

AI 聚合 09/03
HuggingFace

Institutional Newspapers Pipeline: Deriving billions of high quality tokens from historical newspapers

Historical newspapers are an abundant record of public life, but their dense, irregular and sometimes noisy layouts m...

AI 聚合 09/03
HuggingFace

CRISP: Cliff-awaRe Input-adaptive Sparse Prefilling with Structural-Mass-Motivated Routing

The attention prefilling phase of long-context LLM inference scales quadratically, making self-attention a severe com...

AI 聚合 09/03
HuggingFace

A Glance Is All You Need: Single-Pass Fine-Grained Image Captioning with SimLoss

An image may be worth a thousand words, but most captioning models describe it in only a few. Modern vision-language ...

AI 聚合 09/03
HuggingFace

VibeVoice-ASR-Streaming Technical Report

Traditional speaker-attributed ASR systems treated ASR and speaker diarization as two separate tasks. Recently, end-t...

AI 聚合 09/03
HuggingFace

SnapBench: Benchmarking Snap-and-Ask Multimodal Retrieval for Mobile Interactions

Mobile AI acts as a visual oracle, empowering users to snap a picture of something and ask for information. Snap-and-...

AI 聚合 09/03
HuggingFace

ZipTok3D: High-Fidelity 3D Tokenization with Compact Token Prefixes

Compact token sequences are essential for efficient 3D generation. However, existing 3D tokenizers typically organize...

AI 聚合 09/03
HuggingFace

Autoregressive Mosaics: Probing 2D Spatial Reasoning in Text-Only Language Models

Large language models (LLMs) trained only on text and code can sometimes generate programs that draw recognizable ima...

AI 聚合 09/03
HuggingFace

Beyond Visual Similarity: Entity-Aligned Retrieval for Knowledge-Based Visual Question Answering

Knowledge-Based Visual Question Answering (KB-VQA) relies on retrieving external information to answer queries involv...

AI 聚合 09/03
HuggingFace

Exploring Collaboration between a language and a non-language agent

LLMs are increasingly deployed as orchestrators that coordinate specialized subagents to solve complex tasks through ...

AI 聚合 09/03
HuggingFace

Ignorance or Incompetence? Constructing Knowledge-Gated, Verifiable Tasks for LLM Agents

Professional agent tasks often depend on conventions that are absent from public corpora, yet benchmarks rarely contr...

AI 聚合 09/03
HuggingFace

From Production Traffic to Post-Training: Building a Self-Hosted LLM That Covers the Corporate Request Mix

Data-residency constraints force enterprises to self-host LLMs, but continuous adoption of newer models without decom...

AI 聚合 09/03
HuggingFace

Learning Where Outcomes Change:Credit-Addressable Reasoning for Multimodal Geometry

Multimodal geometry reasoning requires VLMs to extract precise visual relations and preserve them through multi-step ...

AI 聚合 09/03