HuggingFace

The Missing Temporal Link: Temporal Context Routing for Script-Driven Audio-Video Generation

Joint audio-video generation models have made substantial progress in visual quality and audio-visual synchronization...

AI 聚合
2026-09-04
HuggingFace

Beyond Retrieval: Progressive Latent Memory Evolution for Streaming Video Understanding

Streaming video understanding requires multimodal large language models (MLLMs) to process continuous visual inputs a...

AI 聚合
2026-09-04
HuggingFace

Puffin-World: Scaling a Unified Multimodal Model with Native 3D World States

We propose Puffin-World, a unified multimodal architecture that integrates physical understanding, spatial simulation...

AI 聚合
2026-09-04
HuggingFace

Knowing When Not to Reuse: Conditional Experience Transfer in Autonomous LLM Post-Training

Large language models offer broad capabilities, but adapting them to evolving domains, tools, and requirements often ...

AI 聚合
2026-09-04
HuggingFace

PACE: Towards Surfacing Hidden Conflicts in User Requests

Personalized assistants should not only comply with user requests but also assess whether those requests are appropri...

AI 聚合
2026-09-04
HuggingFace

Percolation Dynamics in Optimization : Variance Cascades and Discrete Scale Invariance

We study the dynamics of Stochastic Gradient Descent (SGD), which is known to steer deep neural networks toward invar...

AI 聚合
2026-09-04
HuggingFace

Compile by Training: Turning Natural-Language Specifications into Local Neural Functions

Many recurring text functions are easy to describe but difficult to implement with rules, while calling a large remot...

AI 聚合
2026-09-04
HuggingFace

Let Confidence Change, Not the Prediction: Prediction-Preserving Repair for Post-hoc Calibration

Post-hoc calibration corrects reported confidence, yet a multiclass calibrator can also change the associated top-1 p...

AI 聚合
2026-09-04
HuggingFace

Scal3R: Learning Efficient Multi-Relative Pose Query for Scalable Online 3D Reconstruction

Online 3D reconstruction models perform poorly on long videos. This happens because regressing poses relative to a fi...

AI 聚合
2026-09-04
HuggingFace

Using Grounded Theory for Agent Behavior Analysis at Scale

Understanding agent behavior requires methods that scale to thousands of trajectories and surface new patterns in lon...

AI 聚合
2026-09-04
HuggingFace

FoldingAgent: Inferring Parametric Origami Procedures from Demonstration Videos

We present FoldingAgent, an agentic framework for inferring explicit parametric folding programs directly from origam...

AI 聚合
2026-09-04
HuggingFace

NeoMME: A Single-Tower Multimodal-Native Multilingual Foundation Encoder for Efficient Fine-Tuning and Inference

Multimodal models often build on architectures designed for generative vision-language modeling, typically combining ...

AI 聚合
2026-09-04
首页 上一页 第 3 / 211 页 下一页 末页