HuggingFace

Video-Oasis: Rethinking Evaluation of Video Understanding

The inherent complexity of video understanding makes it difficult to determine whether Video-LLM benchmark performanc...

AI 聚合
2026-07-10
HuggingFace

UniClawBench: A Universal Benchmark for Proactive Agents on Real-World Tasks

The rapid development of large language models and multimodal large language models has accelerated the emergence of ...

AI 聚合
2026-07-10
HuggingFace

Why Can't I Open My Drawer? Mitigating Object-Driven Shortcuts in Zero-Shot Compositional Action Recognition

Zero-Shot Compositional Action Recognition (ZS-CAR) requires recognizing novel verb-object combinations composed of p...

AI 聚合
2026-07-10
HuggingFace

A Quantized Native Runtime for On-Device Semantic Audio Generation

Semantic audio applications increasingly require controllable generation on commodity and embedded hardware rather th...

AI 聚合
2026-07-10
HuggingFace

Enhancing In-context Panoramic Generation via Geometric-aware Pretraining

In this work, we present Canvas360, a two-stage framework for in-context panoramic generation that combines geometry-...

AI 聚合
2026-07-10
HuggingFace

LongE2V: Long-Horizon Event-based Video Reconstruction, Prediction, and Frame Interpolation with Video Diffusion Models

Recovering high-quality video from sparse event streams is a challenging task. Regression methods often blur textures...

AI 聚合
2026-07-10
HuggingFace

DrugGen 2: A disease-aware language model for enhancing drug discovery

Current computational approaches for drug design typically focus on generating molecules conditioned on specific targ...

AI 聚合
2026-07-10
HuggingFace

Linear Attention Architectures: Mechanisms, Trade-offs, and Cross-Layer Routing

Self-attention lets each token retrieve information from the full context, but its quadratic cost in sequence length ...

AI 聚合
2026-07-10
HuggingFace

Automating the Design of Embodied Agent Architectures

Embodied agents are typically built as hand-designed compositions of perception, memory, planning, and action modules...

AI 聚合
2026-07-10
HuggingFace

Imagined Rollouts are Kinematic, Not Dynamic: A Diagnosis of Long-Horizon World-Model Failure

Long-horizon failure in world models is conventionally attributed to compounding error, a generic framing that does n...

AI 聚合
2026-07-10
HuggingFace

Sparse Delta Memory: Scaling the State of Linear RNNs through Sparsity

Linear attention models allow a fixed state size and a fixed amount of compute per token. However, due to their limit...

AI 聚合
2026-07-10
HuggingFace

Single-Rollout Asynchronous Optimization for Agentic Reinforcement Learning

Reinforcement learning (RL) is becoming increasingly important for post-training large language models (LLMs). Previo...

AI 聚合
2026-07-10
首页 上一页 第 142 / 216 页 下一页 末页