HuggingFace

VGI-BENCH: Probing Visual Intelligence in Video Generation Models

Recent studies suggest that video generation models can exhibit certain forms of zero-shot visual reasoning through g...

AI 聚合
2026-08-27
HuggingFace

JIT-Agent: Scaling Harness Intelligence via Just-in-Time Harness Evolution

Agent capability is not determined by the model alone. The agent harness, encompassing memory management, planning st...

AI 聚合
2026-08-27
HuggingFace

V-Rubrics: Visual Faithfulness via Rubric-Based Reinforcement Learning

Vision-language models can produce fluent answers that are insufficiently grounded in the visual evidence: a single u...

AI 聚合
2026-08-27
HuggingFace

Open-MOPD: Diagnosing and Fixing Capability Imbalance in Multi-Teacher On-Policy Distillation

Multi-teacher on-policy distillation (M-OPD) has emerged as a promising paradigm for consolidating domain-specialized...

AI 聚合
2026-08-27
HuggingFace

StreamPI: Streaming Multimodal Temporal Modeling for Vision-Language-Action Models

Vision-Language-Action (VLA) models have demonstrated effectiveness in robot manipulation, yet state-of-the-art model...

AI 聚合
2026-08-27
HuggingFace

MA-VLA: Multi-Arm Vision-Language-Action Model for Collaboration and Compositional Generalization

Multi-arm collaboration is becoming a core capability in embodied manipulation. Recent vision-language-action (VLA) m...

AI 聚合
2026-08-27
HuggingFace

WarpSAC: Towards the Pinnacle of Scalable Off-policy RL by Rethinking Exploration and Exploitation

Massively parallel simulation changes the data regime in which off-policy reinforcement learning (RL) is trained, cha...

AI 聚合
2026-08-27
HuggingFace

Gated Recurrent Transformers: Expressive Depth through Recurrent Modulation in Transformers

Scaling transformer language models creates an inherent tension between expressivity and memory efficiency. While uni...

AI 聚合
2026-08-27
HuggingFace

VoiceMem: Streaming Dual-Brain Memory for Real-Time Interaction

Conversational systems, such as duplex speech language models (SLMs), still lack a streaming, accurate, and empatheti...

AI 聚合
2026-08-27
HuggingFace

Stream4D: 4D-Consistency for Streaming Autoregressive Diffusion Video Models

Streaming autoregressive diffusion models enable real-time, long-horizon video generation, but their training objecti...

AI 聚合
2026-08-27
HuggingFace

Rubrics as Visual-Repair Context for Self-Evolving UI-to-Code Generation

Large vision-language models have shown strong progress in UI-to-code generation, yet their test-time self-evolution ...

AI 聚合
2026-08-27
HuggingFace

FIRM-Video: Check Before You Score for Reliable Text-to-Video Reward Modeling

Reliable reward models are essential for text-to-video evaluation and alignment. However, the trade-off between evalu...

AI 聚合
2026-08-27
首页 上一页 第 24 / 212 页 下一页 末页